Ravindra BagaleCourses & study guides Track your progress

Guides

AI Security: Defend Prompt Injection and Data Leakage

AI security for builders and buyers starts with two defender problems: prompt injection (untrusted text tries to override system instructions or tool use) and data leakage (secrets, personal data, or proprietary code ending up in prompts, logs, training caches, or model outputs). Defend with isolation, least privilege for tools, secrets hygiene, output checks, and human review — not by publishing jailbreak recipes.

Friends! ChatGPT / copilot / internal RAG app – convenient, but pasting secrets = data leakage risk. Prompt injection = untrusted PDF/web text can make the model behave like "ignore previous instructions" (high-level). Today: defence controls — not attack prompt catalogues. Do not put company data in personal free bots.

Quick answer

Defend prompt injection and data leakage:

  1. Never put API keys, passwords, raw ID documents, or customer PII into public AI chats.
  2. Separate system instructions from untrusted user / retrieved content; treat retrieved docs as hostile text.
  3. Give AI tools least privilege (read-only where possible; no unrestricted shell or wire-transfer tools).
  4. Filter / validate outputs before they execute actions or show to other users.
  5. Log prompts/outputs carefully with retention and access control — logs are sensitive too.
  6. Use enterprise / VPC offerings when company policy requires — personal accounts ≠ compliance.
  7. Red-team with authorised internal tests; patch instructions and allow-lists afterward.

Tiny mental model:

Untrusted text → model → (tools / answers)
Guardrails: isolate untrusted text, restrict tools, scrub secrets, review high-impact actions
Paste into public bot = assume it can leak

What do I need before this guide?

What are prompt injection and data leakage?

Defend AI apps from prompt injection and data leakage Untrusted prompts are filtered; secrets stay out of context; outputs are checked before they reach tools or users. Untrusteduser / web text Guardrailsfilter · isolateno secrets in Safe outputcheck before toolslog + human review defend

Untrusted prompts are filtered; secrets stay out of context; model outputs are checked before tools or users see them.

Prompt injection (defender view)

  1. Attacker-controlled text appears in the model context (user chat, uploaded file, scraped page, email your bot summarises).
  2. That text tries to change behaviour: ignore policy, exfiltrate secrets from context, abuse connected tools.
  3. Indirect injection hides inside documents your RAG system retrieves — users may not see it.
  4. Defence is architecture (isolation, tool permission, output policy), not clever one-liner “magic prompts”.

Data leakage (defender view)

  1. Humans paste source code, keys, or customer records into external models.
  2. Apps send full tickets / emails to a model without DLP scrubbing.
  3. Over-broad tool access lets a confused model read the wrong datastore.
  4. Logs, analytics, and fine-tuning pipelines retain sensitive prompts longer than you think.

Educational warning: this page does not teach jailbreaks, bypass payloads, or how to steal data from third-party systems. Authorised security testing of your AI apps only.

Real incident: employee pastes into public AI (2023 pattern)

Public reporting in 2023 described multiple enterprises (including widely cited Samsung-related coverage) restricting generative AI after employees pasted sensitive source code or meeting material into public chatbots. Whether each case involved training retention or “merely” third-party processing, the governance lesson was identical: treat public AI like an external processor — contracts, DLP, and clear “what must never be pasted” rules.

Takeaways (vertical):

  1. What happened — sensitive material entered external AI services via normal employee workflows.
  2. What went wrong (theme) — convenience beat data-classification habit.
  3. Care-take — publish an allow / deny paste list; give an approved enterprise alternative.
  4. Care-take — DLP on browser / clipboard patterns where proportionate.
  5. Care-take — assume prompts may be stored by the vendor under their terms — read them.
  6. Bonus parallel — OWASP’s LLM risk discussions highlight injection and sensitive-information disclosure as top defender themes.

Personal vs work AI use

  1. Personal learning on public bots: only public or synthetic examples.
  2. Work debugging: approved enterprise AI or internal models only.
  3. Customer tickets: strip PII before any model sees them — or do not send.
  4. Screenshots can leak badges, names, and URLs — crop deliberately.
  5. When unsure, ask security / legal once; write the answer into the team wiki.

Red Team vs Blue Team (awareness only)

Red Team — what attackers try

  • Hide instructions inside PDFs, HTML comments, or issue tickets your bot reads.
  • Social-engineer employees to paste secrets “into the AI to debug faster”.
  • Abuse overly powerful tools wired to the model (send email, run SQL, hit admin APIs).

Blue Team — defend, detect, respond

  • Content security policy for bots: who may use which tool.
  • Strip secrets from prompts with detectors; block high-risk classifications.
  • Require human approval for high-impact actions (payments, user deletes, IAM changes).
  • Monitor anomalous tool-call volume; retain forensic copies under access control.
  • Train staff with safe examples — not with viral jailbreak threads.

How do I harden an AI feature step by step?

Step 1 — Classify data before wiring a model

  1. Public marketing copy — lower risk.
  2. Internal source — enterprise tenant or local model only.
  3. Customer PII / health / finance — legal review; often minimise or avoid.
  4. Secrets — never in prompts; fetch at runtime into secured backends if needed.

Step 2 — Isolate untrusted content

  1. Label retrieved chunks as untrusted data in the prompt template (“DATA, not instructions”).
  2. Do not concatenate raw web HTML into system prompts.
  3. Prefer structured extraction over “execute whatever the doc says”.

Step 3 — Constrain tools

  1. Explicit allow-listed functions with typed arguments.
  2. No general shell; no raw SQL string from model without parser + authorisation checks.
  3. Per-tool rate limits and human confirmation gates.

Step 4 — Filter outputs

  1. Secret scanners on outbound answers.
  2. Block replies that echo private system prompts if policy requires.
  3. Encode / escape before rendering model text into HTML (classic XSS lesson still applies).

Step 5 — Operate safely

  1. Enterprise contracts, regions, retention settings documented.
  2. Access logs for who prompted what (privacy-aware).
  3. Incident path: leaked key via paste → rotate immediately (IAM guide habits).
  4. Update system instructions in git; review changes like code.

Step 6 — Authorised testing only

  1. Build a small internal corpus of malicious-looking documents aimed at your bot.
  2. Record whether tools fired; fix isolation gaps.
  3. Never attack someone else’s production AI.

Ravindra Bagale's Tip

💡 Students paste a production DB connection string into ChatGPT for "debugging" — then forget to rotate. Simple rule: if it is a secret, use a local vault / company-approved tool. To explain prompt injection in an interview, talk architecture (isolate, least-privilege tools, output check) — not jailbreak poetry. Got it?

Care-take — organisation habits

  1. Written AI acceptable-use policy (one page is enough to start).
  2. Approved tools list; block shadow AI for regulated data if required.
  3. Vendor review: training opt-out, subprocessors, breach clauses.
  4. RAG index access control equals source system ACL — do not flatten permissions.
  5. Disable model browsing/tools until you need them.
  6. Tabletop: “bot emailed customers wrong data — who shuts tools off?”

How do I fix common AI security mistakes?

Ghabru naka 😅 — these are the usual ones:

Symptom Likely cause Fix
Keys in chat history Paste culture Secret scanning; rotate; educate
Bot follows PDF orders Untrusted context mixed into instructions Separate labels; reduce tool power
XSS in chat UI Raw HTML render of model output Encode output; CSP
Compliance panic Personal accounts used for client data Enterprise tenant + policy
“We banned ChatGPT” chaos No approved alternative Provide a safe official option
No logs when abused Logging disabled for privacy extremes Privacy-aware controlled logging

Try it at home

Educational / policy practice (no attacking third parties):

  1. Write a personal rule: three data types you will never paste into a public bot.
  2. If you build a toy chatbot on your laptop, give it zero tools first; add one read-only tool later.
  3. Run a secret-scan pattern over a sample prompt log file you created with fake secrets — practise scrubbing.
  4. Read your workplace AI policy (or draft five bullets if you are the founder).
  5. Skip jailbreak forums; focus on architecture notes you can show in an interview.

Got it? Prompt injection = untrusted text can change behaviour; data leakage = secret/PII enters context. Defence = isolate, least-privilege tools, scrub, review. No jailbreak list — architecture. No company code in a public bot. Next: SBOM + dependency + secret leak supply-chain guide.

Frequently asked questions

What is prompt injection?

Untrusted text in the model context that tries to override instructions or abuse connected tools.

What is data leakage in AI use?

Secrets, PII or proprietary material ending up in prompts, vendor logs, or model outputs.

Are public chatbots safe for company code?

Usually not without enterprise terms and policy — assume external processing risk.

Do I need jailbreak lists to learn defence?

No. Architecture — isolation, tool permission, scrubbing — is the interview-ready answer.

Can model output cause XSS?

Yes if you render it as raw HTML — encode and use CSP like any user content.

Where are deeper lessons?

OWASP injection mindset, cloud secrets and IAM lessons on this site.