AI security 1 min read
Prompt Injection Survival Guide for Builders
A short, practical checklist to keep your agent endpoints from becoming a screenshot of shame.
Guardrails
- Sandwich prompts: system preamble + postamble that cannot be overridden.
- Strict tools: validate arguments; whitelist functions and schema.
- No raw tool output to users; sanitize and summarize.
- Rate limit + auth every agent endpoint; log inputs/outputs for forensics.
Quick detections
- High-perplexity inputs; known jailbreak keywords.
- Unusually long tool arguments; base64 blobs; obfuscated code.
- Sudden role shifts (“you are now…”).
Response plan
- Fail safe: refuse or require human review on suspicious inputs.
- Strip instructions from user content before reuse.
- Regression tests for jailbreak prompts before each deploy.
TL;DR
Assume users (or their docs) will try to hijack your agent. Constrain tools, validate input, and log everything.