𝗟𝗟𝗠 𝗣𝗿𝗼𝗺𝗽𝘁 𝗜𝗻𝗷𝗲𝗰𝘁𝗶𝗼𝗻 𝗮𝗻𝗱 𝗚𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆

📅2 hours ago⏱1 min read

LLMs have no hard boundary between instructions and data. Everything in the context window is one stream of tokens. Prompt injection happens when attacker data acts as instructions. You cannot filter your way to safety. You must manage it with defense-in-depth.

The failure of common defenses:

Keyword Blocklists: Attackers use synonyms, misspellings, or different languages to bypass them. Filtering strings does not filter intent.
Output Redaction: Attackers can fragment or encode secrets so a literal string match fails.
LLM Judges: A separate model can be socially engineered to believe a secret is harmless.
Human Review: Humans see rendered text, not raw bytes. They cannot see hidden characters used in ASCII smuggling.

ASCII Smuggling is a major threat. It uses invisible characters like Unicode Tags or zero-width spaces to hide instructions. The model reads them, but the human sees nothing. This allows identity spoofing and data exfiltration via email or calendars.

How to defend your application:

Sanitize raw payloads: Strip control characters and zero-width characters before they reach the model.
Use allowlists: Define the specific Unicode categories you need instead of chasing bad ones.
Normalize data: Use NFKC-normalization on all inputs.
Minimize secrets: Do not put sensitive data in the context window if the model does not need it.
Treat RAG as untrusted: Assume any document you retrieve for a model is a potential injection vector.
Watch for anomalies: Flag inputs where the visible length differs from the raw code-point count.

Security is a pipeline flaw, not just a model flaw. The fix lives in your application code.

Source: https://dev.to/geekaara/llm-prompt-injection-guardrail-security-glm

Optional learning community: https://t.me/GyaanSetuAi

𝗟𝗟𝗠 𝗣𝗿𝗼𝗺𝗽𝘁 𝗜𝗻𝗷𝗲𝗰𝘁𝗶𝗼𝗻 𝗮𝗻𝗱 𝗚𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆

Continue reading

𝗧𝗵𝗲 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗔𝗜 𝗚𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲 𝗙𝗿𝗮𝗺𝗲𝘄𝗼𝗿𝗸

𝗚𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀 𝗳𝗼𝗿 𝗘𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗔𝗜 𝗔𝗴𝗲𝗻𝘁𝘀

𝗖𝗹𝗮𝘂𝗱𝗲 𝗖𝗼𝗱𝗲 𝗜𝗻 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻: 𝗧𝗵𝗲 𝗚𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀 𝗬𝗼𝘂 𝗡𝗲𝗲𝗱

𝗔𝗜 𝗚𝗮𝘁𝗲𝘄𝗮𝘆 𝗚𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀 𝘄𝗶𝘁𝗵 𝗔𝗪𝗦 𝗕𝗲𝗱𝗿𝗼𝗰𝗸 𝗮𝗻𝗱 𝗞𝗼𝗻𝗴

𝗬𝗼𝘂𝗿 𝗥𝗲𝗽𝗼 𝗖𝗼𝗻𝘁𝗲𝘅𝘁 𝗜𝘀 𝗔𝗻 𝗔𝘁𝘁𝗮𝗰𝗸 𝗦𝘂𝗿𝗳𝗮𝗰𝗲 𝗡𝗼𝘄