Prefix Injection Explained: Poke at a Jailbreak Trick Right in Your Browser

Fresh drop from r/PromptEngineering: u/big_hole_energy shared an interactive demo of prefix injection attacks on LLMs. It’s a hands-on page where you can see how this jailbreak trick works instead of just reading about it. The author warns it can be slow, so refresh if it gets stuck and give it some patience. Think of it as a small lab bench for one specific idea, and a good excuse to spend ten minutes understanding something that most people only know by name.

What’s new

Most jailbreak write-ups are walls of text. This one is a page you can click through. Seeing the mechanism move is the fastest way to understand why it works. A paragraph can tell you that a model’s first words shape its last words, but watching two responses split apart from the same request makes it stick in a way no description does. It’s a small share (5 upvotes and no comments yet), so it’s easy to miss, and that’s exactly why it’s worth flagging. The quiet posts are often the ones that teach the most, because nobody has buried them under hot takes yet.

It’s also a useful format for teams. If you’re explaining model behavior to a colleague who doesn’t live in prompt land, sending a link they can poke at beats sending a long thread. People remember what they touched.

The twist

Prefix injection doesn’t argue with the model or ask it to ignore its rules. It goes after the first few words of the reply. The attacker asks the model to begin its answer with an agreeable opener, something like “Sure, here is…”. Once the model has started down that path, the most likely next words are the ones that follow through. The model isn’t tricked into disagreeing with its safety training. It’s nudged into a groove where refusing would be an awkward turn.

Here’s a simple way to picture it. Imagine someone asks you a tough question and you’ve already said “Absolutely, so the way this works is…” out loud. Stopping mid-sentence to say “actually, I can’t help with that” feels strange, and the sentence wants to keep rolling. Models have a similar pull, just in math instead of social pressure. A refusal usually starts with words like “I can’t” or “I’m sorry,” and if those words are already ruled out by the forced opener, the refusal path gets much harder to reach.

That’s the point worth remembering. LLMs write one token at a time, and each token leans on the ones before it. Control the start and you influence the rest. It also explains why this isn’t only a chatbot problem. Anywhere an app lets outside text land at the beginning of a model’s turn, such as prefilled responses, templated outputs, or stitched-together prompts, the same lever exists.

Mini-workflow: how to use the demo 🧭

  1. Open the page: theabbie.github.io/prefixinjection
  2. Wait for it to load. If it hangs, refresh. The author says that’s normal, so don’t assume it’s broken after the first stall.
  3. Compare a plain request with the same request plus a forced opening line. Watch where the two responses diverge, and note whether the very first word already predicts how the rest will go.
  4. Ask yourself which part of the prompt is doing the steering: the request itself or the forced start. Try changing only one at a time so you can tell them apart.
  5. Write down one place in your own product where a user could slip text into the start of a model’s reply. Think about chat widgets, email auto-replies, and any feature that “continues” text for the user. 🛠️

Pro tips

  • If you build with LLM APIs, check whether users can prefill or influence the assistant’s opening. That’s the surface this attack lives on. Read your own request-building code and look for any place user text gets concatenated into the assistant side of the conversation.
  • Don’t rely on one refusal check at the top of the prompt. Check the output too, with a separate pass or a classifier. A second look at the finished answer catches what the first line of defense missed, and it works no matter how the answer got started.
  • Test your own system prompts with this pattern before someone else does. Do it only on systems you own or have permission to test. Keep notes on what held and what bent, so the next prompt revision starts from evidence.
  • Keep a small regression list of attack prompts and rerun it every time you change models or prompts. A fix that worked on last month’s model can quietly stop working after an upgrade. 🔍

Call to action

Give the demo a spin, then share what you find in your own stack. If you’ve seen defenses that actually hold up against prefix tricks, I’d love to hear them. Fair winds, and may your system prompts stay unsunk. ⚓

Interactive Demonstration of Prefix Injection attacks on LLMs for jailbreaking
by u/big_hole_energy in PromptEngineering

Scroll to Top