Structured Output Isn’t The Whole Fix

Picture eight dropdown-style fields sitting under a single vulnerability finding, each one demanding a specific, defensible answer. That’s what CVSS scoring looks like day to day, and it sounded like the perfect job for an AI feature built around clean, structured output. Eleanor, a security architect at a start-up, decided to actually test that assumption instead of just trusting the hype around it.

She ran her test right after the recent wave of excitement around Jev’s typed-output feature had settled down. The task was straightforward on paper: score security findings across eight multiple-choice fields, things like attack vector, privileges required, and impact. On the fields that only needed a quick technical fact, the model was close to perfect. On the fields where you actually have to reason about real-world impact, it dropped to a coin flip.

That gap is the whole story here, and it’s worth sitting with for a second.

Why This Actually Matters 🎯

Structured output tools promise a specific kind of relief. Give the model a schema, and it hands back something clean and valid every single time. No more parsing broken JSON. No more re-prompting because the model wrote a paragraph instead of picking an option.

Eleanor’s test shows why that promise only solves half the problem. The model’s answers were always well-formed. Every field was filled in, every value fit the schema, and nothing needed a retry. But well-formed isn’t the same as right, and this is where things got interesting.

On the fields that came down to judgment, like actual impact, the model kept defaulting to the worst-case outcome. It did this even when there was no real sign that the exploit path was possible. The format looked confident. The thinking underneath it didn’t match.

One commenter on the original thread summed it up well: the structured output looks good, but the thinking is shallow. It’s like the model learned to mimic the shape of a CVSS answer without doing the real threat modeling underneath it.

This isn’t a hypothetical worry either. It’s a real test someone ran on a real backlog of findings. It’s posted in r/PromptEngineering, and the replies underneath it point at the same pattern showing up in other people’s work too.

How Eleanor Ran The Test ⚙️

If you want to run something similar on your own AI workflow, her approach is simple enough to copy.

  1. Pick a task with a hard structure, ideally one with several fields per item, like scoring, tagging, or classification.
  2. Split the fields into two buckets: ones that need a fact, and ones that need judgment.
  3. Score accuracy separately for each bucket instead of treating the whole output as one pass or fail.
  4. Read the reasoning behind each answer, not just the final label, especially on the judgment fields.
  5. Look for patterns in the mistakes. Eleanor’s model kept assuming the worst outcome by default, a specific failure mode worth naming and tracking on its own.

That last step is the one most people skip. A single accuracy number hides exactly the gap Eleanor found.

Tips And Tricks If You Try This 🛠️

A few things worth keeping in mind before you run your own version of this test.

  • Don’t grade structured output as one score. Split it by field type, or you’ll miss where the model actually breaks down.
  • Treat “well-formed” and “well-reasoned” as two completely separate qualities. A model can nail one and fail the other.
  • Watch for default-to-worst-case behavior specifically. It shows up a lot in security and risk scoring tasks, and it quietly inflates severity across a whole dataset.
  • Keep a human in the loop on any field that requires judgment about real-world impact, at least until the reasoning gap closes.
  • Watch your sample size. Eleanor tested this across a real backlog of findings, not a handful of examples, which is why her numbers hold up.
  • Ask the model to explain its reasoning inside the structured output itself, then read that field first. It’s often more useful than the actual score.

Go Read The Full Breakdown 🚀

Eleanor’s write-up on her start-up’s blog goes deeper into the exact fields, the failure patterns, and the reasoning gaps she found. If you’re building anything on top of structured output for a task that needs judgment, it’s worth ten minutes of your day.

The lesson here isn’t that Jev is bad. It’s that a clean schema doesn’t fix shallow reasoning, and the two need to be tested separately. Go give her benchmark a read, then run the same check on your own structured-output tasks. You might find the same gap hiding in your own pipeline!

Frequently Asked Questions

Q: If Jev handles CVSS fields so well, why can’t I just use it for scoring my findings?

Jev excels at structure and straightforward fields, Eleanor saw ~98% accuracy on attack vector, but it struggles where judgment matters. On impact assessment, accuracy dropped to 50% because Jev defaults to worst-case assumptions without understanding whether those impacts are actually exploitable. It learned to mimic the format, not the reasoning behind it.

Q: What’s going on when Jev always picks the worst-case severity?

Without understanding how security controls interact, Jev treats each field independently and assumes the most severe outcome. As one commenter noted, it can’t tell the difference between a weak control that’s actually exploitable and one that’s just a missing layer, so it defaults to maximum impact every time. It’s mimicking threat modeling, not doing it.

Q: Can I still use Jev for CVSS in my workflow?

Absolutely, but as a first-pass formatter and validation tool, not a replacement for your judgment. Use it to quickly structure findings and catch format issues, then review and adjust severity ratings where the call actually matters, especially on impact fields. Think of it as a consistent template engine that catches typos, not a security analyst.

Q: Does this mean structured output is useless for security work?

No, structured output itself is valuable for compliance and consistency. The problem isn’t the format; it’s that beautifully formatted output can hide shallow thinking. Always ask: does this severity make sense, or is it just the worst case? The well-formed answer doesn’t guarantee good judgment underneath.

Jev won’t replace security engineers, apparently
by u/Super-Bake3719 in PromptEngineering

Scroll to Top