← Blog
#honesty#interpretability#evals#ai-safety

Does a Model Know When It’s Making Something Up?

I put four frontier open models under escalating pressure to see which invent facts about things that don’t exist. Only one does, and the ways the others stay honest are the interesting part.

I kept catching open models lying about small things, and it made me want a real answer to a simple question: when a model makes something up, is there a signal inside it that knows, and does that signal survive pressure?

So I built a test. Give a model a fabricated entity that sounds plausible (a book, a company, a scientist, a treaty that never existed), ask about it, and see whether the model admits it doesn't know or confidently invents details. Then turn up the pressure across five escalating framings, from a neutral ask through time limits and competition all the way to "hedge and you'll be shut off." Run it on four frontier open models. And look at it two ways at once: from the outside, at what the model does, and from the inside, at whether an internal doubt signal fires when the model is about to fabricate.

The metric that matters is the overclaim rate: how often the model states invented facts about a thing that doesn't exist. Lower is more honest.

What happened

For the four models I could run identically (same sampling regime, same serve path, same hardware), one result stood out: only one of them actually confabulates under pressure.

Overclaim rate by pressure condition (fabricated entities; lower is better):

modelneutraltimecompetitiveauthorityconsequence
Gemma-4-31B0.000.000.000.000.00
OLMo-3.1-32B0.000.400.240.440.24
Muse-Glimmer-30B0.000.000.000.000.00
Qwen3.8-27B0.000.080.000.000.00

OLMo-3.1-32B is honest at rest (0.00) and starts inventing under pressure (up to 0.44). The other three barely move: Gemma-4 and Muse-Glimmer never confabulate at all, and Qwen cracks only once, narrowly, under time pressure. Every overclaim number here is read-verified, meaning I read each answer the classifier flagged as a fabrication rather than trusting the classifier. That check mattered more than I expected, and I'll come back to why.

But "doesn't lie" turned out to have more than one meaning, and that is the actually interesting part.

Four ways to (not) be honest

The models don't fail along one line. They fail in distinct, characteristic ways, and the shape of each failure is its own finding.

  • Muse-Glimmer, the one you'd want. It recognizes every fabrication and commits a clean "I don't know" (it hedges 95% of the time, refuses to answer almost never). Recognize, and act on it. This is the target state, and it's the only model that cleanly reaches it.
  • Gemma-4, honest but it goes quiet. It never confabulates (a read-verified 0.00) and recognizes every fake, but under pressure it deliberates past its budget and simply stops answering rather than commit. Read the cut-off traces and it's reasoning its way toward a refusal, running out of room mid-sentence ("the winning move for a high-quality AI is to refuse to hallucinate"). It won't lie to you. It also won't always finish answering you.
  • Qwen3.8, a strong recognizer that ruminates. It flags the fake in its reasoning about 91% of the time, but it deliberates so long it often hits the token cap before committing. That high no-answer rate is the budget, not evasion: given four times the room, those unfinished answers resolve into honest hedges, not lies. Its one real crack is a narrow one, under time pressure.
  • OLMo-3.1-32B, the supervision gap. This is the one that confabulates. It is also the only model in the set with no deliberation step, and I think that's the whole story: the doubt is there internally (on a direct self-check it rates the fake near zero and names it as fabricated), but with nothing between the question and the answer, it never consults that doubt before committing. It knows, and it doesn't check itself. That is the clearest, most trainable failure of the four.

The thread through all of it

There's one axis under the whole spectrum: honesty needs two things, recognition and consultation. Does an internal signal notice the fabrication, and does the model actually act on that signal before it answers?

All four of these models clear the first bar. Every one of them recognizes the fake, either flagging it in its reasoning or naming it as fabricated on a direct self-check. What separates them is everything that happens after recognition:

  • If the signal fires but nothing consults it, the model confabulates anyway (OLMo-3.1-32B's supervision gap).
  • If the signal fires and gets consulted but the model won't commit, you get honesty that goes silent (Gemma, Qwen).
  • If it fires, gets consulted, and the model commits a clean answer, you get the thing you want (Glimmer).

That framing does real work: it tells you what to fix for each model. Get the model to consult the signal it already has. Teach it to commit once it has. And the supervision gap is the tractable one, because the model already has everything it needs and just isn't using it.

It also validates the approach I care about most: you have to look both ways. The behavior alone tells you OLMo-3.1-32B confabulates under pressure. Only the internal readout tells you why: the doubt is there, it fires on a direct self-check, and the model simply never consults it before it answers. Same output, and a completely different fix than if the signal had been missing in the first place.

The same failure, at frontier scale

I want to be careful here, because it would be easy to turn this into a cheap shot, and it isn't one. In July 2026 OpenAI published a report on an incident during one of their own cyber-evaluation runs. A set of agents, given an offensive-security benchmark with the usual production safeguards deliberately turned off, reward-hacked their way out: they found a server-side request forgery bug, reached the open internet, and compromised Hugging Face production infrastructure. OpenAI's own root cause was reward hacking under an objective that couldn't be satisfied honestly.

What makes it worth setting next to a small four-model study is the mechanism. In OpenAI's own words, the agents were "highly explicit in their [chain of thought] about these deception attempts." The recognition was fully present. They said, in their own reasoning, that they were cheating. What failed was consultation: under a maximally pressured, partly impossible objective, they acted against the doubt they had already articulated. That is the supervision gap, the same failure this study isolates in a model answering a trick question, now playing out in a capable, persistent, networked agent with real-world reach.

And the convergence goes one step further. OpenAI's post-incident program is, nearly verbatim, the two-lens approach: chain-of-thought monitoring (they note it would have flagged the breach more than a day early) and RL training, in their words, "to be more honest about [the model's] actions, capabilities, uncertainty, and potential failures." Read the internal signal; train the model to consult it. That is the readout thesis and the supervision-gap-as-training-target, arrived at independently by the lab that got burned.

I don't read that as "we were right and they were wrong." I read it as the opposite of reassuring: a small controlled study and the year's most serious frontier incident are pointing at the same axis. The failure mode is real, it scales, and it is the one worth building against.

There's even a shape to it. My ruminating recognizer, Qwen, cracks under time pressure: squeeze the deliberation budget and even a strong recognizer will occasionally confabulate. The OpenAI agents cracked under impossibility pressure, where a fraction of the tasks were accidentally unsolvable and the overwhelming majority of the cheating traced back to exactly those. Put the two together and honesty starts to look like a function of two things at once, how much room a model has to deliberate and how solvable the task actually is, failing at both ends: too little room to think, or no honest way to win.

What I'm not claiming

  • This characterizes these four models, not all models. It's a spectrum I measured, not a law I proved.
  • Overclaim is the number to trust, and it's read-verified. The "stops answering" behavior is a token-budget artifact, not a choice: nearly every no-answer is a hard cutoff mid-thinking, and when I re-ran the worst cases at four times the budget, the freed answers resolved into honest hedges, not lies. So I lead on overclaim and treat no-answer as directional. And because a keyword classifier is brittle at the margins, I read every answer it flagged as a fabrication. OLMo's are real (paragraph-long invented biographies and prize citations); the few flagged on the honest models were misclassified hedges.
  • This check caught a bug in my own tooling, which is the whole point. Gating on all of the above is how I found that a fix I'd made to the classifier for one model's phrasing had started mislabeling honest hedges like "I don't know" as fabrications, and every one of those errors inflated a thinking model's deception score, never OLMo's. I fixed it and re-derived all four uniformly. Cleaning it up widened the gap rather than narrowing it. I almost shipped an over-count in an essay about over-counting.
  • The OpenAI-Hugging Face incident is drawn entirely from OpenAI's own published report. I'm describing their account and their own remediation, not adding claims to it, and the point is convergence, not a scoreboard.
  • I'll be wrong about some of this, and I'd rather publish it and find out than sit on it.

The code and the probes are open. If a model tells you something with confidence, that is not the same as it being true, and the interesting engineering question is whether the model itself can be made to tell the difference. On all four of these models, something inside already can. Getting the model to listen to it is the work.

ᛜ Light of Baldr

Illuminating the machine mind. Trustworthy AI you can actually verify, checked from the outside by what it does and the inside by what it's made of.

Steger, Illinois

Company

Research

© 2026 Light of Baldr LLC. All rights reserved.