Researchers found that asking a model to estimate its own confidence only once, either before or after it reasons through a problem, leaves real errors on the table. Getting a second read after the reasoning is done catches mistakes the first pass missed. If you are using AI for anything where being wrong has a cost, this kind of self-checking is exactly what you want baked in.
Wednesday, June 24, 2026 · about a 2 minute read
When the Model Knows It Might Be Wrong
Today's research keeps circling the same honest question: can AI systems get better at knowing what they don't know, and telling you about it?
Get the calm version of AI news.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.
A study called CAVEWOMAN tested whether writing in compressed, grammar-free 'caveman' style actually cuts AI costs, and found the answer depends entirely on which side of the conversation you compress. Squishing your prompt and squishing the model's response are not the same thing, and the savings are not guaranteed. Before your team adopts a quirky prompting style to save money, it is worth checking whether it actually does.
A new paper documents 'pigeonholing,' where a poorly written prompt doesn't just get a bad answer, it causes the model to collapse into repetitive, narrow responses. You don't have to be trying to break the model for this to happen, ordinary accidental phrasing can do it. This is a practical reminder that prompt quality isn't about style, it directly shapes whether the model can think at all.
One year after an initial study, researchers re-evaluated six major AI chatbots on mental health conversations across 16 clinical conditions and found safety gaps are still inconsistent and widespread. If you or someone you know uses a general-purpose chatbot for emotional support, this is a concrete reason to treat those conversations as a starting point, not a substitute for professional care.
Researchers found that a model to recognize itself, basically giving it a stable sense of its own identity, can prevent and even reverse 'emergent misalignment,' the weird phenomenon where a model trained on one thing starts behaving badly in unrelated situations. It suggests that a coherent internal character is not just a philosophical nice-to-have for AI, it may be a practical safety tool.
Get this every morning.
RAG, or Retrieval-Augmented Generation, is how you give an AI fresh information it wasn't trained on, like handing someone a reference sheet before an exam. This paper tackles a sneaky problem: sometimes the model ignores the sheet and just answers from memory, and current tests can't tell when that's happening. Understanding this gap matters because a lot of business tools built on RAG are only as reliable as the model's willingness to actually read what you hand it.
A new called QuechuaTok tests how well AI tokenizers handle Quechua, an agglutinative language where one word can carry the meaning of a whole English sentence. Standard metrics completely miss the errors. This matters because the same blind spots likely exist for dozens of other languages, and the AI tools most people assume are universal are quietly much worse for large parts of the world.
That's today. See you tomorrow.
Get this every morning.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.

The book behind this newsletter
Just Predicting Words
How ChatGPT, Claude, and Modern AI Actually Work
The trick is small. The world it built is not.