When you give a chatbot important details piece by piece across a long conversation, its accuracy can drop by 65%, even though all the information is technically still there in the window. If you use an AI assistant for anything that builds up over multiple messages, like a project plan or a medical question, this is a real and current limitation you should know about.
Friday, June 12, 2026 · about a 2 minute read
AI That Listens, Learns, and Sometimes Gets Distracted
Today's research keeps bumping into the same honest problem: these models are impressive until you hand them something slightly messy, and then things get interesting.
Get the calm version of AI news.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.
Simon Willison, a developer whose day job involves building with these models, describes Claude's newest version as relentlessly proactive, meaning it reaches for tools and strategies on its own without being asked. That is useful when it works and a little hard to predict when it does not, which is a useful framing for anyone thinking about putting AI into a real workflow.
Researchers found that if you ask an AI to do something harmful in a language other than English, the safety guardrails are noticeably weaker, because safety training has been concentrated almost entirely in English. If you or your organization use AI tools with a global team, the safety behavior your English-speaking colleagues see may not be what everyone else gets.
A new was built specifically to test AI shopping assistants on multi-turn conversations, the kind where you say 'I need something for a wedding' and then 'actually it is outdoor' and then 'my budget changed.' Current models handle these real shopping conversations worse than the clean single-question tests suggest, so the gap between demo and reality is still wide.
Get this every morning.
Researchers found that the format you use to feed information to a model, how you structure and present the text, changes the model's output independently of what the information actually says. Think of it like reading a memo versus reading a legal contract. Same words, different shape on the page, and your brain pays differently. Models do the same thing, and that is worth knowing if you ever craft prompts or build anything on top of a retrieval system.
A new called Polar tests AI models for political bias across multiple countries and languages, not just American English political categories. As these models get used in news tools and civic applications around the world, knowing whether they lean one way in one country and another way elsewhere is a genuinely important question that nobody has had a clean way to measure until now.
A paper argues that when researchers see interesting patterns inside a model's internal states, like evidence it is doing something that looks like reasoning, those patterns are not the same thing as proof that reasoning is actually happening. It is a small distinction that sounds philosophical until you realize it affects how much you should trust a model when it confidently shows its work.
That's today. See you tomorrow.
Get this every morning.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.

The book behind this newsletter
Just Predicting Words
How ChatGPT, Claude, and Modern AI Actually Work
The trick is small. The world it built is not.