Mistral released Shieldstral, a small open-weights model built specifically to flag harmful content in text and images. If you build anything on top of AI, this is the kind of cheap, inspectable filter that can sit between your users and the model without you having to trust a black box.
Wednesday, August 5, 2026 · about a 2 minute read
AI Gets a Bouncer, a Memory Leak, and a Four-Thousand-Year Side Quest
Today's news is really about one question: what happens when the tools get more capable but the guardrails lag behind.
Get the calm version of AI news.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.
Interpol says AI is now driving more than half of cybercrime in Africa, from phishing to scam voice calls. The practical upshot: the scam email that used to announce itself with bad grammar no longer does, so the old advice of 'just look for typos' is quietly obsolete.
Simon Willison shipped 0.32, a command-line tool that now surfaces reasoning traces and connects to server-side tools, meaning you can watch the model think out loud before it answers. For anyone who uses AI in their daily work, seeing that trace is the difference between trusting an answer and just hoping it is right.
Researchers built TabletCraft, a system that translates ancient Akkadian cuneiform in both directions and even renders the actual wedge-shaped characters. Half a million clay tablets in museums have been basically unreadable to anyone without a very specific PhD, and that wall just got a lot shorter.
Get this every morning.
A new study found that language models encode whether a statement is true or false as a kind of direction inside their own internal math, like a hidden compass needle pointing toward truth. This matters because it means the model is not always lying when it hallucinates: sometimes it knows the right answer internally but still says the wrong thing out loud, which tells us the problem is in how it converts internal knowledge to words, not just in what it learned.
A systematic study found that the benchmarks used to rank AI models are hitting a ceiling, with top models scoring so high that differences between them become meaningless noise. When the scoreboard stops working, it gets harder to know whether a new model is genuinely better or just better at the test.
Researchers tested whether doctors preferring one AI medical response over another actually means the preferred response is safer, and found the answer is often no. If you ever see 'clinician-preferred AI' as a selling point for a health product, this paper is a useful reason to ask a follow-up question.
That's today. See you tomorrow.
Get this every morning.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.

The book behind this newsletter
Just Predicting Words
How ChatGPT, Claude, and Modern AI Actually Work
The trick is small. The world it built is not.