The Next TokenLearnBookArchive
Saved

Tuesday, July 7, 2026 · about a 2 minute read

Smaller, Cheaper, and Quietly Smarter

Today the story is not about one giant leap. It is about a dozen small bets that AI gets more useful when it gets leaner, more honest about what it does not know, and closer to where people actually live.

Get the calm version of AI news.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Hacker NewsBusiness
GLM 5.2 and the coming AI margin collapse

A well-argued post is making the rounds about how cheap, capable models from China are squeezing the profit margins of every AI company. If the cost of intelligence keeps falling this fast, the business models that everyone built in 2023 are going to need a serious rethink, and that includes the price you pay for the tools you use at work.

Read
Simon WillisonModels
tencent/Hy3

Tencent released Hy3, a 295-billion- model under an open Apache 2.0 license, meaning anyone can download and use it freely. More open, capable models mean more options for companies that do not want to be locked into one vendor, which is good news if you have ever felt stuck paying for something you could not inspect or control.

Read
Hacker NewsModels
Small AI Models Gain Traction In places with unreliable networks

Small AI models are gaining real traction in places with unreliable internet, including hospitals and pharmacies in low-connectivity regions, because they can run on a single device without a cloud connection. This matters because it is a reminder that the most important AI deployments of the next decade may happen somewhere with no reliable signal, not in a San Francisco office.

Read
arXiv cs.CLModels
Gemma 4 Technical Report

Google released its Gemma 4 technical report, detailing a new family of models built for efficiency and reasoning, meaning they can handle text and images together. Open-weight models you can run yourself are getting seriously good, which changes the calculus for any team deciding whether to build on a closed API or host something locally.

Read

Get this every morning.

arXiv cs.CLResearch
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

Researchers found that LLMs are consistent, meaning they give the same verbal answer repeatedly, but miscalibrated, meaning the confidence in that answer does not match reality. Think of it like a friend who always sounds certain when giving directions but is wrong about one in four turns. Knowing this should change how much you trust an AI explanation of a risk or a probability, even when it sounds perfectly confident.

Read
Want the slow, plain-English version of why this matters? This is exactly the kind of idea the book was written to unpack, one light-switch analogy at a time.JPWExplained properly in the book
arXiv cs.CLSafety
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

A new paper shows that frontier models can do multi-step reasoning using meaningless filler like dots, with no visible chain of thought that humans can check. It is a small, strange finding, but it raises a real question: if the thinking is hidden inside tokens that carry no words, how do we know what the model is actually doing?

Read

That's today. See you tomorrow.

Get this every morning.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Just Predicting Words book cover

The book behind this newsletter

Just Predicting Words

How ChatGPT, Claude, and Modern AI Actually Work

The trick is small. The world it built is not.

PaperbackKindleAudiobook · SpotifyAudiobook · Google Play