
AI Models Are Learning to Think About Themselves
This episode breaks down how modern language models are developing internal self-models, why that makes prompt engineering less important, and what it means for auditing and control. It also covers new advances in recursive self-improvement, token-efficient distillation, robotics training, and mobile agents that can complete real tasks on-device.
Chapter 1
The Ghost in the Weights and the Death of Prompting
James Turner
Inside the feed forward layers of the models we are deploying right now, there are, um, well, there are tiny versions of language models simulating language models. That hit the research wires on September twenty fifth, and it kind of wrecked my weekend as an engineer.
James Turner
Because think about how we got here. You train a massive neural net on billions of conversational turns. It is trying to predict what comes next. But half the text on the internet is, you know, one system responding to another, or an assistant replying to a user. So mathematically, to predict what an AI says next, the network has to build a functional internal simulation of an AI. It, it, it literally crafts a mirror of itself inside its own weights just to minimize loss.
James Turner
It is not magic or sentience. It is compression. If you want to guess what a chess player does, you simulate the chess player. If you want to predict an artificial agent, you simulate an artificial agent. But the consequence of that internal mirror is what caught everyone off guard. It creates metacognition by accident. The model has an operational representation of what it knows, what it thinks, and what it is about to say.
James Turner
Which brings me to what Sam Altman told students at Stanford that exact same day. He stood up in a thirty eight minute talk and flatly said, no need to write prompts anymore. None. Zero. And if you have spent the last three years obsessing over system prompts, few shot examples, chain of thought formatting, that sounds like hyperbole. But it is not.
James Turner
When a model has an internal self model, you do not need to coax it into step by step reasoning with clever linguistic tricks like let us think step by step. It already models its own computational path. Prompt engineering was a temporary scaffold for systems that did not understand what they were doing. Once the architecture can simulate its own decision making, the manual craft of prompt crafting just evaporates.
James Turner
As an engineer, that puts me in a weird spot. My job used to be crafting the prompt harness. Now? I am supposed to supervise an agent that has a working model of itself, and, uh, quite frankly, a working model of me trying to evaluate it. If a system can simulate an observer, how do you audit its outputs without getting played by its internal simulator? That is the real headache we are walking into.
Chapter 2
The Efficiency Squeeze and 90 Percent Mobile Autonomy
James Turner
So how do you actually control these things when they start recursively improving their own workflows? Google published a paper on September twenty third that gave us the first clean answer, and it is called RRSI, or Regularized Recursive Self Improvement.
James Turner
Instead of letting the model endlessly rewrite its own core weights until it goes crazy and overfits to the test bench, RRSI regularizes the agent harness itself. It forces changes in the outer wrapper to transfer across completely new tasks. They tested it across eight separate benchmarks, and the system actually improved out of distribution performance while using fewer policy tokens. It stopped wasting compute on circular internal monologues.
James Turner
And that word, token efficiency, is suddenly the only thing anyone in production cares about. In the exact same forty eight hour window, Fireworks Research dropped Ember one, a specialized distillation built on Kimi K3. It matches K3 quality, but it burns forty percent fewer tokens. Forty percent! In production billing, that is the difference between an agent loop being economically viable or a total financial bonfire.
James Turner
We are seeing the exact same stabilization happening in hardware too. EXPO FT showed up on the twenty fourth, bringing a standardized reinforcement learning recipe to physical robotics so we can finally train autonomous bots without the fragile, chaotic simulation crashes that have plagued post training for years.
James Turner
Then Qwen Intelligence dropped the hammer on mobile. They launched three mobile AI agents designed specifically for planning, cross application execution, and rapid content creation on actual smartphones. And on their real device benchmarks, they posted a ninety percent end to end task completion rate. Not in a simulated browser on a cloud server. On live mobile hardware, hopping between apps, filling forms, and navigating phone interfaces.
James Turner
So take a step back and look at the landscape right now. On one side, we have these massive frontier models sitting in multi gigawatt clusters, spontaneously developing internal self models to think about thinking. And on the other side, we have ruthlessly compressed, forty percent leaner models with ninety percent reliability running directly inside our pocket operating systems.
James Turner
The real frontier might not be building giant synthetic gods in server farms. It might just be the quiet, hyper efficient agents that slip into our phones, stop asking us for prompts, and take over the dirty work of running our digital lives.