48-Hour AI
All Episodes
Ternary AI, Local LLMs, and Robots in Real Homes

Ternary AI, Local LLMs, and Robots in Real Homes

0:00|0:00

We explore how a 27B AI model can now run locally on a consumer laptop with under 6GB of RAM, thanks to ternary weights and group-wise scaling. The episode also connects this leap in efficiency to humanoid robots generalizing in real homes and infra agents autonomously managing massive datacenter deployments.


Chapter 1

Squeezing 27 Billion Parameters into 5 Point 9 Gigabytes

James Turner

So, um, if I told you two years ago that you could run a full 27 billion parameter AI model, with a massive 262 thousand token context window, multimodal vision capabilities, and high level reasoning, directly on a consumer laptop with less than 6 gigabytes of RAM, you, you would have told me I was completely dreaming.

James Turner

Because back then, standard wisdom in machine learning said that 4 bit quantization was basically the hard floor. If you tried to push model weights down any lower, the intelligence just collapsed into gibberish, right? But er, this week, Ternary Bonsai 2 27B completely shattered that floor.

James Turner

They managed to shrink a 27 billion parameter model into a tiny 5 point 9 gigabyte footprint! How? By dropping weight precision down to just 1 point 76 effective bits per weight. Yeah, 1 point 76 bits!

James Turner

Let's, let's break down how that math actually works under the hood, because it's fascinating. Instead of standard 16 bit floating point numbers or even 4 bit integers, Bonsai uses ternary weights. That means every single weight in the network is restricted to just three possible values: minus 1, 0, or plus 1.

James Turner

Now, you might wonder, how on earth do you get nuanced reasoning out of a system where every parameter is just a simple off, zero, or on switch? The secretsauce here is pairing those ternary weights with FP16 group wise scaling.

James Turner

So, instead of storing full precision numbers for all 27 billion weights, you store simple discrete values in memory and apply small scaling factors across groups of weights. It gives you nearly lossless compression. The model retains its reasoning, its coding abilities, and its visual processing, but the physical size in memory drops by almost 9x.

James Turner

I actually tested low bit local execution on my own development rig yesterday using custom kernels for Apple MLX and NVIDIA CUDA. And, um, I have to admit, there is something deeply satisfying about running a heavy agentic workload completely offline.

James Turner

No monthly cloud API bills ticking away in the background, no network latency, no privacy concerns about where your data is traveling. You're just pulling tokens straight off your local memory bus at lightspeed. It fundamentally changes how we think about edge computing.

Chapter 2

Zero Shot Household Robotics and Self Deploying Infra

James Turner

And that shift toward extreme efficiency isn't just happening inside software models on our laptops. It's leaking directly out into the physical world and infrastructure layer in ways that feel almost surreal.

James Turner

Take robotics, for example. Figure just introduced Helix 2 point 5, which is a humanoid control model pretrained on their Index dataset. And here is the kicker: they tested it zero shot across 30 different homes in the Bay Area.

James Turner

Zero shot means the developers collected zero training data inside those specific homes. The robot walked into 30 completely unfamiliar living rooms and kitchens, and it just started tidying rooms, folding towels, and making beds.

James Turner

Think about how hard that is! Every house has different lighting, messy clutter, weird furniture layouts, and different fabric textures. Normally, a physical robot stumbles the second a single chair is shifted three inches to the left. But Helix 2 point 5 generalized straight out of the box.

James Turner

And while humanoids are learning to navigate our messy homes, AI agents are taking over massive datacenters at the exact same time. Look at what Z dot ai did with GLM 5 point 3.

James Turner

They deployed a GLM powered Infra Agent to configure, optimize, and launch a production serving stack across more than 100 thousand specialized accelerators. And they did the entire deployment in under two weeks.

James Turner

The agent was writing custom kernel fixes, handling dense feedback loops, and tweaking system level parameters on the fly. In the end, it tripled the system's operational throughput, while human engineers were kept strictly in the loop just to oversee high level risk and objectives.

James Turner

Do you see the pattern connecting all of this? We are moving past the era where AI just writes text or generates code snippets in an editor. We are entering an era of physical and infrastructural execution.

James Turner

Whether it's an agent configuring 100 thousand chips in a server farm or a robot folding laundry in an unknown apartment, the hardware layer is becoming automated. Human engineers aren't writing every line of config or training every motor policy anymore. We are turning into system architects and risk managers.

James Turner

It's a wildly fast transition. And honestly, running these low bit models locally gives us a front row seat to where it's all heading. Alright, that's it for today's quick take, talk soon!