Skip to content
Writing

Essay

We built cybernetics. We called it AI.

A four-rung ladder of intelligence, a quiet retreat from the word AGI, and $720B of capex pointed at a story the people telling it have stopped believing.

I picked up Learning Deep Representations of Data Distributions sometime in June, meaning to skim it. Buchanan, Pai, Wang and Ma, out of Berkeley, and prepared for a course on the principles of deep learning and intelligence in general. Something had trended, I wanted enough context to nod at it, the usual arrangement. I got through one chapter and then it sat at the back of my head for two months like a stone in a shoe.

This post is what eventually fell out. It didn't arrive all at once, so I'm going to lay it out roughly the way it actually came.

The book sets up a ladder. Four rungs.

At the bottom, phylogenetic intelligence. This is the species-level reinforcement learning of evolution. The only way to update the weights is for a generation to die. Slow, costly, and it runs on death. Nature's idea of a training loop.

Then ontogenetic intelligence, which shows up when nervous systems appear roughly 550 million years ago. Now a single animal can learn inside its own lifetime, from its own senses, without the species having to kill off a cohort to try again. Much faster. You can imagine the relief.

Then societal intelligence, which arrives with language and later writing. A group of humans learns faster than any one human alone. Civilization as distributed cognition.

And at the top of the ladder, artificial intelligence in the sense the 1956 Dartmouth group meant when they coined the term. Not animal cognition running on silicon. Something newer. The ability to form abstractions, reason with symbols, generate hypotheses. Knowledge that isn't downstream of experience.

Participants of the 1956 Dartmouth Summer Research Project on Artificial Intelligence.
Fig. 01Dartmouth, 1956. Where the term was coined, and the rung of the ladder we did not actually climb.

The quiet claim the book is making, and it took me a re-read to catch it, which tells you how carefully I read the first time, is that what we've built in the last decade of the so-called AI renaissance isn't the top of that ladder. It isn't McCarthy's program at all. It's Norbert Wiener's, from 1948. Cybernetics. Closed-loop feedback, encoding, decoding, prediction, control. Animal-level cognition at industrial scale, because we finally have enough silicon and enough data to physically realize the machines Wiener could only sketch by hand. Their phrase for the current moment is the renaissance of cybernetics. I don't think I can unread that.

First edition of Norbert Wiener's Cybernetics: or Control and Communication in the Animal and the Machine, 1948.
Fig. 02Wiener, 1948. Hermann & Cie in Paris, with The Technology Press and Wiley. The sketch it took eighty years and a mountain of silicon to run.

I closed the book, went back to work, and said the word "agentic" four times in standup the next morning without irony. Which is fine. It's my job. I build multi-agent systems on top of frontier LLMs, I'm reasonably good at it, and I bill by the hour for exactly the thing I'm about to complain about. I just couldn't stop seeing the shape of the thing anymore.

The first thing that started landing differently was something I'd half-noticed for months without stacking it up. Sometime in the last year, the people actually running this industry quietly stopped saying "AGI."

Sam AltmanOpenAI
“not a super useful term.”
Dario AmodeiAnthropic
“I’ve always disliked it.”
Marc BenioffSalesforce
the AGI narrative is “hypnosis.”
Daniela AmodeiAnthropic
“outdated.”
Satya NadellaMicrosoft

any self-declared AGI is “benchmark hacking.”

Which is funny, because Microsoft literally co-wrote the AGI clause in OpenAI’s own contract.

Individually, each of these is a reasonable adult being precise about terminology. Put them in a row and it reads differently. Five years of "the singularity is right there," and then, roughly in unison, a careful walk-back from the word itself. You don't walk back from a story you're confident in. You walk back when the story is under strain.

The second thing was the money, which I went looking at properly in July, mostly out of morbid curiosity.

The build-out isn't slowing. It's accelerating. 2026 combined capex across Microsoft, Amazon, Alphabet, Meta, and Oracle is estimated around $720B. Capex is running at 34% of revenue. Free cash flow across the group is going negative for the first time in 35 years. OpenAI's telegraphed datacenter pipeline alone is $1.4T. Meanwhile the entire pure-play AI vendor cohort, every name you've heard of, is on track for somewhere south of $35B in combined 2026 revenue.

Supplier-side capex is more than twenty times customer-side revenue. I did that division twice on my phone, the way you recheck an ATM balance.

A hyperscale AI datacenter build-out under construction.
Fig. 03$720B of 2026 capex, and free cash flow across the group going negative for the first time in 35 years.

If this were a research push, the marginal dollar would go to the research problems that are actually open. Causal reasoning. Persistent memory. World models. Compositional generalization. Embodied grounding. Energy efficiency at anything close to a human brain. None of that is where the money is going. The money is going into GPUs that depreciate 20% a year, into racks, serving inference for productivity software. Coding assistants, customer support, marketing copy. Useful, real, sometimes profitable. Nothing in there is on the road to McCarthy.

Goldman has a phrase for the customer side of all this. "AI productivity beneficiaries." It's an accurate phrase. It is also not a phrase anybody has ever used about a scientific revolution. Nobody called the people downstream of general relativity relativity productivity beneficiaries.

Stanford's 2026 AI Index puts enterprise adoption at 78%, with measurable gains concentrated in a narrow band: code, support, content. That is the shape of a cybernetic deployment. Closed loops, fast feedback, recoverable failures. Animal-level cognition applied to industrial-scale tasks, and it works. It's just not what was on the box.

The third thing hit in a cinema hall, which is not the most expected place to read papers, I'll grant you. I was benchmarking model generations on my phone during the interval.

The improvement-per-dollar curve had stopped looking like a hockey stick. It looked, on a linear plot, tired.

The Chinchilla math makes it concrete. Training a trillion-parameter model at Chinchilla-optimal ratios needs about 20T tokens of high-quality text. Estimates of the total open web sit somewhere between 10T and 50T. There's a ceiling and we're in the room with it.

Four-panel Chinchilla scaling analysis: IsoFLOP loss curves with the compute-optimal locus, loss versus training compute against an irreducible floor, optimal parameter and token counts, and loss gained per decade of compute.
Fig. 04The same argument, worked out. (a) IsoFLOP curves and the compute-optimal locus. (b) Loss against compute, flattening onto an irreducible floor at E = 1.69. (c) Optimal parameters and tokens scaling as C^0.45 and C^0.55. (d) What a decade of extra compute actually buys: 0.190 in loss at 10^21 FLOPs, 0.066 at 10^24, 0.023 at 10^27.

The industry response has been honest engineering, to be fair. Mixture-of-experts. Inference-time compute. Reasoning models that "think longer." Synthetic data. Distillation. PEFT. All of it good work. All of it trying to extract more from the same paradigm. None of it asking whether the paradigm itself is the constraint. There's a long-horizon execution paper from a few months back showing that the same model which nails a hard problem in one shot will fail on a longer, easier one. Reliable multi-step execution isn't a scaling problem. It's an architecture problem, and architecture problems don't respond to capex.

And then there was LeCun, who had already left before I started the book, but who I only really understood in retrospect.

He walked out of Meta in January after twelve years at FAIR. Two months later his new lab, AMI, raised $1.03B in seed at a $3.5B valuation. Largest seed round in European history, which is a sentence that would have sounded made up in 2019.

The quote from him in MIT Tech Review is the one I keep coming back to. LLMs, he said, are now technology development, not research. He told academics, in print, to stop working on them.

Plenty of smart people disagree with him about world models. Hassabis publicly does. Amodei does. Altman does. Reasonable people can fight all day about JEPA versus autoregressive scaling and I don't think the technical question is settled.

But the technical question isn't the point of his exit. His argument is institutional, and the institutional one is harder to dodge. What gets called "AI research" in 2026 is mostly the engineering of a known paradigm at industrial scale. That is necessary work. Someone has to do it. It just isn't where a fundamentally different kind of intelligence gets discovered. A scientist scaling a known thing has stopped doing science.

One of the people who built this field walked out of its biggest lab to do research that lab was no longer interested in funding. That is not a story about Meta.

Last night all of it finally sat down in the same room together, which is why this exists.

It's almost light out here at 5am. The agent run I was babysitting finished a while ago. The client will open the output in the morning and it'll work. It usually works. That's the part that makes this awkward to write, because nobody is being defrauded here. The thing does the thing.

But I watched it run with that ladder in my head and I couldn't stop seeing what it actually is. Perceive, encode, predict, decode, act, take the feedback, adjust. That's the whole architecture. It is a beautiful closed loop and it is a 1948 idea with better hardware.

On that ladder, on a good day, I'm building extraordinary cybernetics. On a bad day I'm tuning prompts. Not artificial intelligence in the sense the Dartmouth workshop meant. And I don't think anyone in this industry is, at the scale the funding suggests they are.

What bothers me isn't that we built cybernetics. Cybernetics is a real accomplishment. Wiener saw something profound about closed-loop feedback in animals in 1948, and it took eighty years and a mountain of silicon to make his sketch run, and it runs. Good. What bothers me is the pretense that we built the other thing. That hundreds of billions of dollars are being allocated against a story the people telling it have quietly stopped believing.

There's one line from the book I keep coming back to. I'm paraphrasing, but it says something like: we cannot know what remains to be done in the science of intelligence until we honestly understand what we've already done. We haven't. We're too busy building.

There's a version of my next ten years where I stay in the industry, keep shipping agents, watch the valuations. There's another where I go work on the parts of intelligence that aren't just closed loops in silicon. I don't know yet which one I end up in. Writing this out over two months helped me see which one I'm actually curious about.

Look at what we built. Ask honestly what it is. Then go figure out what's missing.

Aryan Gurav — MumbaiMore writingTop