AI is either as good as it gets, or the worst it'll ever be
Why cost, not capability, decides what AI becomes.
Contents
There are two ways to look at AI. One is that this is roughly as good as it gets, the curve is flattening, and we are looking at the best these systems will ever be. The other is that this is the worst version you will ever use again, a clumsy early prototype you will laugh at in three years.
The loudest objections to AI are that it is expensive and that it is dumb. It just predicts the next word. It is a statistical model with no understanding, a stochastic parrot, fancy autocomplete. Every one of those descriptions is a claim about how the machine works. None of them is a claim about whether what comes out is useful.
In 1984, Edsger Dijkstra caught the shape of this discussion. Asked whether machines can think, he wrote, it is about as relevant as asking whether submarines can swim. A submarine does not swim, it crosses the ocean anyway.
That is the whole rebuttal to the parrot. You can be completely right that it is “just” next-token prediction and it changes nothing about whether the contract it drafted holds up, whether the code it wrote runs, whether Pelaris works. People are building real things on this today. Whatever AI “really” is under the hood is a question for philosophers, and I have products to ship and enterprises to change.
So for a builder, capability is already off the table. The model is good enough for the things most builders need. The only question left is whether it can get value, and that is a question about cost.
Whether AI earns its place in a product is no longer a question of how smart it is. It is a question of how much it costs to run.
The floor has already moved, and it does not move back
Start with the pessimistic case, because even there the world has changed permanently. Suppose capability stops dead today. No more frontier jumps, no smarter models, this is the ceiling. The floor has still moved.
Three years ago, software could not reliably draft a contract, summarise a research paper, write a working function from a plain-English description, or hold a coherent conversation about your codebase. Now, summarising and generating text has become boring. Any developer (or even non-developers) can call that capability for a fraction of a cent to build applications. This is not a forecast. It already happened.
A capability freeze does not un-happen what has already shipped.
The floor is already high enough to build a real product on, and everything above it is upside. So the genuinely interesting question is not whether the floor moved. It is whether the floor is the ceiling.
The ridicule has been relentless, and aimed at the wrong axis
None of this stopped the dismissals. Since ChatGPT launched on 30 November 2022 and reached an estimated 100 million monthly users within about two months, the fastest ramp for a consumer app at the time, the dominant register around large language models has been that they are less impressive than they look.
The most durable was “stochastic parrots”, from Bender, Gebru and colleagues in 2021, arguing these models stitch together word sequences “without any reference to meaning”. Then came the popular shorthand: fancy autocomplete, pattern-matching with no comprehension. Gary Marcus has argued since March 2022 that pure scaling would “hit a wall” and the models would plateau. As recently as late 2024, reporting on OpenAI’s Orion suggested scaling had stalled and the strategy had stopped working.
Two things are true about these critiques at once. They were often right about a specific weakness at a specific moment, and they were aimed at a question that had already stopped mattering. The parrot argument is a claim about mechanism. Whether you can build a business on the output is a claim about results. The critics won the first argument and lost track of the second, which was the only one the rest of us were having.
And even on their own chosen ground, capability, the ceiling calls kept arriving just before the next jump. The pattern is the point, not any single benchmark.
- GPT-4 (March 2023) lifted MMLU from GPT-3.5’s 70.0% to 86.4%, per OpenAI’s technical report. The much-quoted “90th percentile bar exam” claim from that report was shown to be overstated, closer to the 62nd to 69th percentile against first-time takers, so the MMLU leap is the cleaner figure to lean on.
- ARC-AGI, a benchmark built specifically to be a wall for LLMs, fell in December 2024 when OpenAI’s o3 scored 75.7% at low compute and 87.5% at high compute against GPT-4o’s roughly 5%. ARC Prize called it “a surprising and important step-function increase”. The 87.5% run was also enormously compute-expensive, which is a useful tell for a piece about cost.
- The International Mathematical Olympiad fell in July 2025, when Google DeepMind’s Gemini Deep Think reached gold-medal standard, scoring 35 out of 42 in the official contest window.
Hallucination is real and remains the honest, major failure, and it is why “good enough” always needs a “for what”. It is good enough for a large and growing class of tasks, and demonstrably not good enough where a confident fabrication is fatal. But the calls that the ceiling had been reached kept landing just before the next jump. I am not saying the next jump is guaranteed, and it does not need to be, because the argument does not rest on it.
The releases themselves are arriving faster too. Frontier models used to land years apart. Now they come in a steady stream, from more labs at once.
The gap between genuine capability jumps has compressed from years to months, and the field has widened from one lab to half a dozen. Each dot is a real model. The density on the right is the argument.
Watch a man eat spaghetti
The cleanest visceral proof of “it only got better” is not a benchmark table. It is a meme. In March 2023 a Reddit user posted a clip of Will Smith eating spaghetti generated with ModelScope, and it was grotesque. His face morphed, his hands fused with the fork, the pasta floated. Forbes called it an “eldritch abomination”. The community adopted the prompt as an informal benchmark, a way to track how fast AI video was improving.
It improved fast. By May 2025, Google’s Veo 3 had effectively passed the test, producing a photorealistic version with natural chewing, believable fork-and-pasta physics, and, an industry first for a consumer model, natively generated synced audio (if oddly crunchy chewing sounds). Forbes ran the headline that Google had passed it, judging the output far closer to real footage while noting tells still remained. Two years. Eldritch abomination to near-indistinguishable.
Will Smith himself joined in. In February 2024 he posted his own parody to Instagram, pretending to be the AI version, captioned “This is getting out of hand!”. He was not wrong.
None of that settles whether the model understands a fork, and it does not need to. The same cheap, capable generative stack now reaches well past video. It is remaking text and images too, where the effect has been to explode content, not kill it, and it has got cheap enough that anyone can spin up an entire silly persona site in an afternoon, as I did with Bradsolutely. Whatever it is, you can build with it now.
Revolutions feel slow, then sudden
There is a reason the ridicule cycle keeps mistiming the trajectory. General-purpose technologies do not arrive on a smooth ramp. They incubate quietly, then climb almost vertically, and the people living through the quiet phase mistake it for the whole story.
The personal computer is the cleanest example. The hardware all arrived in a tight window: the Altair 8800 in January 1975, the Apple II in 1977, the IBM PC in August 1981. Then adoption crawled. US household ownership was only around 15% in 1989 and did not cross 50% until roughly 2000, per Our World in Data. Twenty-five years from launch to majority. And the incumbent did not collapse gradually. Smith Corona, the last US typewriter maker, filed for Chapter 11 on 5 July 1995, a full decade after PCs already dominated word processing. Slow, slow, slow, then the floor falls out from under the old thing all at once.
The detail worth noticing is that the climb keeps getting steeper. The PC took about 25 years to reach a household majority. Smartphones went from 35% of US adults in 2011 to around 91% by 2024.
AI itself is an example of this. Generative AI stems from a 2017 Google research paper, Attention Is All You Need, and the revolution was slow: five years of curiosities like Facebook researchers quietly redirecting a 2017 experiment after two negotiation bots drifted into an efficiency shorthand, the story the press turned into “Facebook shut down an AI that invented its own language”, and a Google engineer going public that LaMDA was sentient in June 2022, a claim the field rejected, until ChatGPT exploded at the end of 2022.
Each general-purpose technology incubates, then climbs, and the curve compresses with each generation. The reason it compresses is the same reason this whole piece turns on cost: each new technology gets cheap faster than the one before it, and cheap is what pulls the climb forward.
Capability decides what AI can do. Cost decides how much of it you can afford to use.
Cost is the real AI variable
Even if you grant the pessimists everything, even if capability froze exactly where it is today, AI would still deliver enormous value, on one condition. The cost has to keep falling.
A literal, total research freeze would slow the cost fall, since a chunk of it comes from distillation, and distillation needs a live frontier to distill from. Freeze everything and you keep the hardware gains, roughly 30 to 40% a year, and lose most of the algorithmic tenfold. Cost still falls. It just falls slower. The argument does not need the fast rate, only a falling one, and a falling one is what every mechanism below delivers.
The logic is simple. A frozen-capability model that costs a meaningful amount per query only justifies itself for high-value tasks. The same frozen model at a hundredth of the price justifies itself for a hundred times more uses, including ones nobody bothers to attempt today because the economics do not clear. For most products, cost is the binding constraint, not intelligence. Cheaper inference is also what makes a return easier to find, because it lowers the price of an experiment: you can afford to try ten uses to land the two that pay off. That fits what the 95% no-return numbers actually show. Model capability is rarely what holds a deployment back. Process, interface and cost are, and cost is the one falling fastest.
The good news is that the price of a fixed unit of AI capability has been collapsing.
The headline framing is a16z’s “LLMflation”: the cost of a model at a given quality tier falls by roughly tenfold a year. Their cleanest datapoint sits at a low bar. A GPT-3-quality model (MMLU around 42) cost about $60 per million tokens at launch in late 2021 and about $0.06 by late 2024, a thousandfold drop in three years. Hold the bar higher and the fall is smaller but still steep: on the same a16z data, GPT-4-class quality is down around 60 times since early 2023. The load-bearing number is the range, not the headline, and Epoch AI’s independent analysis supplies it: a median fall of around 50 times per year across benchmarks, from 9 to 900 times depending on the task, and accelerating since January 2024. The market felt it directly when DeepSeek’s R1 launched in January 2025 priced at roughly a 27th of o1’s per-token rate for comparable reasoning, and wiped $600 billion off Nvidia in a single day.
Falling cost is the base case, not optimism
It would be easy to read those numbers as a recent fluke. The empirical fall is not in doubt, a16z and Epoch both measure it, and the mechanism is concrete: better algorithms, distillation and cheaper hardware, which I come back to below. What the history adds is not proof, it is precedent. Steep, predictable cost decline is what maturing technologies tend to do, and the shape has a name.
It is the experience curve, or Wright’s Law: cost falls a consistent percentage for every doubling of cumulative production. The Ford Model T followed it almost exactly, the touring car dropping from around $850 at launch to $290 by 1925 as production scaled. The same shape shows up everywhere you look, documented in Our World in Data:
- Solar PV modules fell about 99.6% from $106 per watt in 1976 to $0.38 per watt in 2019, a roughly 20% drop for every doubling of installed capacity (Swanson’s Law).
- Lithium-ion batteries fell more than 99%, from about $9,200 per kWh in 1991 to roughly $78 per kWh in 2024, at a learning rate near 19%.
- The cost of artificial light fell about 99.97% over two centuries, from $785 per thousand lumen-hours in 1800 to 23 cents by 1992 in real terms, work the economist William Nordhaus used to argue our price indexes badly understate real progress.
- DNA sequencing fell from $95 million per genome in 2001 to roughly $525 by 2022, explicitly outpacing Moore’s Law after next-generation sequencing arrived in 2008.
These are not proof that inference obeys the same learning rate. They are physical-goods curves driven by cumulative production, and AI inference falls for its own reasons, mostly software. But the shape rhymes, and a steep, sustained decline is the single most common behaviour of a maturing technology. That is the company AI’s cost curve is keeping.
Betting that AI inference keeps getting cheaper is betting with the grain of how technologies mature, not against it.
What falling cost unlocks
If cost keeps falling and capability also keeps improving, the combination is what changes the product surface. Cheap, capable inference is what makes always-on agents, long-horizon reasoning, and per-task economics viable. The interesting uses are not the ones we run today and wish were cheaper. They are the token-heavy ones nobody runs yet, because at today’s prices the economics do not work. Take an always-on coach that re-reasons over an athlete’s entire history on every interaction, exactly the kind of thing I want Pelaris to do. Run that for every user, every day, and the per-query cost sinks it. At a tenth of the price it clears, so I am building for it now and letting the cost line come to meet the feature. It is the cost side of a point I have made before, that for most products the interface, not the model is the bottleneck now.
The mechanics of the cost fall are worth understanding, because they are concrete rather than hopeful. Epoch AI found the compute needed to hit a fixed performance level halves roughly every eight months, far faster than Moore’s Law, and performance per dollar improves around 30 to 40% per year on top of that. Better algorithms and cheaper chips stack. Distillation does the rest: today’s frontier capability becomes tomorrow’s cheap commodity, and DeepSeek’s distilled 32B model already beats GPT-4o and outscores o1-mini on maths benchmarks like AIME, at a fraction of the size. Competition, as the DeepSeek price shock showed, supplies the pressure.
Counter-Case: AI Training Costs
The cost story does not run only one direction, and pretending otherwise would undercut the rest of the discussion. This is also where I want to be exact about which cost claim I am making, because there are two, and only one of them is solid.
The solid one is per-query cost, the price to run a fixed unit of capability. That is the number a16z and Epoch measure, the number builders pay. It is falling fast and reliably.
The claim I am not making is that total AI spend, for businesses and for AI training, falls. It does not follow, and I would not defend it. Frontier training costs are rising about 2.4 times per year and are projected to exceed $1 billion per model by 2027, so pushing the frontier gets more expensive even as serving a fixed capability gets cheaper. The two trends move in opposite directions. Data-centre electricity demand is set to roughly double to around 945 TWh by 2030 with AI as the main driver, though the IEA’s own caveat is that the binding constraint is local grid concentration, not global supply. And there is Jevons paradox, which Microsoft’s Satya Nadella invoked after DeepSeek: cheaper AI does not guarantee lower total spend, because it unlocks new uses faster than it saves money on old ones.
That is the honest boundary of the argument. Total spend rising is a macro question, and it does not stop any individual product from clearing its own economics.
The bet worth making
That leaves builders and enterprises with an asymmetry. If you design for today’s prices and today’s quality, you are designing for the worst version of the inputs you will ever have. Building on falling-cost, rising-capability infrastructure is a different discipline from building on a fixed platform, and is hard for teams to adjust to.
So, is this the best AI is going to be, or the worst it will ever be? Ask it and you are discussing capability, which stopped being the primary question the moment you could build something real on the output. Dijkstra thought the question itself was beside the point. The model does not think. Like his submarine, it crosses the ocean anyway. The only variable left that can decide what AI becomes is cost, and the cost that matters to anyone deciding what to build is falling with the grain of every technology that came before it.
Toby Keith once sang that he ain’t as good as he once was. That is not a problem these models have. They have only run in one direction, up, while the cost of running them keeps running the other way, down. Neither of those is the thing worth deciding. The thing worth deciding is what you are going to build on top of them and how you will use them to transform enterprises.
Read next
Bradley Hunt
AI, engineering & leadership