← Back to all posts
Opinion

Fast AI Won't Speed Up the Old Jobs. It Will Reorganize What Gets Built.

July 6, 2026 · Opinion
Fast AI Won't Speed Up the Old Jobs. It Will Reorganize What Gets Built.

TL;DR

On a recent episode of No Priors, Cerebras CEO Andrew Feldman made an argument worth sitting with, because it is about where AI goes next rather than which model won this week. His claim: fast inference is not an upgrade to today's AI products, it is a phase change that will reorganize what gets built. His framing quote is the one to keep. "How big is the market for slow search? Zero. How big is the market for dial-up internet? Zero. That is how big the market for slow inference will be." We think the core of that is right and the timeline is the part to be skeptical about. Speed is quietly becoming the axis that decides which AI products are even possible, but the "entirely new business models" it promises are, as always, easy to assert and hard to schedule. Here is our read.

every platform shift runs in two phases PHASE 1 · REPLACE the old job, done faster chat · coding · design tools (where AI is now) PHASE 2 · REORGANIZE the job itself changes new workflows · new businesses (where the value is) Precedent: the PC replaced typewriters (phase 1), then the cloud and SaaS reorganized how work happens (phase 2). The second jump is the big one.
Feldman's real point: AI is still in phase one, doing familiar jobs faster. The productivity leap comes in phase two, when work reorganizes around it.

Why speed is suddenly the whole game

For years Cerebras was a curiosity: a company building a single chip the size of a dinner plate, roughly 46,000 square millimeters against everyone else's postage-stamp GPUs, claiming inference 15 to 20 times faster and getting mostly shrugs. Then, as Feldman tells it, two things changed at once. Models got smart enough to use every day around 2025, and the moment something is in your daily workflow, latency stops being a nicety and becomes existential. Nobody waits for a slow website, and nobody will wait for a slow agent. That shift is why Cerebras went from novelty to the largest IPO of 2026 (it raised $5.5B in May and popped to roughly a $70B market cap) on the back of a 750-megawatt, $20B-plus OpenAI deal and an AWS partnership.

Here is the deeper reason speed matters more than it used to, and it is the part the soundbite skips. Reasoning and agentic models do not answer in one shot. They think in loops, call tools, check their own work, and burn tokens by the thousand-fold compared to a single chat reply. When a task is one model call, a slow model is annoying. When a task is a hundred chained calls, a slow model is unusable, because the latency compounds a hundred times. So test-time compute, the whole "let the model think longer" era, quietly turned inference speed from a spec-sheet bragging right into the thing that gates which agentic products can exist at all. That is the strong version of Feldman's claim, and it holds up.

Where it is already going: watch the coders

If you want a preview of phase two, Feldman handed over a concrete one from inside his own company. Cerebras went from spending under $1,000 per engineer per month on AI tokens to roughly $25,000 to $30,000 in about eight months. And a subset of his engineers, the ones whose brains fit the new shape, went from being 10x developers to something like 100x by running eight or ten agents around the clock and shifting their job from writing code to governing agents: one to build, one to QA, one to paper over the models' habits. That is not "coding, but faster." That is a different job with a different unit of work, which is exactly what phase two looks like when it arrives.

AI token spend per engineer at Cerebras, ~8 months apart (per Feldman) ~8 months agounder $1k / mo now$25-30k / mo The heaviest users went from 10x to 100x by governing fleets of agents, not by typing faster. Vendor anecdote, but a telling one.
A 25x jump in per-engineer inference spend in under a year is what the reorganization looks like on an expense report.

The Netflix tell, and what it actually predicts

Feldman's favorite analogy is Netflix: it mailed DVDs, assumed its rival was Blockbuster, and when the internet got fast it did not become a better DVD mailer, it became a streaming service and then a studio, an entirely new business. The useful part of the analogy is the prediction embedded in it. The first thing a step-change in speed does is make the old thing cheaper and faster, which is obvious and visible. The valuable thing it does comes later and looks absurd in advance: a DVD-by-mail company buying film studios made no sense in 2005. Translate that forward and it says the AI products that matter most in a few years are the ones that will sound ridiculous to pitch today, because they assume inference is so fast and so cheap that you would call a model hundreds of times where you now call it once. Continuous background agents, real-time multi-agent systems, software that simulates and self-checks before it answers: uneconomic and too slow now, ordinary later.

Our take: right thesis, salesman's clock

Two things can be true. The thesis is right, and the man delivering it is talking his book. A hardware CEO telling you that speed is about to reorganize the economy is a bit like a barber telling you that you need a haircut, which does not make him wrong but does mean you should check the mirror yourself. Three honest asterisks temper the pitch. First, the "new business models" claim is the oldest song in infrastructure; railroads, electricity, and the cloud all sang it, and all delivered, on a timeline far longer and messier than any founder's slide predicted. Second, the 100x story is survivorship talking: Feldman freely admits most of his own company, himself included, is "limping along" trying to make these tools fit their jobs. The reorganization is real at the frontier of adopters and barely started at the median desk. Third, the same token explosion that signals the shift is also detonating budgets, which is why companies from Tesla to Uber spent this year slapping caps on AI spend. We wrote about that in our piece on the AI spending caps, and it is the counterweight to the speed story: phase two only arrives if inference keeps getting both faster and cheaper at once, not just faster.

Where we think this lands

Our prediction, stated so you can hold us to it. Over the next year or two, the AI products that win are the ones architected for abundant, fast inference, agents that loop, verify, and run continuously, rather than thin chat wrappers around a single call. Tokens-per-second and time-to-first-token graduate from engineer trivia to headline product specs, the way page-load time did for the web. "Agent governor" becomes a real and well-paid skill. And a handful of genuinely new businesses appear that look, from here, faintly ridiculous, which is precisely how you will know they are the phase-two ones. Feldman is selling the shovels, so discount the urgency. But the direction he is pointing, away from faster old jobs and toward reorganized new ones, is the right direction to be looking.

Key Takeaways

  • Cerebras CEO Andrew Feldman argues fast inference is a phase change, not an upgrade: like fast internet turning Netflix from a DVD mailer into a studio, speed reorganizes what gets built rather than just accelerating it.
  • The strong version holds: reasoning and agentic models chain hundreds of calls, so latency compounds and inference speed now gates which agentic products can exist at all.
  • The leading indicator is coding: Cerebras' per-engineer token spend jumped from under $1k to $25-30k a month in eight months, and its heaviest users shifted from writing code to governing fleets of agents.
  • The skeptic's asterisks: Feldman sells the hardware, the "new business models" promise always runs on a longer clock than pitched, most workers are still "limping along," and the token boom is also blowing up budgets.
  • Our call: near-term winners are products built for abundant fast inference (looping, self-verifying, always-on agents); speed metrics become headline specs; and the biggest new businesses will look absurd right up until they are obvious.

Sources: No Priors podcast (Andrew Feldman interview), TechCrunch (Cerebras IPO), Yahoo Finance (OpenAI & AWS deals)

AI infrastructureCerebrasinferenceagentic AIfuture of AIAndrew FeldmanNo PriorsAI economics
CONSOLE
$