I’ve learned to be suspicious of any tech conversation that starts with a countdown. “Two years to AGI.” “Five years until everything changes.” It’s a familiar rhythm: a little awe, a little fear, and a lot of implied inevitability. So when Amazon’s Peter DeSantis—one of the people sitting closest to the machinery of modern AI—basically says “not so fast,” I find myself leaning in.
Speaking at VivaTech in Paris, DeSantis laid out a view that feels almost unfashionably grounded: yes, progress is real, and yes, it’s fast—but **we’re still nowhere near the kind of AI that truly rearranges daily life at scale**. Not because of a lack of ambition, but because the stack—models, chips, systems, interfaces—still needs several major leaps, not just incremental polish.
Why “better” isn’t the same as “transformative”
DeSantis’ central claim is blunt: **AI needs “a few more orders of magnitude” of improvement** before it becomes “really interesting” in the deeper, society-shifting sense people like to promise. Over the past year or two, he noted, we’ve seen roughly an order-of-magnitude gain in efficiency for what models can do. The problem is that one order of magnitude is impressive on a chart and underwhelming in a world where humans expect fluidity, context, and speed without thinking about it.
What I hear in that framing is a quiet pushback against the idea that today’s generative AI is already the destination. It’s a powerful tool, sure. But if you’ve ever watched a model hesitate, lose the thread, or respond like it’s playing catch-up to your intent, you’ve felt the gap DeSantis is pointing at. **“Impressive” is not the same as “embedded in everything.”**
At Amazon, DeSantis oversees an organization that connects foundational AI models, custom silicon, and even quantum computing—essentially, the layers that determine whether AI is a clever demo or a durable platform. With decades inside the company, from early cloud infrastructure to chip strategy, his perspective carries the kind of operational skepticism hype tends to sand down.
Transformers won’t be the last word
One of the more telling parts of his argument is architectural. Today’s leading models are dominated by transformers, and they’ll keep improving. But DeSantis doesn’t sound convinced that transformers, as we currently use them, are the endgame for the kinds of interactions people ultimately want.
He expects **new model architectures beyond transformers**—not as a vague “something better,” but as a necessity driven by real constraints: responsiveness, cost, and the ability to deal with the messy, layered nature of human communication. If you’re waiting for a single model family to keep scaling until it becomes a universal mind, his view suggests you may be waiting for the wrong thing.
And here’s the point I find most reassuring: DeSantis argues that **humans will remain central to the most complex AI innovations for the foreseeable future**. That doesn’t mean “humans in the loop” as a buzzword. It means the hardest problems—judgment, meaning, responsibility, design choices that reshape work—don’t disappear just because a model gets better at predicting the next token.
The 40-millisecond problem: AI has to keep up with how we talk
We rarely think about the timing of conversation, but DeSantis does. He described human interaction as running on something like a **40-millisecond clock**. In real dialogue, we don’t just exchange words; we exchange signals: pauses, half-starts, little acknowledgments, shifts in tone, the micro-moments that tell you whether someone is confused, joking, offended, intrigued, or about to interrupt.
If AI is going to feel like a natural collaborator—especially in voice, video, or embodied contexts—it can’t respond like it’s reading from a slow teleprompter. It needs to “see” and interpret gestures, hesitations, and backchanneling in real time. And that’s not just a model challenge; it’s a systems challenge.
In plain terms: **latency becomes meaning**. A delayed response isn’t merely slower; it changes the social feel of the interaction. The future DeSantis gestures toward is one where AI isn’t impressive because it writes paragraphs, but because it can participate at human tempo without breaking the spell.
- Faster inference so interaction doesn’t stall
- Richer sensory understanding across speech, vision, and context
- New software approaches to orchestrate real-time behavior
- New hardware pathways that make that speed affordable
That last word—affordable—hangs over the whole conversation. The magic doesn’t scale if the price doesn’t come down.
Chips and models: the underrated feedback loop
DeSantis’ most operational point is also the most important for anyone who assumes “AI progress” is mainly about bigger models. He argues that the next wave depends on chips and models evolving together—tightly, deliberately, and ahead of time.
If chip teams don’t communicate what’s coming, model designers can’t prepare architectures or training methods to exploit it. Then, when the chips finally arrive, everyone spends months retrofitting software and research to match the hardware. In his telling, that lag is avoidable—and expensive.
When the collaboration works, it creates what he describes as a flywheel: **better models lead to better chips, which lead to lower cost and better efficiency, which then enables better models again**. That compounding effect is how a technology stops being a luxury and starts being infrastructure.
And there are already signs of it. DeSantis pointed to startups building “world models”—systems focused on simulating aspects of physics and environments rather than simply generating text—that are choosing Amazon’s Trainium and achieving close to double the compute efficiency compared with typical industry baselines. Even if you don’t care about the branding, the pattern matters: specialization and co-design are becoming the real accelerants.
So where does that leave the rest of us?
I walk away from DeSantis’ view with an oddly calming takeaway: **this is a starting line, not a finish line**. If you’re dazzled (or anxious) about what AI is today, his argument suggests we’re still in the phase where the medium is being invented—where interaction speed, architectures, and compute economics haven’t settled.
And if you’re wondering whether AI will replace human input in the most complex decisions, his answer is effectively: not in the way the loudest forecasts imply. Humans aren’t a temporary patch until the models “grow up.” Humans are part of the design space—because the hardest problems aren’t only technical. They’re social, organizational, and moral.
So here’s the question I can’t stop thinking about: when AI finally gets those “few more orders of magnitude,” will we be ready—not just with faster chips and smarter architectures, but with clearer ideas about what we actually want it to do?
Comments