
This year's AI headlines have been about intelligence. Foundation models like Astra and Fable solving Millennium Prize Problems. Startups building domain-specific foundation models for cybersecurity, law, or finance. Companies putting together “sovereign” intelligence strategies to own their own AI.
But a new story is becoming even more important: value. It started with open-weight models from companies like Kimi and Qwen, but now even OpenAI and Anthropic’s releases are largely focused on cost-per-intelligence/task, with deep discounts and token efficiency – like Luna with a 90% price cut over two generations.
Why the change? Because most new models are already smart enough to do most jobs. There are some exceptions, like math and physics. But for the vast majority of work, intelligence is now a commodity – shifting value back from models to the applications that surround them.
A test to measure intelligence is only good if it’s hard. A third-grade math test is great for evaluating third graders but useless for eighth graders, who should ace it every time.
AI models do the same to their benchmarks. A new test is made, and models do terribly. They get smarter and the scores improve. Then at some point, they score near 100%, and improvements no longer matter: the benchmark is “saturated”. A new, harder test is created, and the process repeats.

These benchmarks are artificial, but the same progression is true for real human work too. When it comes to most knowledge work – writing SQL queries, matching financial transactions, summarizing notes, etc. – models have learned all they need to know. Who cares about the next Fable release when Sonnet does the job just fine?
If the models are more than smart enough, why hasn’t AI automated all of our jobs? Because most work is not intelligence-constrained, it’s context-constrained.
We've used this example before: if you dropped Einstein into a back-office role at a large company, he wouldn’t be very effective. He’s not familiar with the processes; he doesn’t know the details of customers and vendors; he’s not aware of the historical business trade-offs. He’s more than smart enough, but doesn’t have the right context.
For many enterprise AI applications, swapping in the newest model no longer moves the needle much on quality. Far more gains come from better context engineering.
Cost and performance are often a context problem too. The smartest frontier model might complete the job, but spend a majority of the time (and cost) re-discovering how to do it each run. Better context engineering can dramatically transform the economics in 3 ways:
With excellent open-source models and new post-training techniques, it’s tempting to want your own custom model. We think post-training will be an important part of the AI toolkit, but not for the reason most people expect.
The key benefit of post-training is cost & performance. It allows you to run a workflow on a smaller model, and at scale, the savings or latency improvements can justify real investment. Custom models will not meaningfully increase model capabilities, especially for tasks when most models already meet the quality threshold.
When the decision is now about cost & performance rather than product differentiation, post-training is just one type of serving optimization. We think most teams will be happy to outsource it to companies like Applied Compute, Prime Intellect, and Baseten who compete on the intelligence-performance-cost balance as their whole business.
One of the worst insults you could call a startup was an “AI wrapper” – an application that gets whittled down to nothing as models get smarter.
That may have been true of some early AI apps, where careful prompting was needed to overcome model shortcomings. But when intelligence is no longer the limiting factor, smarter models don’t eat the application. Quite the opposite: context becomes increasingly valuable as agents operate, learn, and document information over time. Building these systems requires a deep understanding of the problem space, product intuition, and feedback loops that allow the best teams to improve faster than their competitors.
If you believe in this future, I'd love to hear from you: at@theoryvc.com.