I would like to share a broader thought about how I think AI development should progress in the long term.
I see an interesting parallel with the evolution of CPUs.
In the earlier years of personal computing, progress was often associated heavily with clock speed:
MHz → higher MHz → GHz → higher GHz
For a period, it felt as though a faster processor mainly meant increasing clock frequency.
But eventually the CPU industry ran into practical limits such as heat, power consumption and diminishing returns. Progress then shifted toward better architectures, multiple cores, better performance per watt, specialised hardware, and eventually combinations such as performance cores and efficiency cores.
In other words, the industry gradually moved from:
“How fast can one core run?”
toward:
“How much useful work can the whole system perform efficiently?”
I think AI may eventually need to make a similar transition.
Today, improvements in AI are still often associated with:
larger models
more parameters
larger context windows
more GPUs
more inference compute
more reasoning tokens
longer reasoning time
These approaches can absolutely improve capability, but continuously increasing compute may not be the most sustainable long-term direction by itself.
I think an equally important target should be:
How can we obtain the same or better intelligence while using much less compute, memory, energy and time?
Some useful measures of progress might therefore become:
useful work per token
useful work per GPU-hour
useful work per joule
useful work per second
reasoning quality per unit of compute
successful agent tasks per unit of compute
For example, imagine a difficult task that today requires:
15 minutes + very high reasoning effort + 100 units of compute
A future model that achieves a better result in:
2 minutes + 20 units of compute
would, in my view, represent a major form of progress even if the model itself were not dramatically larger.
This is also why I find developments such as context compaction, more intelligent memory management, reusable agent skills, retrieval, mixture-of-experts approaches, specialised models and better reasoning efficiency particularly interesting.
Rather than processing everything with maximum compute all the time, a future AI system might dynamically decide:
which information deserves active attention
which context can be compressed
which task needs a powerful reasoning model
which task can use a smaller or faster model
which previous work can be reused instead of recomputed
which subtasks can run independently
This could become somewhat analogous to modern processors using different cores for different workloads.
A future AI system might have:
high-performance reasoning for difficult problems
efficient processing for routine work
specialised models/tools for particular tasks
memory and retrieval so previous work does not need to be repeatedly recomputed
context management that keeps only the most useful information active
I believe this direction could have several benefits:
lower inference cost
lower electricity consumption
reduced data-centre requirements
faster responses
longer-running agents
higher usage limits
lower API costs
better accessibility of advanced AI
more sustainable scaling
The long-term goal, in my opinion, should not simply be:
“Use more compute to obtain more intelligence.”
It should increasingly become:
“Obtain more intelligence from every unit of compute.”
Just as CPUs eventually moved beyond the simple clock-speed race, I think AI may need its own transition from a mainly compute-scaling era toward an era of architectural and efficiency improvements.
This does not mean stopping the development of larger or more powerful models. Both approaches can happen together.
My hope is that future generations of AI can become simultaneously:
more capable + faster + cheaper + more energy-efficient.
I would be interested to hear how others think AI efficiency should be measured and whether “useful intelligence per unit of compute” should become a much more important benchmark for future models.