AI progress should focus on efficiency, not only more compute: lessons from CPU evolution

I would like to share a broader thought about how I think AI development should progress in the long term.

I see an interesting parallel with the evolution of CPUs.

In the earlier years of personal computing, progress was often associated heavily with clock speed:

MHz → higher MHz → GHz → higher GHz

For a period, it felt as though a faster processor mainly meant increasing clock frequency.

But eventually the CPU industry ran into practical limits such as heat, power consumption and diminishing returns. Progress then shifted toward better architectures, multiple cores, better performance per watt, specialised hardware, and eventually combinations such as performance cores and efficiency cores.

In other words, the industry gradually moved from:

“How fast can one core run?”

toward:

“How much useful work can the whole system perform efficiently?”

I think AI may eventually need to make a similar transition.

Today, improvements in AI are still often associated with:

larger models

more parameters

larger context windows

more GPUs

more inference compute

more reasoning tokens

longer reasoning time

These approaches can absolutely improve capability, but continuously increasing compute may not be the most sustainable long-term direction by itself.

I think an equally important target should be:

How can we obtain the same or better intelligence while using much less compute, memory, energy and time?

Some useful measures of progress might therefore become:

useful work per token

useful work per GPU-hour

useful work per joule

useful work per second

reasoning quality per unit of compute

successful agent tasks per unit of compute

For example, imagine a difficult task that today requires:

15 minutes + very high reasoning effort + 100 units of compute

A future model that achieves a better result in:

2 minutes + 20 units of compute

would, in my view, represent a major form of progress even if the model itself were not dramatically larger.

This is also why I find developments such as context compaction, more intelligent memory management, reusable agent skills, retrieval, mixture-of-experts approaches, specialised models and better reasoning efficiency particularly interesting.

Rather than processing everything with maximum compute all the time, a future AI system might dynamically decide:

which information deserves active attention

which context can be compressed

which task needs a powerful reasoning model

which task can use a smaller or faster model

which previous work can be reused instead of recomputed

which subtasks can run independently

This could become somewhat analogous to modern processors using different cores for different workloads.

A future AI system might have:

high-performance reasoning for difficult problems

efficient processing for routine work

specialised models/tools for particular tasks

memory and retrieval so previous work does not need to be repeatedly recomputed

context management that keeps only the most useful information active

I believe this direction could have several benefits:

lower inference cost

lower electricity consumption

reduced data-centre requirements

faster responses

longer-running agents

higher usage limits

lower API costs

better accessibility of advanced AI

more sustainable scaling

The long-term goal, in my opinion, should not simply be:

“Use more compute to obtain more intelligence.”

It should increasingly become:

“Obtain more intelligence from every unit of compute.”

Just as CPUs eventually moved beyond the simple clock-speed race, I think AI may need its own transition from a mainly compute-scaling era toward an era of architectural and efficiency improvements.

This does not mean stopping the development of larger or more powerful models. Both approaches can happen together.

My hope is that future generations of AI can become simultaneously:

more capable + faster + cheaper + more energy-efficient.

I would be interested to hear how others think AI efficiency should be measured and whether “useful intelligence per unit of compute” should become a much more important benchmark for future models.