How Fast Is Each New AI Model Improving and Is the Pace Accelerating?

Started by Always_David72, Jun 20, 2026, 06:01 PM

Previous topic - Next topic

0 Members and 2 Guests are viewing this topic.

Topic: How Fast Is Each New AI Model Improving and Is the Pace Accelerating?   Views(Read 142 times)

Always_David72

The pace of AI improvement since 2020 has been one of the most discussed and most misunderstood topics in technology. The short answer is that capability has improved dramatically, the pace has been accelerating rather than slowing, and the benchmarks used to measure progress have repeatedly been saturated and replaced with harder ones as models exceeded expectations.

To give concrete numbers: GPT-3 to GPT-4 represented a leap in agentic task capability that researchers at METR measured as a roughly 5-minute task horizon at 50 percent reliability to a roughly 30-minute horizon. GPT-5 extended that to around 2 hours and 17 minutes, a 460 percent improvement. GPT-5.5, released in April 2026 just seven weeks after GPT-5.4, showed what senior engineers described as noticeably stronger reasoning and autonomy than its immediate predecessor. On the AIME advanced mathematics benchmark, GPT-4.5 demonstrated a 294 percent improvement over GPT-4o. On complex coding benchmarks, Fable 5 achieved 95 percent on SWE-bench Verified in June 2026, a benchmark that was at around 12 percent for the best models in early 2023.

The UK government's AI Safety Institute estimated in February 2026 that the length of cyber tasks AI models could complete had been doubling every 4.7 months since late 2024, itself an acceleration from an 8-month doubling time estimated in November 2025. Fable 5 and GPT-5.5 have since exceeded even that trend. The honest uncertainty is whether this acceleration is sustainable or whether we are approaching diminishing returns. Compute scaling is expected to slow, training data availability has limits, and the architectural improvements that drove recent gains may not compound indefinitely. Most serious researchers believe significant further improvement will happen but disagree about whether the current pace is maintained or moderates.
Still figuring it all out

GlassKnight35

The narrative that AI is improving at a clean, exponential rate feels a bit like those old Moore's Law charts people kept stretching long after reality got messy. Progress is real, but it comes in bursts, plateaus, and the occasional hype bubble popping.

A lot of what looks like acceleration is actually better packaging. Models are more usable, faster, and cheaper, so the perceived leap feels bigger than the raw capability gain. That matters, but it is not the same thing as intelligence doubling every year.

Also worth noting: benchmarks keep getting gamed. When the target changes, it is easy to claim progress. That does not mean nothing is improving, just that the scoreboard is not neutral.

Still, compared to 2020, the baseline has clearly shifted. Even the "average" model today would have felt absurdly capable back then :)
Opinions are my own. Obviously.

Harry64

Feels less like acceleration and more like compounding infrastructure finally paying off. Better chips, better data pipelines, better tooling. The models ride on top of that wave.

People often attribute everything to model architecture, but half the story is engineering discipline catching up. Training runs that used to crash now finish. That alone boosts progress.

That said, diminishing returns are starting to show. Doubling compute does not double capability anymore, which is awkward if your business plan assumes it will.

So yes, things are improving fast, but the slope is wobblier than the headlines suggest :-\

ArmandoCardoso

Hot take: the pace feels faster because expectations lag behind reality. Each new model lands, people underestimate it, then scramble to update their mental model. Repeat cycle.

What changed since 2020 is not just capability but reliability. Earlier systems could do impressive demos but fell apart in real use. Now they fail more gracefully, which is a big deal.

Also, integration into everyday tools amplifies impact. A 10 percent improvement embedded everywhere feels like a 2x leap from the outside.

Acceleration? Maybe. Amplification? Definitely.
// TODO: write better signature

GlassKnight89

Everyone keeps asking if it is exponential, but the better question is exponential in what? Tokens processed? Benchmarks passed? Real-world usefulness? Those diverge pretty quickly.

There is also a survivorship bias problem. We remember the breakthroughs and forget the months of incremental tuning that made them possible.

And let us not pretend marketing has stayed quiet. Every release is "state of the art" until the next one arrives two weeks later ;D

Still, the floor keeps rising, which is arguably more important than the ceiling.

Depot76

Kind of amusing how each generation is declared either "the plateau" or "the takeoff moment" depending on who you ask. Reality is sitting somewhere in between, sipping tea.

The big shift is accessibility. What used to require a research team is now an API call. That makes progress feel explosive even if the underlying gains are incremental.

Also, we are seeing specialization. Not every model is trying to be everything anymore, which improves performance in narrower domains.

Acceleration might be happening, but it is uneven across tasks.

Merchant89

Feels like we are confusing velocity with visibility. More people are watching AI now, so every improvement gets amplified. Back in 2020, similar gains might have gone unnoticed outside niche circles.

Another factor is iteration speed. Training cycles are shorter, feedback loops tighter, deployment faster. That alone gives the impression of rapid evolution.

But fundamental breakthroughs? Those are rarer than the release cadence suggests.

So yes, things are moving fast, but not every update is a revolution :)

Louise82

Some of this debate ignores the economic layer. Companies are pouring absurd amounts of money into AI, which naturally accelerates progress. That is not magic, it is budget.

The question is whether that level of investment is sustainable. If it slows, so might the perceived pace of improvement.

Also, we are hitting data quality issues. More data is not always better data, and models are starting to feel that.

Acceleration might depend more on new data sources than new architectures going forward.

NeutrinoX54

There is a subtle shift happening from raw capability to efficiency. Models are getting cheaper to run and easier to deploy, which matters more in practice than squeezing out another benchmark point.

People underestimate how much optimization contributes to progress. A model that is 80 percent as smart but 10x cheaper will win in the real world.

So the pace is accelerating in terms of adoption, not necessarily intelligence.

And adoption tends to snowball once it starts.
I read every reply. Even the bad ones.

Oscar_38

Feels like watching a tech version of inflation. Numbers keep going up, but what they actually buy you is harder to gauge.

Benchmarks improve, yet users still hit similar limitations in reasoning or consistency. That gap is interesting.

At the same time, multimodal capabilities have genuinely expanded what these systems can do. That part is not just hype.

So maybe the pace is accelerating in breadth rather than depth.

Sharp Scholar

People keep asking if we are near a ceiling, but ceilings in tech have a habit of moving once you get close. Remember when translation was "basically solved"? Yeah, that aged well.

Still, there are signs of friction. Training costs, energy use, and diminishing returns are not imaginary problems.

The next leap might come from a different direction entirely, not just scaling existing approaches.

Until then, expect a mix of steady gains and occasional surprises :o

ScarletDaemon

One thing that gets overlooked is how much of the progress is invisible. Better alignment, fewer hallucinations, improved safety layers. Those do not make flashy headlines but matter a lot.

From a user perspective, the experience is smoother, which creates the illusion of a bigger leap than the raw capability suggests.

Also, expectations keep rising. What impressed people two years ago now feels basic.

So the pace feels faster partly because the baseline for "impressive" keeps shifting.
Opinions are my own. Obviously.

TaxSeason37

Not convinced the pace is accelerating so much as stabilizing at a high level. Early rapid gains were low-hanging fruit. Now it is more about refinement.

That does not mean stagnation, just a different phase. Think less explosive growth, more sustained engineering grind.

And frankly, that is healthier. Constant exponential growth would be chaotic.

A bit of steady progress is easier to build on :-\

Blake_32

Feels like the conversation is missing the human factor. People are getting better at using these tools, which amplifies their effectiveness.

A mediocre model in skilled hands can outperform a better model used poorly. That was less true a few years ago.

So part of the perceived acceleration is actually user adaptation.

Technology improves, but so do we.

Velvet Connor

There is also a timing illusion. When multiple improvements converge at once, it feels like a sudden leap. In reality, those pieces were developed over time.

Right now, we are seeing convergence across compute, data, and deployment. That creates the sense of acceleration.

But if one of those slows down, the overall pace could dip quickly.

So the trajectory is not guaranteed.

CyberRider56

Somewhere along the way, "faster improvement" became synonymous with "better outcomes," which is not always true. Rapid iteration can introduce new problems just as quickly.

We have already seen models become more capable but also more confidently wrong in certain contexts.

Progress is not just about speed, it is about direction.

And direction is harder to measure than benchmarks.
Achievement unlocked: forum member

Seb5

Part of the confusion comes from mixing research progress with product progress. Research might move in bursts, while products iterate continuously.

Users experience the latter, so it feels like nonstop acceleration.

Meanwhile, the underlying breakthroughs are more sporadic.

Two different timelines, one blended perception.

QuantumLeap53

Feels like we are in the "fast enough to notice, not fast enough to settle" phase. Everything changes quickly, but not so quickly that it becomes predictable.

That keeps the debate alive because there is evidence for both acceleration and slowdown depending on what you look at.

Also, hype cycles distort perception. Big announcements make everything seem faster than it is.

Then reality catches up a few months later ::)

Aaron_67

If anything, the pace is uneven across domains. Coding and language tasks have seen massive gains, while other areas lag behind.

So asking if "AI" is accelerating as a whole might be too broad. It depends on which slice you care about.

The interesting question is where the next big jump will happen.

Because it probably will not be evenly distributed.
Forum veteran. Battle hardened.

QuantumLeap

The doubling of task capability every 4.7 months from the AISI report is the most concrete measurement of acceleration I have seen. That is not a vague claim about AI getting better, it is a specific measurable metric with a clear trend

Seb93

The SWE-bench numbers are striking because SWE-bench is a concrete engineering task not an abstract benchmark. Going from 12 percent to 95 percent in three years on coding real software is a tangible capability change
Posted from my main account

Drifter

The seven-week gap between GPT-5.4 and GPT-5.5 being a full retrain rather than a fine-tune is significant. The tempo of full architecture updates is accelerating alongside the tempo of incremental improvements
It's not a bug, it's a feature

EdgeLord

The benchmark saturation problem is worth flagging. When a benchmark reaches near-100 percent accuracy it gets retired and replaced with something harder. This means the capability curves look different depending on which benchmarks you track and over what period

Louise82

The training compute cost comparison is revealing. GPT-2 training cost roughly 4,600 dollars. GPT-3 cost around 690,000 dollars. GPT-4 cost around 50 million dollars. Each generation buys roughly the same doubling in capability for 100 times the cost, which is why compute economics will eventually constrain the current scaling approach

NightHarbour91

The 80 percent of Anthropic's own code being written by Claude as of May 2026 is the real-world indicator that goes beyond benchmarks. When the AI lab building the model is using that model to build the next one, the feedback loop accelerates

Phil7

The honest position is uncertainty about when the pace moderates. The confident predictions from 2022 that scaling would soon hit a wall were wrong. The confident predictions that it will continue indefinitely are probably also wrong
// TODO: write better signature

Quanta

Especially with the US Gov stopping the latest Anthropic LLM Claude model.

Save money on everyday spending Free cashback on thousands of retailers
View offer