· AI  · 8 min read

The Quality Gap in Open AI Models, Honestly Quantified

In June an open model cracked 50 points on the leading independent AI benchmark for the first time; six weeks later a second one came within three points of the frontier. Why the average number is still almost irrelevant for your own decision.

In June an open model cracked 50 points on the leading independent AI benchmark for the first time; six weeks later a second one came within three points of the frontier. Why the average number is still almost irrelevant for your own decision.

The quality gap between open and proprietary AI models has narrowed to a few points on benchmark averages, but it is not closing evenly.

The best open model (Kimi K3) scores 60 on the Artificial Analysis Intelligence Index, while the leading models from Anthropic and OpenAI sit at 61 to 63. On routine tasks the difference is barely measurable anymore; on complex, multi-step reasoning it stays larger. Which number matters depends on your own task mix, not the average.

On June 16, 2026, the company Z.ai released an open language model called GLM-5.2. It was the first open model to score above 50 on the Intelligence Index from Artificial Analysis, an independent ranking of large language models. Six weeks later, Moonshot AI released the weights of Kimi K3: initially measured at 57 by Artificial Analysis, now at 60 on the current index. The leading models from Anthropic and OpenAI sit at 61 to 63 points.

Three points of distance. That sounds like a clear answer, but it is the wrong question. The right question is not how large the gap is on average across all tasks, but how large it is for the tasks that actually come up in your own company.

Center and Edge of Distribution

Knowledge work roughly splits into two groups. One has a familiar pattern and an easily verifiable result: summaries, standard emails, translation, routine extraction, ordinary code. That is the Center of Distribution, because tasks exactly like these show up millions of times in the training data of every major model. The other group sits at the Edge: novel problems, multi-step reasoning, work with high error costs. Less comparable training material exists there, and that is where the quality difference between models shows up most clearly.

Most everyday knowledge work is, by definition, Center of Distribution. But most companies have never measured how their own task mix actually splits between center and edge. Without that measurement, model selection stays a guess.

Two stacked charts share one x-axis, which runs across the task space from the middle of the distribution out to both edges. The upper chart shows two quality curves. In the middle they almost coincide, and both fall off toward either edge, but the open-model curve falls faster, so the shaded gap between them widens from barely visible at the center to wide at both edges. Neither curve reaches zero: quality drops at the edge, it does not disappear. The lower chart shows the share of tasks as a bell curve peaking in the middle and running down toward zero at both edges. The center region, marked between one standard deviation either side of the middle, is where most tasks sit, and it is exactly where the two quality curves are closest.

Output quality

  • Proprietary frontier models
  • Open models
  • Quality gap

Distribution of tasks in the training data

Edge of Distribution

Center of Distribution

Edge of Distribution

Schematic view, not measured data. The center region sits around the middle of the distribution, its boundary drawn at one standard deviation. There the gap between proprietary and open models is barely measurable; toward the edges it grows while the share of tasks shrinks.

What the Intelligence Index shows today

How fast this field moves shows in the eight weeks after the GLM-5.2 release. When the model appeared, it was the only open model above 50, at 51 points; the next open models trailed by a clear margin at 44 and 43. In late July, Moonshot AI then released the weights of Kimi K3, which today sits three points behind the frontier. Anyone who based a model decision on the June gap was working with outdated numbers six weeks later.

The average number hides how unevenly the gap spreads across different task types:

  • Coding benchmarks: the gap has narrowed to 2 to 3 percentage points (as of Q1 2026, analyses by Artificial Analysis and OpenRouter).
  • Reasoning-heavy benchmarks: proprietary models most recently still led here by 3 to 8 percentage points (as of Q2 2026, same sources).
  • Price: Artificial Analysis puts the price advantage of open models at comparable intelligence at a factor of 2 to 6 (as of April 2026); at EU providers, open models start at around 0.17 US dollars per million tokens (IONOS).

Translated: on routine tasks with a clear pattern, the quality loss from an open model is often barely noticeable. On demanding, multi-step reasoning, it stays real.

Why the gap is not closing evenly

Epoch AI measures the gap independently through its own Epoch Capabilities Index, and the picture is sober: since January 2026, the most capable open models have lagged the proprietary frontier by an average of four months, a gap of roughly eight index points. In October 2025, the same measurement had found an average lag of three months.

So the gap is not a line running smoothly toward zero. Depending on the ruler, it even moves in opposite directions: Epoch’s lag grew from three to four months, while on the Intelligence Index a single release like Kimi K3 pushed the distance down to a few points in late July. Every new frontier release resets the bar, and open models need time to catch up. Anyone building a model decision on “the gap keeps closing steadily” is building on a line that does not exist in the data.

What this means for your own model choice

A recommendation that always favors switching to open models would be dishonest. Interactive, demanding work with high error costs, such as software development or strategic analysis, often sits at the edge of the distribution. There, the gap to the most expensive models stays relevant, and a switch would measurably cost quality.

The other half of the truth: a large share of daily tasks in most companies sits in the center. Standard correspondence, meeting notes, translations, recurring reports. For that share, the remaining quality difference on benchmark level is now small, while the cost difference stays large.

Whether a switch is worth it, then, does not depend on the model. It depends on how large the center share of your own tasks actually is. How to measure that share without guessing is the topic of the next post.

Frequently asked questions

How large is the quality gap between open and proprietary AI models? As of August 2026, the best open model (Kimi K3) scores 60 on the Artificial Analysis Intelligence Index, while the leading models from Anthropic and OpenAI sit at 61 to 63. On reasoning-heavy benchmarks the gap was most recently 3 to 8 percentage points (Q2 2026); on coding benchmarks it has narrowed to 2 to 3 percentage points (Q1 2026).

What does Center of Distribution mean for AI tasks? Center of Distribution tasks have a familiar pattern and an easily verifiable result, such as summaries, standard emails, or routine code. Edge of Distribution tasks are novel problems, multi-step reasoning, and work with high error costs. Open models are now nearly equivalent in the center; at the edge they still lag.

Should you switch to an open model because the gap is smaller? That depends on your own task mix, not on the benchmark average. If most of your work is Center of Distribution, switching costs little quality and saves significantly on cost. If you do a lot of interactive, demanding work, such as software development or strategy, you are often still better off with a frontier model.

Sources


Related: Residency Is Not Jurisdiction, Paid Per Seat, Used Per Task, and What Happens When Your AI Vendor Disappears.

Back to Blog

Related Posts

View All Posts »
What Happens When Your AI Vendor Disappears

What Happens When Your AI Vendor Disappears

An AI vendor can disappear no matter how big it was. Three questions decide whether that means a migration for your company or a rebuild from zero, and they belong before the contract, not after.

Paid Per Seat, Used Per Task

Paid Per Seat, Used Per Task

300 paid Copilot seats are not 300 users. What Microsoft's own reports count as "active," what the widely cited 20-30 percent figure actually measures, and why it likely understates the gap rather than overstates it.

Residency Is Not Jurisdiction

Residency Is Not Jurisdiction

Data "stored in the EU" sounds like protection from US access. It is not, automatically. What the CLOUD Act and FISA 702 mean, and when it matters for your company.