cyberivy
Open-Weight ModelsMozillaKimi K3GLM 5.2AI EconomicsOpen Source AIArtificial AnalysisMETR

Open-weight AI models are now only four months behind

September 15, 2026

Große blaue Buchstaben AI stehen zwischen verschlungenen schwarzen Linien auf einer digitalen Fläche

Mozilla's new report puts the lead of closed frontier models at 4.4 months. For routine tasks, open-weight models can be substantially cheaper.

What this is about

A new report from Mozilla puts the performance gap between leading closed AI models and the best open-weight models at just 4.4 months. Ars Technica previewed the findings on September 15, 2026. The comparison matters economically: the report says Moonshot AI's Kimi K3 scores only three points behind Anthropic's closed Fable 5 model on the Artificial Analysis Intelligence Index while costing about 30 percent as much.

The conclusion is not that open models are always better. It is that the premium for a closed frontier model increasingly pays off only for particular workloads. For recurring standard work, Mozilla argues that organizations should evaluate open models as the default starting point.

What the comparison actually measures

Mozilla combines several perspectives. One is METR's time-horizon measure, which asks how long a human expert would need for a task that a model can complete with a 50 percent success rate. According to the data cited in the report, the best closed model can currently handle tasks about 1.7 times as long as those handled by the best open model.

In simplified terms, if an open model reliably completes a seven-hour task, the closed model reaches about twelve hours. Four months later, the observed trend suggests that the open model reaches the previous twelve-hour mark. A second comparison uses Terminal-Bench 2.1 with a neutral harness. There, Z.ai's open GLM 5.2 finished less than one point behind Claude Opus 4.7 and 4.8 while costing roughly five times less per completed task.

Why it matters

Companies no longer have to choose only between maximum performance and complete in-house development. They can route work by difficulty: cheaper open models for routine tasks and closed frontier models for long, difficult, or time-critical cases. The report cites DoorDash as an example of this split.

There is also a strategic dependency. Eight of the ten models with the highest token volume on OpenRouter in August 2026 offered open weights, but many of the strongest open models come from China. Mozilla therefore calls for more public compute programs, neutral foundations, and European and US alternatives. Open weights do not mean full openness either: training data, data pipelines, and training code often remain undisclosed.

In plain language

Think of renting a car. The expensive, fully equipped vehicle may be worth it for a difficult mountain journey. For the daily trip to the supermarket, a much cheaper car is often sufficient. The question is not which vehicle is best overall, but which one can complete the actual route reliably.

A practical example

A software team processes 10,000 support cases each month. Of these, 9,000 involve known questions, summaries, and simple classification; 1,000 require long context chains or high-risk decisions. If an open model handles routine cases for one-fifth of the cost per completed task, the team can cover most work cheaply and route only hard cases to a closed frontier model.

Before switching, the team would need to test both options on its own data: success rate, latency, infrastructure cost, privacy, and operational monitoring. A model's list price is not the full bill.

Scope and limits

  • Benchmarks only partly reflect real workflows; a small score gap can be important or irrelevant depending on the task.
  • Open weights can reduce license and usage costs but add expenses for hardware, operations, updates, security, and specialists.
  • The 4.4-month figure describes an observed average trend, not a guarantee that every open model will catch every closed model on that schedule.

The comparison also does not establish whether training data is lawful, balanced, or fully documented. Organizations should therefore avoid choosing on benchmark and price alone.

SEO & GEO keywords

Mozilla State of Open Source AI, open-weight models, Kimi K3, GLM 5.2, Claude Opus, Artificial Analysis, METR, Terminal-Bench 2.1, AI costs, open AI, model comparison, European AI

💡 In plain English

Mozilla says closed frontier models are now only about 4.4 months ahead on average. For many routine tasks, open-weight models can therefore be the more economical choice.

Key Takeaways

  • Mozilla puts the gap at 4.4 months.
  • Kimi K3 trails Fable 5 by three points on the cited index while costing about 30 percent as much.
  • A neutral Terminal-Bench comparison places GLM 5.2 less than one point behind Claude Opus 4.7 and 4.8.
  • Open weights do not automatically mean open training data or training code.
  • Routing work by difficulty may be the most economical strategy.

FAQ

What does a 4.4-month lead mean?

It estimates the time gap between capability frontiers in the observed trend. It does not apply to every task or every model.

Are open models always cheaper?

Not necessarily. Operations, hardware, security, and specialist staff can consume part of the price advantage.

Are open-weight models fully open source?

Usually not. Weights may be available while training data, pipelines, and training code remain closed.

Sources & Context