New report: Australia tops 121 economies for AI use, new index shows.
Read The Trust Dividend

Responsible AI Australia · Interactive Tracker

The State of the Models

Every frontier AI model that matters, on one page. Who builds them, how they score, and whether you can ever hold the weights yourself.

0.0%

Best open-weight GPQA score

Moonshot's Kimi K3, whose weights anyone can download, now answers PhD-level science questions within 0.6 points of the best private model. The gap between open and closed AI has never been thinner.

Figures last verified 4 August 2026 · Maintained by Responsible AI Australia · Independent evaluations preferred over vendor claims

Leaderboard

Pick a benchmark. See who leads, and who you could run yourself.

Bars are coloured by access type: violet models are rented through an API, amber models publish their weights under a restricted licence, and green models are genuinely open source. Filter to see how close the downloadable models have come.

198 multiple-choice questions written by PhD scientists in biology, physics and chemistry, designed so a web search alone cannot answer them. A rough proxy for expert-level scientific reasoning.

1
Gemini 3.1 ProGoogle DeepMind
94.1%
2
GPT-5.5OpenAI
94.0%
3
Kimi K3Moonshot AI
93.5%
4
GLM-5.2Zhipu AI
91.2%
5
DeepSeek V4-ProDeepSeek
90.1%

Not shown (no published score on this benchmark): Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Grok 4.5, Muse Spark 1.1, Qwen3.8 Max, Llama 5, MiniMax M3, Qwen3-235B. A missing score means the developer has not published a credible figure, not that the model scored zero.

Open vs private

Three ways to ship a frontier model.

The label on a model matters as much as its score. It decides whether you control your own AI supply chain, what happens to your data, and what a vendor can change without asking you.

Private

Weights stay on the developer's servers. You rent access through an API or app, and the developer can change, restrict or retire the model at any time.

You get
The strongest raw capability today, plus safety filtering, uptime and support handled for you.
You give up
No control. Pricing, behaviour and availability can change under you, and your data leaves your infrastructure.

Tracked here: Claude Opus 5, Claude Fable 5, GPT-5.6 Sol

Open weight

The trained weights are published and can be downloaded and run on your own hardware, but the licence carries restrictions, so it does not meet the open-source definition.

You get
Run it on your own hardware, fine-tune it on your own data, and keep sensitive information in-house.
You give up
Licence terms still bind you, and some uses (or jurisdictions) may be excluded. Read the licence, not the headline.

Tracked here: Kimi K3, Llama 5

Open source

Weights are released under an OSI-approved licence such as MIT or Apache 2.0. Anyone can use, modify and redistribute them, including commercially, with almost no strings attached.

You get
Full freedom to use, modify, redistribute and commercialise. No vendor can take the model away.
You give up
You carry the operational and safety burden yourself, and the released weights still sit a step behind the private frontier.

Tracked here: DeepSeek V4-Pro, GLM-5.2, MiniMax M3

One honest caveat: even MIT-licensed models release the trained weights, not the training data or full training code. Fully open models, where all three are published, remain rare research projects. For most organisations the licence on the weights is what matters in practice, so that is what this page classifies.

Compare models

The full table, sortable.

Click any column to sort, and any row for the story behind its numbers, including which figures are independently verified and which are the developer's own.

Licence
Gemini 3.1 ProGoogle DeepMindPrivateProprietaryMay 202694.1
GPT-5.5OpenAIPrivateProprietaryApr 202694.080.6100.0$7.78
Kimi K3Moonshot AIOpen weightModified MITJul 202693.5$4.33
GLM-5.2Zhipu AIOpen sourceMITJun 202691.299.2$1.18
DeepSeek V4-ProDeepSeekOpen sourceMITApr 202690.180.6
Claude Opus 5AnthropicPrivateProprietaryJul 202697.0$7.22
Claude Fable 5AnthropicPrivateProprietaryJul 202695.099.7$14.44
GPT-5.6 SolOpenAIPrivateProprietaryJul 202696.2$7.78
Grok 4.5xAIPrivateProprietaryJun 2026$2.44
Muse Spark 1.1MetaPrivateProprietaryJun 2026$1.58
Qwen3.8 MaxAlibabaPrivateProprietaryJul 2026
Llama 5MetaOpen weightLlama Community LicenceApr 2026
MiniMax M3MiniMaxOpen sourceApache 2.0May 2026
Qwen3-235BAlibabaOpen sourceApache 2.02025

Tap a row for context on its figures. Price is a blended USD rate per 1M API tokens where a like-for-like figure exists. A dash means no published, credible figure.

The open surge

The story of 2026 is how little daylight is left.

Two years ago, downloadable models trailed the private frontier by a wide, comfortable margin. That margin is now measured in tenths of a point. Moonshot released Kimi K3's full 2.8 trillion-parameter weights on 27 July 2026, with the best GPQA Diamond score ever published for open weights. DeepSeek's MIT-licensed V4 line matches last year's private flagships on real software-engineering work at a fraction of the price, and Zhipu's GLM-5.2 essentially solves competition maths.

The traffic runs both ways. In April 2026 Meta, the loudest advocate for open AI, split its strategy: Llama 5 stayed open weight while its new flagship, Muse Spark, became the company's first closed model. Alibaba ships its best Qwen capability to the hosted Max tier before the open releases. Openness is a strategy, not an identity, and it can be reversed.

For Australian organisations the practical takeaway is choice. Serious capability no longer requires a proprietary API, but governance does not come bundled with a licence file. A downloadable model shifts the responsibility for safe deployment from the vendor to you.

0.0 pts

GPQA gap between the best open-weight and best private score

0 of 14

tracked models have downloadable weights

$0.00

per 1M tokens for MIT-licensed GLM-5.2, versus $7.22 for the private SWE-bench leader

Methodology

How this tracker stays honest, and current.

Figures are reviewed on every major model release and re-verified at least monthly. This page was last verified on 4 August 2026.

  • Independent evaluations (Epoch AI and Scale AI, surfaced through LM Council) are preferred over developer-reported figures. Where only the developer's number exists, the model's row note says so.
  • A blank cell means no credible published figure exists. We never estimate, interpolate or average incompatible benchmark variants.
  • Benchmark scores are narrow measures of capability, not of safety, reliability or fitness for your use case. A model that tops a leaderboard can still be the wrong choice for a regulated Australian deployment.
  • Model and provider logos are the trademarks of their respective owners, reproduced for identification only via the MIT-licensed lobehub icon set.