Responsible AI Australia · Interactive Tracker
The State of the Models
Every frontier AI model that matters, on one page. Who builds them, how they score, and whether you can ever hold the weights yourself.
0.0%
Best open-weight GPQA score
Moonshot's Kimi K3, whose weights anyone can download, now answers PhD-level science questions within 0.6 points of the best private model. The gap between open and closed AI has never been thinner.
Figures last verified 4 August 2026 · Maintained by Responsible AI Australia · Independent evaluations preferred over vendor claims
Leaderboard
Pick a benchmark. See who leads, and who you could run yourself.
Bars are coloured by access type: violet models are rented through an API, amber models publish their weights under a restricted licence, and green models are genuinely open source. Filter to see how close the downloadable models have come.
198 multiple-choice questions written by PhD scientists in biology, physics and chemistry, designed so a web search alone cannot answer them. A rough proxy for expert-level scientific reasoning.
Not shown (no published score on this benchmark): Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Grok 4.5, Muse Spark 1.1, Qwen3.8 Max, Llama 5, MiniMax M3, Qwen3-235B. A missing score means the developer has not published a credible figure, not that the model scored zero.
Open vs private
Three ways to ship a frontier model.
The label on a model matters as much as its score. It decides whether you control your own AI supply chain, what happens to your data, and what a vendor can change without asking you.
Private
Weights stay on the developer's servers. You rent access through an API or app, and the developer can change, restrict or retire the model at any time.
- You get
- The strongest raw capability today, plus safety filtering, uptime and support handled for you.
- You give up
- No control. Pricing, behaviour and availability can change under you, and your data leaves your infrastructure.
Tracked here: Claude Opus 5, Claude Fable 5, GPT-5.6 Sol
Open weight
The trained weights are published and can be downloaded and run on your own hardware, but the licence carries restrictions, so it does not meet the open-source definition.
- You get
- Run it on your own hardware, fine-tune it on your own data, and keep sensitive information in-house.
- You give up
- Licence terms still bind you, and some uses (or jurisdictions) may be excluded. Read the licence, not the headline.
Tracked here: Kimi K3, Llama 5
Open source
Weights are released under an OSI-approved licence such as MIT or Apache 2.0. Anyone can use, modify and redistribute them, including commercially, with almost no strings attached.
- You get
- Full freedom to use, modify, redistribute and commercialise. No vendor can take the model away.
- You give up
- You carry the operational and safety burden yourself, and the released weights still sit a step behind the private frontier.
Tracked here: DeepSeek V4-Pro, GLM-5.2, MiniMax M3
One honest caveat: even MIT-licensed models release the trained weights, not the training data or full training code. Fully open models, where all three are published, remain rare research projects. For most organisations the licence on the weights is what matters in practice, so that is what this page classifies.
Compare models
The full table, sortable.
Click any column to sort, and any row for the story behind its numbers, including which figures are independently verified and which are the developer's own.
| Licence | |||||||
|---|---|---|---|---|---|---|---|
| Gemini 3.1 ProGoogle DeepMind | Private | Proprietary | May 2026 | 94.1 | – | – | – |
| GPT-5.5OpenAI | Private | Proprietary | Apr 2026 | 94.0 | 80.6 | 100.0 | $7.78 |
| Kimi K3Moonshot AI | Open weight | Modified MIT | Jul 2026 | 93.5 | – | – | $4.33 |
| GLM-5.2Zhipu AI | Open source | MIT | Jun 2026 | 91.2 | – | 99.2 | $1.18 |
| DeepSeek V4-ProDeepSeek | Open source | MIT | Apr 2026 | 90.1 | 80.6 | – | – |
| Claude Opus 5Anthropic | Private | Proprietary | Jul 2026 | – | 97.0 | – | $7.22 |
| Claude Fable 5Anthropic | Private | Proprietary | Jul 2026 | – | 95.0 | 99.7 | $14.44 |
| GPT-5.6 SolOpenAI | Private | Proprietary | Jul 2026 | – | 96.2 | – | $7.78 |
| Grok 4.5xAI | Private | Proprietary | Jun 2026 | – | – | – | $2.44 |
| Muse Spark 1.1Meta | Private | Proprietary | Jun 2026 | – | – | – | $1.58 |
| Qwen3.8 MaxAlibaba | Private | Proprietary | Jul 2026 | – | – | – | – |
| Llama 5Meta | Open weight | Llama Community Licence | Apr 2026 | – | – | – | – |
| MiniMax M3MiniMax | Open source | Apache 2.0 | May 2026 | – | – | – | – |
| Qwen3-235BAlibaba | Open source | Apache 2.0 | 2025 | – | – | – | – |
Tap a row for context on its figures. Price is a blended USD rate per 1M API tokens where a like-for-like figure exists. A dash means no published, credible figure.
The open surge
The story of 2026 is how little daylight is left.
Two years ago, downloadable models trailed the private frontier by a wide, comfortable margin. That margin is now measured in tenths of a point. Moonshot released Kimi K3's full 2.8 trillion-parameter weights on 27 July 2026, with the best GPQA Diamond score ever published for open weights. DeepSeek's MIT-licensed V4 line matches last year's private flagships on real software-engineering work at a fraction of the price, and Zhipu's GLM-5.2 essentially solves competition maths.
The traffic runs both ways. In April 2026 Meta, the loudest advocate for open AI, split its strategy: Llama 5 stayed open weight while its new flagship, Muse Spark, became the company's first closed model. Alibaba ships its best Qwen capability to the hosted Max tier before the open releases. Openness is a strategy, not an identity, and it can be reversed.
For Australian organisations the practical takeaway is choice. Serious capability no longer requires a proprietary API, but governance does not come bundled with a licence file. A downloadable model shifts the responsibility for safe deployment from the vendor to you.
0.0 pts
GPQA gap between the best open-weight and best private score
0 of 14
tracked models have downloadable weights
$0.00
per 1M tokens for MIT-licensed GLM-5.2, versus $7.22 for the private SWE-bench leader
Methodology
How this tracker stays honest, and current.
Figures are reviewed on every major model release and re-verified at least monthly. This page was last verified on 4 August 2026.
- Independent evaluations (Epoch AI and Scale AI, surfaced through LM Council) are preferred over developer-reported figures. Where only the developer's number exists, the model's row note says so.
- A blank cell means no credible published figure exists. We never estimate, interpolate or average incompatible benchmark variants.
- Benchmark scores are narrow measures of capability, not of safety, reliability or fitness for your use case. A model that tops a leaderboard can still be the wrong choice for a regulated Australian deployment.
- Model and provider logos are the trademarks of their respective owners, reproduced for identification only via the MIT-licensed lobehub icon set.
