Responsible AI Australia · Interactive Tracker
The State of the Models
Every frontier AI model that matters, on one page. Who builds them, how they score, and whether you can ever hold the weights yourself.
0.0%
Best open-source SWE-bench score
DeepSeek's MIT-licensed V4-Pro, which anyone can download and use commercially, now fixes real GitHub issues within 0.6 points of Claude Opus 5 on the same independent board. The gap between open and closed AI has never been thinner.
Figures last verified 18 September 2026 · Maintained by Responsible AI Australia · Independent evaluations preferred over vendor claims
Leaderboard
Pick a benchmark. See who leads, and who you could run yourself.
Bars are coloured by access type: violet models are rented through an API, amber models publish their weights under a restricted licence, and green models are genuinely open source. Filter to see how close the downloadable models have come.
198 multiple-choice questions written by PhD scientists in biology, physics and chemistry, designed so a web search alone cannot answer them. A rough proxy for expert-level scientific reasoning.
Not shown (no published score on this benchmark): Claude Fable 5.1, Claude Fable 5, Muse Spark 1.3, GLM-5.3, Qwen3-235B. A missing score means the developer has not published a credible figure, not that the model scored zero.
Open vs private
Three ways to ship a frontier model.
The label on a model matters as much as its score. It decides whether you control your own AI supply chain, what happens to your data, and what a vendor can change without asking you.
Private
Weights stay on the developer's servers. You rent access through an API or app, and the developer can change, restrict or retire the model at any time.
- You get
- The strongest raw capability today, plus safety filtering, uptime and support handled for you.
- You give up
- No control. Pricing, behaviour and availability can change under you, and your data leaves your infrastructure.
Tracked here: Claude Fable 5.1, Claude Opus 5, Claude Fable 5
Open weight
The trained weights are published and can be downloaded and run on your own hardware, but the licence carries restrictions, so it does not meet the open-source definition.
- You get
- Run it on your own hardware, fine-tune it on your own data, and keep sensitive information in-house.
- You give up
- Licence terms still bind you, and some uses (or jurisdictions) may be excluded. Read the licence, not the headline.
Tracked here: Kimi K3, Qwen3.8 Max, GLM-5.3
Open source
Weights are released under an OSI-approved licence such as MIT or Apache 2.0. Anyone can use, modify and redistribute them, including commercially, with almost no strings attached.
- You get
- Full freedom to use, modify, redistribute and commercialise. No vendor can take the model away.
- You give up
- You carry the operational and safety burden yourself, and the released weights still sit a step behind the private frontier.
Tracked here: DeepSeek V4-Pro, DeepSeek V4.1-Flash, GLM-5.2
One honest caveat: even MIT-licensed models release the trained weights, not the training data or full training code. Fully open models, where all three are published, remain rare research projects. For most organisations the licence on the weights is what matters in practice, so that is what this page classifies.
Compare models
The full table, sortable.
Click any column to sort, and any row for the story behind its numbers, including which figures are independently verified and which are the developer's own.
| Licence | |||||||
|---|---|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | Private | Proprietary | Sep 2026 | 96.3 | – | – | $14.44 |
| Gemini 3.8 FlashGoogle DeepMind | Private | Proprietary | Sep 2026 | 95.3 | – | 99 | $1.08 |
| Gemini 3.1 ProGoogle DeepMind | Private | Proprietary | May 2026 | 94.1 | 76 | 96 | $3.11 |
| Claude Opus 5Anthropic | Private | Proprietary | Jul 2026 | 94 | 97 | 99 | $7.22 |
| GPT-5.5OpenAI | Private | Proprietary | Apr 2026 | 94 | 80.6 | 100 | $7.78 |
| Grok 4.6xAI | Private | Proprietary | Aug 2026 | 94 | – | 99 | $2.44 |
| Kimi K3Moonshot AI | Open weight | Kimi K3 Licence | Jul 2026 | 93.5 | 93.4 | – | $4.33 |
| GPT-5.6 SolOpenAI | Private | Proprietary | Jul 2026 | 93 | 96.2 | – | $5.78 |
| Grok 4.5xAI | Private | Proprietary | Jun 2026 | 93 | 86.6 | 98 | $2.44 |
| Qwen3.8 MaxAlibaba | Open weight | Qwen3.8 Max Licence | Jul 2026 | 92.6 | – | – | – |
| GLM-5.2Zhipu AI | Open source | MIT | Jun 2026 | 91.2 | 78.7 | 99.2 | $1.73 |
| MiniMax M3MiniMax | Open weight | MiniMax Community Licence | Jun 2026 | 91 | 80.5 | – | – |
| DeepSeek V4.1-FlashDeepSeek | Open source | MIT | Sep 2026 | 90.9 | – | – | $0.40 |
| DeepSeek V4-ProDeepSeek | Open source | MIT | Apr 2026 | 90.1 | 96.4 | – | $1.61 |
| Muse GlimmerMeta | Open source | Apache 2.0 | Aug 2026 | 83.5 | 76 | 94.7 | – |
| Claude Fable 5.1Anthropic | Private | Proprietary | Sep 2026 | – | – | 100 | $14.44 |
| Claude Fable 5Anthropic | Private | Proprietary | Jun 2026 | – | 95 | 99.7 | $14.44 |
| Muse Spark 1.3Meta | Private | Proprietary | Sep 2026 | – | – | 99 | $1.58 |
| GLM-5.3Zhipu AI | Open weight | GLM-5.3 Licence | Aug 2026 | – | – | – | $1.73 |
| Qwen3-235BAlibaba | Open source | Apache 2.0 | 2025 | – | – | – | – |
Tap a row for context on its figures. Price is a blended USD rate per 1M API tokens where a like-for-like figure exists. A dash means no published, credible figure.
The open surge
The story of 2026 is how little daylight is left.
Two years ago, downloadable models trailed the private frontier by a wide, comfortable margin. That margin is now measured in tenths of a point. DeepSeek's MIT-licensed V4-Pro, in its August 2026 build, resolves 96.4 percent of SWE-bench Verified issues on Vals AI's independent board, 0.6 behind Claude Opus 5, at less than a quarter of the price. Moonshot released Kimi K3's full 2.8 trillion-parameter weights in July with a GPQA Diamond score within three points of the best private model, and Zhipu's GLM-5.2 essentially solves competition maths.
The traffic runs both ways. In April 2026 Meta, the loudest advocate for open AI, made its new flagship, Muse Spark, the company's first closed model, then returned to open source in August with the much smaller, Apache-licensed Muse Glimmer. Alibaba released the weights of Qwen3.8 Max in August, but under a custom licence, and Zhipu moved GLM-5.3 from MIT to a licence of its own. Openness is a strategy, not an identity, and it can be reversed in either direction.
For Australian organisations the practical takeaway is choice. Serious capability no longer requires a proprietary API, but governance does not come bundled with a licence file. A downloadable model shifts the responsibility for safe deployment from the vendor to you.
0.0 pts
SWE-bench gap between the best downloadable model and the best private one
0 of 20
tracked models have downloadable weights
$0.00
per 1M tokens for MIT-licensed DeepSeek V4-Pro, versus $7.22 for Claude Opus 5, the private SWE-bench leader
Methodology
How this tracker stays honest, and current.
Figures are reviewed on every major model release and re-verified at least monthly. This page was last verified on 18 September 2026.
- Independent evaluations (Epoch AI, Vals AI and Artificial Analysis, surfaced through LM Council and BenchLM) are preferred over developer-reported figures. Where only the developer's number exists, the model's row note says so. Epoch AI publishes whole-number scores, so those are shown without a decimal rather than given false precision.
- A blank cell means no credible published figure exists. We never estimate, interpolate or average incompatible benchmark variants.
- Benchmark scores are narrow measures of capability, not of safety, reliability or fitness for your use case. A model that tops a leaderboard can still be the wrong choice for a regulated Australian deployment.
- Model and provider logos are the trademarks of their respective owners, reproduced for identification only via the MIT-licensed lobehub icon set.
Sources
- Epoch AI model pages (independent GPQA, SWE-bench and OTIS Mock AIME runs)
- Vals AI, SWE-bench Verified leaderboard (archived, 500 tasks)
- Artificial Analysis, GPQA Diamond evaluations
- LM Council (Epoch AI and Scale AI independent runs)
- LLM Stats leaderboard and release log
- BenchLM composite index
- Hugging Face model cards (licences, parameter counts and developer-reported figures)
- Moonshot AI, Kimi K3 weights release
- MarkTechPost, open trillion-scale MoE comparison
- The Batch, Meta's Muse Spark pivot
