New report: Australia tops 121 economies for AI use, new index shows.
Read The Trust Dividend

Responsible AI Australia · Interactive Tracker

The State of the Models

Every frontier AI model that matters, on one page. Who builds them, how they score, and whether you can ever hold the weights yourself.

0.0%

Best open-source SWE-bench score

DeepSeek's MIT-licensed V4-Pro, which anyone can download and use commercially, now fixes real GitHub issues within 0.6 points of Claude Opus 5 on the same independent board. The gap between open and closed AI has never been thinner.

Figures last verified 18 September 2026 · Maintained by Responsible AI Australia · Independent evaluations preferred over vendor claims

Leaderboard

Pick a benchmark. See who leads, and who you could run yourself.

Bars are coloured by access type: violet models are rented through an API, amber models publish their weights under a restricted licence, and green models are genuinely open source. Filter to see how close the downloadable models have come.

198 multiple-choice questions written by PhD scientists in biology, physics and chemistry, designed so a web search alone cannot answer them. A rough proxy for expert-level scientific reasoning.

1
GPT-6 AstraOpenAI
96.3%
2
Gemini 3.8 FlashGoogle DeepMind
95.3%
3
Gemini 3.1 ProGoogle DeepMind
94.1%
4
Claude Opus 5Anthropic
94%
5
GPT-5.5OpenAI
94%
6
Grok 4.6xAI
94%
7
Kimi K3Moonshot AI
93.5%
8
GPT-5.6 SolOpenAI
93%
9
Grok 4.5xAI
93%
10
Qwen3.8 MaxAlibaba
92.6%
11
GLM-5.2Zhipu AI
91.2%
12
MiniMax M3MiniMax
91%
13
DeepSeek V4.1-FlashDeepSeek
90.9%
14
DeepSeek V4-ProDeepSeek
90.1%
15
Muse GlimmerMeta
83.5%

Not shown (no published score on this benchmark): Claude Fable 5.1, Claude Fable 5, Muse Spark 1.3, GLM-5.3, Qwen3-235B. A missing score means the developer has not published a credible figure, not that the model scored zero.

Open vs private

Three ways to ship a frontier model.

The label on a model matters as much as its score. It decides whether you control your own AI supply chain, what happens to your data, and what a vendor can change without asking you.

Private

Weights stay on the developer's servers. You rent access through an API or app, and the developer can change, restrict or retire the model at any time.

You get
The strongest raw capability today, plus safety filtering, uptime and support handled for you.
You give up
No control. Pricing, behaviour and availability can change under you, and your data leaves your infrastructure.

Tracked here: Claude Fable 5.1, Claude Opus 5, Claude Fable 5

Open weight

The trained weights are published and can be downloaded and run on your own hardware, but the licence carries restrictions, so it does not meet the open-source definition.

You get
Run it on your own hardware, fine-tune it on your own data, and keep sensitive information in-house.
You give up
Licence terms still bind you, and some uses (or jurisdictions) may be excluded. Read the licence, not the headline.

Tracked here: Kimi K3, Qwen3.8 Max, GLM-5.3

Open source

Weights are released under an OSI-approved licence such as MIT or Apache 2.0. Anyone can use, modify and redistribute them, including commercially, with almost no strings attached.

You get
Full freedom to use, modify, redistribute and commercialise. No vendor can take the model away.
You give up
You carry the operational and safety burden yourself, and the released weights still sit a step behind the private frontier.

Tracked here: DeepSeek V4-Pro, DeepSeek V4.1-Flash, GLM-5.2

One honest caveat: even MIT-licensed models release the trained weights, not the training data or full training code. Fully open models, where all three are published, remain rare research projects. For most organisations the licence on the weights is what matters in practice, so that is what this page classifies.

Compare models

The full table, sortable.

Click any column to sort, and any row for the story behind its numbers, including which figures are independently verified and which are the developer's own.

Licence
GPT-6 AstraOpenAIPrivateProprietarySep 202696.3$14.44
Gemini 3.8 FlashGoogle DeepMindPrivateProprietarySep 202695.399$1.08
Gemini 3.1 ProGoogle DeepMindPrivateProprietaryMay 202694.17696$3.11
Claude Opus 5AnthropicPrivateProprietaryJul 2026949799$7.22
GPT-5.5OpenAIPrivateProprietaryApr 20269480.6100$7.78
Grok 4.6xAIPrivateProprietaryAug 20269499$2.44
Kimi K3Moonshot AIOpen weightKimi K3 LicenceJul 202693.593.4$4.33
GPT-5.6 SolOpenAIPrivateProprietaryJul 20269396.2$5.78
Grok 4.5xAIPrivateProprietaryJun 20269386.698$2.44
Qwen3.8 MaxAlibabaOpen weightQwen3.8 Max LicenceJul 202692.6
GLM-5.2Zhipu AIOpen sourceMITJun 202691.278.799.2$1.73
MiniMax M3MiniMaxOpen weightMiniMax Community LicenceJun 20269180.5
DeepSeek V4.1-FlashDeepSeekOpen sourceMITSep 202690.9$0.40
DeepSeek V4-ProDeepSeekOpen sourceMITApr 202690.196.4$1.61
Muse GlimmerMetaOpen sourceApache 2.0Aug 202683.57694.7
Claude Fable 5.1AnthropicPrivateProprietarySep 2026100$14.44
Claude Fable 5AnthropicPrivateProprietaryJun 20269599.7$14.44
Muse Spark 1.3MetaPrivateProprietarySep 202699$1.58
GLM-5.3Zhipu AIOpen weightGLM-5.3 LicenceAug 2026$1.73
Qwen3-235BAlibabaOpen sourceApache 2.02025

Tap a row for context on its figures. Price is a blended USD rate per 1M API tokens where a like-for-like figure exists. A dash means no published, credible figure.

The open surge

The story of 2026 is how little daylight is left.

Two years ago, downloadable models trailed the private frontier by a wide, comfortable margin. That margin is now measured in tenths of a point. DeepSeek's MIT-licensed V4-Pro, in its August 2026 build, resolves 96.4 percent of SWE-bench Verified issues on Vals AI's independent board, 0.6 behind Claude Opus 5, at less than a quarter of the price. Moonshot released Kimi K3's full 2.8 trillion-parameter weights in July with a GPQA Diamond score within three points of the best private model, and Zhipu's GLM-5.2 essentially solves competition maths.

The traffic runs both ways. In April 2026 Meta, the loudest advocate for open AI, made its new flagship, Muse Spark, the company's first closed model, then returned to open source in August with the much smaller, Apache-licensed Muse Glimmer. Alibaba released the weights of Qwen3.8 Max in August, but under a custom licence, and Zhipu moved GLM-5.3 from MIT to a licence of its own. Openness is a strategy, not an identity, and it can be reversed in either direction.

For Australian organisations the practical takeaway is choice. Serious capability no longer requires a proprietary API, but governance does not come bundled with a licence file. A downloadable model shifts the responsibility for safe deployment from the vendor to you.

0.0 pts

SWE-bench gap between the best downloadable model and the best private one

0 of 20

tracked models have downloadable weights

$0.00

per 1M tokens for MIT-licensed DeepSeek V4-Pro, versus $7.22 for Claude Opus 5, the private SWE-bench leader

Methodology

How this tracker stays honest, and current.

Figures are reviewed on every major model release and re-verified at least monthly. This page was last verified on 18 September 2026.

  • Independent evaluations (Epoch AI, Vals AI and Artificial Analysis, surfaced through LM Council and BenchLM) are preferred over developer-reported figures. Where only the developer's number exists, the model's row note says so. Epoch AI publishes whole-number scores, so those are shown without a decimal rather than given false precision.
  • A blank cell means no credible published figure exists. We never estimate, interpolate or average incompatible benchmark variants.
  • Benchmark scores are narrow measures of capability, not of safety, reliability or fitness for your use case. A model that tops a leaderboard can still be the wrong choice for a regulated Australian deployment.
  • Model and provider logos are the trademarks of their respective owners, reproduced for identification only via the MIT-licensed lobehub icon set.