Back to recaps

Four frontier launches in eight days: six labs now field a model above 50 on the Artificial Analysis Intelligence Index

Artificial Analysis chart artwork for four frontier model launches in eight days
Image: Artificial Analysis

In an Artificial Analysis article, the benchmarking firm examines four frontier-model launches in eight days and the expansion from two to six labs with models near the top of its Intelligence Index. It compares capability scores, prices, and Kimi K3's position in the field.

Four launches compress the frontier

Section 01

SpaceXAI's Grok 4.5 scored 54 on July 8. OpenAI's GPT-5.6 Sol, Terra, and Luna followed with 59, 55, and 51. Meta's Muse Spark 1.1 reached 51. Moonshot AI's Kimi K3 arrived July 16 at 57.

The top three were Claude Fable 5 at 60, GPT-5.6 Sol at 59, and Kimi K3 at 57. Three labs held those positions. Four of the ten highest-scoring models had arrived since July 8; six had arrived since early June.

Claude Fable 5 kept first place, but its lead fell from four points to one. OpenAI, Meta, and Moonshot AI confirm their releases. Artificial Analysis supplied the cross-model scores and order. They are a July 17 snapshot of its composite Index.

The frontier expands from two labs to six

Section 02

In early June, only Anthropic and OpenAI had a model at 51 or higher. GLM-5.2 added Z AI in mid-June. Grok 4.5 added SpaceXAI, Muse Spark 1.1 added Meta, and Kimi K3 added Moonshot AI. The count reached six labs in six weeks.

Index v4.1 combines nine evaluations. Agent tasks carry 34% of the weight, coding and scientific reasoning carry 24% each, and general capabilities carry 18%. Artificial Analysis estimates uncertainty below plus or minus one point. A displayed 51 marks its chosen cutoff; it does not show a decisive gap from a model displayed at 50.

Kimi K3 enters near the top

Section 03

Kimi K3 scored 57 at launch. Its GDPval-AA v2 result was 1668 Elo, behind Claude Fable 5 at 1760 and GPT-5.6 Sol at 1748. On AA-Briefcase, it reached 1547 Elo. Its Analytical Quality Elo of 1760 was close to Fable 5 at 1764.

GDPval-AA v2 uses 220 professional tasks across 44 occupations and nine industries. Blind pairwise comparisons set an Elo rating against a human baseline of 1,000. AA-Briefcase uses 91 tasks across four private, multi-week project scenarios. It combines rubric results, analysis quality, and presentation.

Artificial Analysis measured Kimi K3 at $0.94 per Index task, against $1.80 for Claude Opus 4.8 at a nearby displayed result. The live leaderboard had already changed Kimi K3's ordinal by July 20. The launch figures remain valid as dated results, not a permanent rank.

Near-frontier intelligence becomes cheaper

Section 04

GPT-5.6 Sol scored one point below Claude Fable 5 and cost $1.04 per Index task, against $2.75 for the leader. Grok 4.5 scored 54 at $0.31, under a third of GPT-5.5's $0.99. At 51, GPT-5.6 Luna cost $0.21 and Muse Spark 1.1 cost $0.26. Both came in below GLM-5.2's earlier $0.32.

Artificial Analysis calculates cost per task from provider prices, token use, reasoning, cache reads, and cache writes across the weighted Index workload. OpenAI prices Sol at $5/$30 per million input/output tokens and Luna at $1/$6. SpaceXAI prices Grok 4.5 at $2/$6.

The two-to-three-times reduction applies to Index v4.1 at the tested settings. Different workloads can use different token volumes and produce a different cost gap.

The charts show the wider shift

Section 05

The first chart places six labs above 50. Its arrows mark six new model entries in eight days, including three GPT-5.6 tiers from one lab. The historical chart shows one or two labs holding the frontier for most of the period since late 2022, followed by four more labs moving within single digits of first place.

The third chart covers coding agents. GPT-5.6 Sol in Codex led at 80. GPT-5.6 Terra in Codex and Claude Fable 5 in Claude Code both sat near 77. Grok 4.5 in Grok Build reached 76, and Muse Spark 1.1 in OpenCode reached 69.

The Coding Agent Index combines DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Each row is a model-plus-agent setup. Tooling and orchestration are part of the result.

The last chart returns to cost. GPT-5.6 Luna, Muse Spark 1.1, and Grok 4.5 sat at or below the previous week's cheapest model at the same displayed level. Kimi K3 and GPT-5.6 Sol landed within three and one points of Claude Fable 5 at roughly a third of its measured cost per task.