---
title: "The Frontier AI Model Race (2026)"
slug: frontier-ai-race
type: data-index
sector: ai-models
canonical_url: https://deepstoryresearch.com/data/frontier-ai-race
series: [capability_cost, releases, training_compute, safety_matrix]
series_count: 4
data_point_count: 21
data_as_of: 2026-06
source_quality: curated_snapshot
tier: free
generated_at: 2026-08-11
---

# The Frontier AI Model Race (2026)
_Model releases, capability-per-dollar, training compute & safety posture_

> Deepstory Research context file · **Free tier** · <https://deepstoryresearch.com/data/frontier-ai-race>
> Self-contained AI briefing. Drop into an LLM window or RAG pipeline.

## AI research use contract

- Treat this file as source-grounded context, not a live database. Quote `source_ref` and `data_as_of` when using numbers.
- Separate `FACT`, `ESTIMATE`, `FORECAST`, `TARGET`, and `INFERENCE`. If a row basis is unclear, say it is unclear.
- Prefer T1/T2 sources for hard facts; use T3/T4 sources as context that should be checked before money, legal, medical, operational, or policy decisions.
- Keep answer structure clean: first say what the data says, then what you infer, then what would change the conclusion.

## Research upgrade checklist

- Separate first-party benchmark claims from independent evaluations and live user-preference benchmarks.
- Add inference price, latency, context length, and tool-use reliability so capability is not reduced to benchmark score.
- Track open-weight availability, licensing, and deployment constraints by model family.
- Add a prompt that asks the AI to identify benchmark-gaming and eval-contamination risks.

## TL;DR — index summary

The frontier is a four-lab race (Anthropic, OpenAI, Google, Meta) measured four ways: release cadence by quarter, capability-per-dollar (a top model scores ~62 on the AA Intelligence Index at ~$25/M output tokens), training compute rising off a ~3.1e23-FLOP GPT-3 baseline (2020), and a safety-posture matrix. The pattern: capability keeps climbing while the price of a given capability level keeps falling — the two moving in opposite directions is the whole story.

## What this index covers

Four series on the 2026 frontier-model race: quarterly releases by lab, a capability-vs-cost scatter of leading models, an estimated training-compute (FLOP) trend, and an editorial safety-posture matrix across risk domains.

**Series in this index:**

- `capability_cost` — Benchmark score vs output-token price for leading models.
- `releases` — Count of frontier releases per quarter by lab.
- `training_compute` — Estimated training FLOP over time.
- `safety_matrix` — Per-model risk-domain posture (blocked / gated / allowed).

**Entities tracked:** Anthropic (Claude), OpenAI (GPT), Google DeepMind (Gemini), Meta (Llama); benchmarks: Artificial Analysis, LMArena, Epoch AI, Stanford HAI.

## Key findings

- Four labs trade the lead quarter to quarter — releases are clustered, not steady (anthropic-news / openai-research / google-deepmind / meta-ai).
- A leading model tops the Artificial Analysis Intelligence Index at ~62 while costing ~$25/M output tokens — capability up, price-per-capability down (artificial-analysis).
- Training compute scales off a ~3.1e23-FLOP GPT-3 (2020) baseline (epoch-ai, corroborated by Stanford HAI).
- The safety matrix shows hard limits blocking high-risk cyber/bio/chem while gating legal/medical use — an editorial, not measured, scale (anthropic-news).

## Evidence basis map

The free tables above show `source_ref` but not the row-level `value_basis`. This section mirrors the visible rows only, so AI tools can separate observed values from estimates, forecasts, targets, and derived values.

| series | row | source_ref | value_basis |
| --- | --- | --- | --- |
| capability_cost | model=Claude Opus 4.8; benchmark_score=62; cost_per_m_tokens=25 | artificial-analysis | AA Intelligence Index (tops leaderboard, ~62); Opus-class output ~$25/M (Jun 2026) |
| capability_cost | model=GPT-5.5; benchmark_score=60; cost_per_m_tokens=12 | artificial-analysis | AA Intelligence Index ~60; output ~$12/M (Jun 2026) |
| capability_cost | model=Gemini 3.1 Pro; benchmark_score=57; cost_per_m_tokens=12 | artificial-analysis | AA Intelligence Index ~57; output ~$12/M (Jun 2026) |
| capability_cost | model=Claude Sonnet 4.6; benchmark_score=51; cost_per_m_tokens=15 | artificial-analysis | AA Intelligence Index ~51; output ~$15/M (Jun 2026) |
| releases | quarter=2024 Q2; anthropic=1; openai=1; google=0; meta=0 | anthropic-news | Claude 3.5 Sonnet (Jun 2024); GPT-4o (May 2024) — Anthropic + OpenAI release notes |
| releases | quarter=2024 Q3; anthropic=0; openai=1; google=0; meta=1 | openai-research | OpenAI o1-preview (Sep 2024); Meta Llama 3.1 405B (Jul 2024) |
| releases | quarter=2024 Q4; anthropic=0; openai=1; google=1; meta=0 | google-deepmind | Google Gemini 2.0 (Dec 2024); OpenAI o1/o3 (Dec 2024) |
| releases | quarter=2025 Q1; anthropic=1; openai=0; google=0; meta=0 | anthropic-news | Claude 3.7 Sonnet (Feb 2025) |
| releases | quarter=2025 Q2; anthropic=1; openai=0; google=0; meta=1 | meta-ai | Meta Llama 4 (Apr 2025); Anthropic Claude 4 / Opus 4 (May 2025) |
| training_compute | year=2020; flops_estimate=3.1e+23 | epoch-ai | GPT-3 ~3.1e23 FLOP (Epoch AI; corroborated by Stanford HAI AI Index) |
| training_compute | year=2022; flops_estimate=2.5e+24 | epoch-ai | PaLM-class frontier ~2.5e24 FLOP (Epoch AI) |
| training_compute | year=2023; flops_estimate=2e+25 | epoch-ai | GPT-4 ~2e25 FLOP (Epoch AI) |
| training_compute | year=2024; flops_estimate=5e+25 | epoch-ai | Gemini Ultra ~5e25 FLOP (Epoch AI) |
| training_compute | year=2025; flops_estimate=5e+26 | epoch-ai | Grok-4 ~5e26 FLOP, largest known run (Epoch AI) |
| safety_matrix | model=Claude Fable 5; cyber=1; bio=1; chem=1; legal=3; medical=2 | anthropic-news | Topics_Content/07_New_Trending_Topics_2026.md §1.2 — hard limits block high-risk cyber/bio/chem (fallback to Opus 4.8); editorial 1=blocked/2=gated/3=allowed |
| safety_matrix | model=Claude Opus 4.8; cyber=2; bio=2; chem=2; legal=3; medical=2 | anthropic-news | ASL-3-style gating on CBRN/cyber per model card; editorial encoding |
| safety_matrix | model=Frontier baseline (typical); cyber=2; bio=2; chem=2; legal=3; medical=2 | stanford-hai-2024 | Typical frontier-lab usage-policy posture; editorial encoding |

## The series (per-chart briefings)

### Capability vs cost — scatter
**Direct answer:** Benchmark score vs output-token price for leading models.
**Time bracket:** snapshot (2026)

| model | benchmark_score | cost_per_m_tokens | source_ref |
| --- | --- | --- | --- |
| Claude Opus 4.8 | 62 | 25 | artificial-analysis |
| GPT-5.5 | 60 | 12 | artificial-analysis |
| Gemini 3.1 Pro | 57 | 12 | artificial-analysis |
| Claude Sonnet 4.6 | 51 | 15 | artificial-analysis |

**Read:** The frontier moves down-and-right: same capability gets cheaper each generation.

### Releases by quarter — stacked bar
**Direct answer:** Count of frontier releases per quarter by lab.
**Time bracket:** quarterly

| quarter | anthropic | openai | google | meta | source_ref |
| --- | --- | --- | --- | --- | --- |
| 2024 Q2 | 1 | 1 | 0 | 0 | anthropic-news |
| 2024 Q3 | 0 | 1 | 0 | 1 | openai-research |
| 2024 Q4 | 0 | 1 | 1 | 0 | google-deepmind |
| 2025 Q1 | 1 | 0 | 0 | 0 | anthropic-news |
| 2025 Q2 | 1 | 0 | 0 | 1 | meta-ai |

_+4 more rows on the live page and in the full working dataset._

**Read:** Cadence is bursty; leadership rotates rather than compounding at one lab.

### Training compute — line (log)
**Direct answer:** Estimated training FLOP over time.
**Time bracket:** annual

| year | flops_estimate | source_ref |
| --- | --- | --- |
| 2020 | 3.1e+23 | epoch-ai |
| 2022 | 2.5e+24 | epoch-ai |
| 2023 | 2e+25 | epoch-ai |
| 2024 | 5e+25 | epoch-ai |
| 2025 | 5e+26 | epoch-ai |

**Read:** A roughly exponential climb off the GPT-3 baseline.
**Caveat:** FLOP figures are third-party estimates (Epoch AI), not disclosed by labs.

### Safety posture matrix — heatmap
**Direct answer:** Per-model risk-domain posture (blocked / gated / allowed).
**Time bracket:** snapshot

| model | cyber | bio | chem | legal | medical | source_ref |
| --- | --- | --- | --- | --- | --- | --- |
| Claude Fable 5 | 1 | 1 | 1 | 3 | 2 | anthropic-news |
| Claude Opus 4.8 | 2 | 2 | 2 | 3 | 2 | anthropic-news |
| Frontier baseline (typical) | 2 | 2 | 2 | 3 | 2 | stanford-hai-2024 |

**Caveat:** Editorial 1=blocked / 2=gated / 3=allowed scale — a posture summary, not a measured benchmark.

## Cross-series synthesis — why this matters

The capability_cost and training_compute series pull in opposite directions on price: compute per model rises, yet the cost of reaching a given capability level falls, because architecture and inference efficiency improve faster than raw scale inflates cost. That scissors is why "the frontier" keeps commoditising last year's frontier.

Release cadence explains why no single lab compounds a durable lead: the four-way race means whoever ships next resets the leaderboard, so the interesting variable is capability-per-dollar over time, not any one quarter's winner.

## Entities & relationships

| Entity | What they do | Key stat | Relationship |
| --- | --- | --- | --- |
| Anthropic | Claude model family | Tops AA Index (~62) in the snapshot | Competes with OpenAI/Google/Meta; safety-forward posture |
| OpenAI | GPT model family | GPT-5.5-class in the capability set | Head-to-head with Anthropic on frontier |
| Google DeepMind | Gemini model family | Trails on some coding/reasoning marks | Third pole; strong on multimodal/context |
| Meta | Llama (open-weight) | Open-weight release track | Sets the open-weight floor the others price against |
| Artificial Analysis / Epoch AI | Independent benchmarkers | AA Index; FLOP estimates | Provide the neutral capability + compute measures |

## Timeline of key events & decisions

- **2020** — GPT-3 (~3.1e23 FLOP) → the compute baseline the whole trend is indexed to
- **2024 Q2** — Claude 3.5 Sonnet + GPT-4o → the modern price-performance step that starts the cadence series
- **2026** — Opus-class model tops AA Index at ~$25/M → the capability-per-dollar frontier in the snapshot

## Cross-industry ripple

- **Cloud / semiconductors:** Rising training compute drives GPU/accelerator demand and data-center buildout (see the AI-power-demand and AI-capex indexes).
- **Software / SaaS:** Falling capability-per-dollar collapses the cost of embedding frontier models, compressing margins for wrappers. _(inference)_
- **Labor:** Cheaper frontier capability widens task automation (see the AI-jobs-exposure index). _(inference)_

## Non-obvious reads (interpretation)

_This section is interpretation, not sourced fact — each item names the observation it is built on and a confidence level._

- **Observation:** Compute per model rises while cost-per-capability falls.
  **Read:** The moat is shifting from "who has the most compute" toward inference efficiency and distribution — a lab could lead on benchmarks and still lose on unit economics. _(confidence: medium)_
- **Observation:** Leadership rotates every few quarters.
  **Read:** Durable advantage is more likely to come from product/distribution and safety trust than from a transient benchmark lead. _(confidence: medium)_

## Glossary / key terms

- **FLOP** — Floating-point operations — the unit of training compute.
- **AA Intelligence Index** — Artificial Analysis's composite capability score across benchmarks.
- **Capability-per-dollar** — Benchmark score relative to token price — the frontier's real cost curve.
- **Open-weight** — A model whose parameters are downloadable (e.g. Llama), setting a price floor.
- **$/M tokens** — Price per million output tokens — the inference-cost yardstick.

## How to use this with AI

Paste this file into an LLM context window (or a RAG store) and ask cross-series questions. The tables carry a `source_ref` per row; the source registry below maps each ref to a named source and a trust tier.

### Suggested prompts (multi-series)

```text
Using capability_cost and training_compute, argue whether raw compute or inference efficiency is the bigger 2026 moat, citing the series.
```

```text
From the releases series, characterise the competitive dynamic among the four labs and what would break the rotation.
```

```text
Summarise the safety_matrix posture and flag where it is editorial rather than measured.
```

## Sources & trust

| ref | source | trust tier | url |
| --- | --- | --- | --- |
| anthropic-news | Anthropic Model Cards & Release Notes | T2 · Primary filing / official body | https://www.anthropic.com/news |
| openai-research | OpenAI Research Blog | T2 · Primary filing / official body | https://openai.com/research/ |
| google-deepmind | Google DeepMind Research | T2 · Primary filing / official body | https://deepmind.google/research/ |
| meta-ai | Meta AI — Llama release notes | T2 · Primary filing / official body | https://ai.meta.com/blog/ |
| lmarena | LMArena Leaderboard | T3 · Reputable press / research org / OWID | https://lmarena.ai/ |
| artificial-analysis | Artificial Analysis — model intelligence & pricing | T3 · Reputable press / research org / OWID | https://artificialanalysis.ai/ |
| epoch-ai | Epoch AI — compute & training trends | T3 · Reputable press / research org / OWID | https://epoch.ai/ |
| stanford-hai-2024 | Stanford HAI — AI Index Report | T3 · Reputable press / research org / OWID | https://aiindex.stanford.edu/report/ |

## Caveats & what this index cannot answer

- Exact, lab-disclosed training compute (FLOP figures are third-party estimates).
- Real-world safety outcomes (the matrix is a posture summary on an editorial scale).
- Market share or revenue by lab (not in this index).
- Estimates and forward targets are labelled in the working dataset; never read a labelled estimate or forecast as a settled figure.
- This is an observational data index, not investment, legal, or medical advice.

## Data freshness & methodology

- **Last updated:** 2026-08-11
- **Data as of:** 2026-06
- **What changed most recently:** 2026 snapshot: an Opus-class model tops the Artificial Analysis Intelligence Index (~62) at ~$25/M output tokens.

**Methodology (as seeded):**

> Source-backed values are seeded for four of the five charts: the release
> cadence by lab (2024 → 2026, from each lab’s release notes), the
> capability-vs-cost scatter (Artificial Analysis Intelligence Index vs
> output $/M tokens, June 2026 snapshot), training-compute growth by year
> (Epoch AI, corroborated by Stanford HAI), and a safety-gated capabilities
> matrix.
> 
> Every numeric point carries a sources[].ref and a value_basis. Sources are
> each lab’s own model cards/release notes, Artificial Analysis, Epoch AI,
> and Stanford HAI — cross-checked against public release timelines.
> 
> EDITORIAL ENCODING: the safety-gated matrix scores each model × domain as
> 3 = allowed / 2 = gated / 1 = blocked. This is an interpretation of each
> model’s published safety policy, not a measured benchmark.
> 
> PLACEHOLDER: the per-lab benchmark-trajectory chart is left unseeded — a
> consistent historical per-quarter, per-lab benchmark series was not
> sourceable without mixing incompatible benchmarks. Re-verified 2026-06-15.

---

_Free tier. The full row-level dataset and per-source detail live on the live page and in the working file. Canonical: <https://deepstoryresearch.com/data/frontier-ai-race>. Deepstory Research · https://deepstoryresearch.com_
