The change log

Field notes.

What changed, when we checked it, and why it matters. A record of this guide’s evidence—not a stream of every model announcement.

Follow via RSS ↗Download the evidence snapshot ↓

An opening snapshot, with the rules attached.

The first edition documents 14 architectures, 18 historical benchmark rows and 12 research milestones. The original wager remains separate from our experiment.

The guide opens with three explicit definitions of Transformer-like: an attention backbone, hybrids included, and linear attention included. The same rules apply across the model catalog and the historical benchmark experiment.

The two benchmark baskets preserve Tables 2 and 3 of the Jamba paper. They are useful because one has a particularly clear boundary case: three of the five long-context tasks are led by Jamba. Counting that hybrid changes the majority. It is an illustration from 2024, not a current frontier ranking.

The catalog includes 2026 documentation for Qwen3.5, Nemotron 3 Super and Gated DeltaNet-2. Their architecture entries do not add their newer performance claims to an older comparison table.

On this review date the original wager page displays Yes. We have not treated that running status as an official final resolution. Future evidence changes and corrections will be recorded here with their source and review date.

What makes a useful update?

A documented architecture, a comparable result, an independent replication, a correction, or an attributable resolution. A new model name by itself does not move the verdict.