The small print, in readable type
Show your
working.
Definitions first. Evidence second. Conclusions with a date attached. Here is how this guide makes its calls.
A field guide with an experiment.
The original wager asks whether Transformer-like models will lead most benchmarked natural-language-processing tasks on January 1, 2027. Jonathan Frankle takes the affirmative; Sasha Rush takes the opposing side. The original page describes a charitable donation of equity as the stake. Read the original terms ↗
This is an independent companion. Neither participant has approved our definitions, chosen our benchmark baskets, or delegated adjudication to this site.
The interactive verdict illustrates how architecture definitions and task selection affect an answer. The model catalog explains documented mechanisms. Neither is a live ranking of every model in the field.
Three explicit definitions.
Attention backbone
Sequence mixing uses conventional softmax attention. GQA, latent attention and mixture-of-experts still qualify. Attention–SSM and attention–linear hybrids do not.
Include hybrids
Also count models that mix conventional attention with state-space, recurrent or linear-attention layers. A little attention is enough to qualify.
Include linear attention
Also count pure linear-attention architectures. This is our broadest rule, not a claim that all recurrent or state-space models are Transformers.
These are editorial conventions, not universal scientific categories. Our narrow rule describes an attention-based sequence backbone, not a byte-for-byte copy of the 2017 architecture. Mixture-of-experts, grouped-query attention and latent attention can all appear within it.
For the broad rule, “linear attention” means an architecture explicitly documented that way. It does not automatically include every recurrent or state-space model with a related mathematical formulation. RWKV and pure Mamba models remain in their documented recurrent or state-space branches here.
Unknown is not a vote for either side.
If reliable documentation does not establish the relevant mechanism, the classification is unknown. A product name, an anonymous claim or a high benchmark score is not enough. A published family description can support a family-level claim while leaving exact layer details undisclosed.
Count tasks. Keep the denominator.
The initial experiment reproduces two tables from the Jamba paper, version 2: 13 academic task or aggregate columns in Table 2, and five long-context QA tasks in Table 3. We do not count either table’s average as an extra task.
- Use the highest reported score within each row. These tables use metrics where higher is better.
- Classify the row’s winner under the selected definition. Excluding its architecture does not delete its score or promote the runner-up.
- Give each task one vote. “Most” means strictly more than half of all tasks in the selected table.
- If tied leaders all qualify, count one qualifying task. If they all fall outside, count one outside task. A tie across the boundary is unresolved; an unknown leader is unknown.
- Keep unresolved and unknown tasks in the denominator. Say “Yes” only if the known qualifying tasks already exceed half. Say “No” only if even resolving every uncertain task in favor would not produce a majority. Otherwise the result is insufficient evidence.
The displayed scores are author-reported and rounded. We do not infer statistically meaningful differences from a narrow numerical lead. Models differ in size, data and training compute, so these are published comparisons, not a controlled experiment isolating architecture.
Task weighting is a choice. MMLU and BBH each count once here even though they contain multiple subtasks; the academic basket also includes coding and math. The two baskets are separate experiments and are never pooled. They are not an agreed operational definition of “most NLP tasks” in 2027.
What is live, and what is curated?
The countdown follows your device’s clock. The architecture cards, research timeline, source review dates and benchmark tables are curated records. No background crawler updates them, and no model API generates a verdict when you visit.
Sources were reviewed on September 28, 2026. Publication dates describe the source, not when this site first observed it. The timeline is a retrospective, not an archive of past leaderboard states.
Each model card links to a paper or a developer’s official documentation. Layer strips are schematic illustrations of sequence-mixing mechanisms. They omit normalization, residual connections, feedforward layers and other components; they must not be read as exact layer counts.
Corrections and future evidence changes belong in the dated field notes. To reproduce today’s evidence, download the versioned evidence snapshot. A changed source or a later checkpoint does not silently rewrite a historical score.
When the countdown reaches zero.
The original wager specifies January 1, 2027 without a time zone. Our display uses 00:00 UTC as a transparent convention, not an additional term of the original bet.
At the deadline, the countdown stops at zero and points here. It does not declare a winner. On September 28, 2026, the original page displayed “Current Status: Yes”; we record that as a dated observation, not a final settlement.
An official resolution will require an attributable statement from the wager’s participants or their designated adjudicator. Any independent assessment on this site must remain separately labeled, with its benchmark selection, rules and unresolved evidence visible.
The question continues after the date. A historical wager can close while research into attention, recurrence and hybrids carries on.
Just enough vocabulary.
- Attention
- A way to combine information from tokens according to how relevant they are to one another. Conventional Transformer attention uses a softmax weighting operation.
- Sequence mixing
- The mechanism that lets information move between positions in a sequence. Attention, recurrence and convolution offer different ways to do it.
- State-space model / SSM
- A model that evolves a state as it processes inputs. Selective variants make parts of that update depend on the current input.
- Linear attention
- A family of attention formulations that avoid forming the full pairwise softmax attention matrix. Many can be computed through recurrent state updates.
- Mixture-of-experts / MoE
- A routing scheme that sends a token to a subset of parameter groups. It can coexist with attention, state-space layers or other mechanisms.
- KV cache
- Stored attention keys and values from earlier tokens, reused during generation. The memory cost depends on the model and attention design.
- Benchmark
- A particular evaluation with a dataset, metric and testing protocol. A score is meaningful only alongside those details.
- F1
- A score balancing precision and recall. The long-context QA table reports F1 on a scale from zero to one.
- Zero-shot / few-shot
- How many worked examples are included in the prompt before the evaluated question. Scores from different prompting setups should not be casually combined.
Go straight to the sources.
Primary documentation used in the guide. Model classifications are our interpretations of these sources.
- Is Attention All You Need? — the original wager ↗
- Attention Is All You Need ↗Vaswani et al. · 2017
- The Llama 3 Herd of Models · §3.2 ↗Meta · 2024
- DeepSeek-V3 Technical Report ↗DeepSeek · 2024
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces ↗Gu & Dao · 2023
- Transformers are SSMs: Structured State Space Duality ↗Dao & Gu · 2024
- Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence ↗Peng et al. · 2024
- Hyena Hierarchy: Towards Larger Convolutional Language Models ↗Poli et al. · 2023
- Jamba: A Hybrid Transformer-Mamba Language Model ↗AI21 Labs · 2024
- Nemotron-H: Accurate and Efficient Hybrid Mamba-Transformer Models ↗NVIDIA · 2025
- Qwen3-Next-80B-A3B-Instruct · official model card ↗Qwen · 2025
- Gated Delta Networks: Improving Mamba2 with Delta Rule ↗Yang, Kautz & Hatamizadeh · 2024
- Qwen3.5-397B-A17B · official model card ↗Qwen · 2026
- NVIDIA Nemotron 3 Super · research release ↗NVIDIA · 2026
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention ↗Hatamizadeh, Choi & Kautz · 2026
- Efficiently Modeling Long Sequences with Structured State Spaces ↗Gu, Goel & Ré · 2021
A small site, a small footprint.
Reading and changing rules require no account. Quiz progress is stored locally in your browser and can be cleared with “Start over.” Sharing copies a link or result; it does not publish it anywhere. The site does not collect email addresses or run advertising analytics. Hosting providers may retain ordinary request logs.
Help keep the guide going.
This guide is free to read and explore. If you find it useful, you can support the field guide on Buy Me a Coffee ↗. Contributions help with research, writing and upkeep. Support is entirely optional and does not influence the evidence, classifications or conclusions.
Contributions support this independent site and are separate from the charitable donation described in the original wager. The support link takes you to Buy Me a Coffee, which handles checkout and payment details under its own terms and privacy policy.