Replacement
State-space and recurrent models can process language without conventional attention. The open question is how far that advantage travels across tasks, scales and training budgets.
Explore attention-free models →A field guide to the 2027 Transformer bet
A five-year bet. A moving definition.
An architecture that keeps changing shape.
Will Transformer-like models still lead most NLP benchmarks on January 1, 2027?
01 / Make the call
What counts as a Transformer? Change the rules and watch the same published results tell a different story.
Also count models that mix conventional attention with state-space, recurrent or linear-attention layers. A little attention is enough to qualify.
Jamba paper, Table 2 · March–July 2024
Under your rules
tasks led by a qualifying architecture
“Most” means more than half. This is a result within one paper’s comparison, not a verdict on the 2027 wager.
Try long-context QA, then switch between “Attention backbone” and “Include hybrids.” Jamba’s three task wins change sides.
Five models in the Jamba paper’s 2024 academic comparison. Each reported task or aggregate receives one vote; this selection includes coding and math. Models retain their published scores when a rule excludes their architecture; they do not disappear from the competition.
| Task | Llama 2 13B | Llama 2 70B | Gemma 7B | Mixtral 8×7B | Jamba | Leader qualifies? |
|---|---|---|---|---|---|---|
| HellaSwagSentence completion · 10-shot | 80.7 | 85.3 | 81.2 | 86.7 | 87.1 | Yes |
| WinoGrandePronoun resolution · 5-shot | 72.8 | 80.2 | 72.3 | 81.2 | 82.5 | Yes |
| ARC-EasyScience questions · 0-shot | 77.3 | 80.2 | 81.5 | 77.6 | 73.5 | Yes |
| ARC-ChallengeScience questions · 25-shot | 59.4 | 67.3 | 53.2 | 66.0 | 64.4 | Yes |
| PIQAPhysical commonsense · 0-shot | 80.5 | 82.8 | 81.2 | 83.0 | 83.2 | Yes |
| Natural QuestionsClosed-book QA · 5-shot | 37.7 | 46.9 | 32.6 | 44.8 | 45.9 | Yes |
| TruthfulQATruthfulness · 0-shot | 37.4 | 44.9 | 44.8 | 46.8 | 46.4 | Yes |
| BoolQYes/no questions · 10-shot | 81.7 | 85.0 | 87.2 | 88.4 | 88.2 | Yes |
| QuACConversational reading · 0-shot | 42.7 | 42.4 | 39.2 | 40.9 | 40.9 | Yes |
| GSM8KGrade-school math · 3-shot CoT | 34.7 | 55.3 | 54.5 | 60.4 | 59.9 | Yes |
| HumanEvalCode generation · pass@1 | 18.3 | 29.9 | 32.3 | 34.8 | 29.3 | Yes |
| MMLUKnowledge aggregate · 5-shot | 54.8 | 69.8 | 64.3 | 70.6 | 67.4 | Yes |
| BBHReasoning aggregate · 3-shot | 39.4 | 51.2 | 55.1 | 50.3 | 45.4 | Yes |
Author-reported, rounded scores; no claim of statistical significance or equal training budgets. Read the source ↗ · Download CSV ↓ · How we count
02 / The interesting part
State-space and recurrent models can process language without conventional attention. The open question is how far that advantage travels across tasks, scales and training budgets.
Explore attention-free models →Combine a few attention layers with a different sequence model. If that wins, did the Transformer survive—or did its replacement keep the useful bits?
Explore hybrid models →A leaderboard tells you who scored highest. It does not tell you where an architectural family ends. “Uses attention” and “is a Transformer” are different claims.
Read our working definitions →03 / What to watch
Three kinds of evidence worth following. A new release matters when it changes one of these arguments.
In the record. Mamba, Hyena and recurrent models document alternatives to conventional attention. Source ↗
What is still needed. A clearly defined task basket, comparable evaluations against strong contemporaries, and enough wins to establish a majority. Efficiency alone does not settle a quality-based wager.
In the record. Jamba already provides a historical example of a hybrid leading some tasks in a published comparison. Source ↗
What is still needed. An agreed rule about whether the hybrid counts. New benchmark scores cannot resolve a disagreement about the category itself.
In the record. Linear attention and state-space duality show why an architecture’s name is not a sufficient description. Source ↗
What is still needed. A rule about the mechanism being counted, plus documentation of the actual model. Mathematical equivalence and practical implementation are different kinds of evidence.
* The original wager names a date, not a time zone. This site uses midnight UTC for the countdown. On 28 September 2026, the original page displayed “Yes.” That is a dated status, not a final adjudication. Read the resolution policy.