{"version":"2026-09-28.1","reviewed":"2026-09-28","scope":"Curated architecture guide and historical paper comparisons; not a live frontier leaderboard or official adjudication","originalWager":{"url":"https://www.isattentionallyouneed.com/","observedStatus":"Yes","observedAt":"2026-09-28","deadlineDate":"2027-01-01","displayTimeZone":"UTC","finalResolution":null},"definitions":[{"id":"strict","label":"Attention backbone","short":"Attention backbone","description":"Sequence mixing uses conventional softmax attention. GQA, latent attention and mixture-of-experts still qualify. Attention–SSM and attention–linear hybrids do not."},{"id":"hybrids","label":"Include hybrids","short":"Hybrids included","description":"Also count models that mix conventional attention with state-space, recurrent or linear-attention layers. A little attention is enough to qualify."},{"id":"broad","label":"Include linear attention","short":"Linear attention included","description":"Also count pure linear-attention architectures. This is our broadest rule, not a claim that all recurrent or state-space models are Transformers."}],"models":[{"id":"transformer","name":"The original Transformer","by":"Vaswani et al.","year":2017,"family":"transformer","blocks":["A","A","A","A","A","A","A","A"],"summary":"The reference point. Attention replaces recurrence and convolution in the sequence backbone.","mixing":"Multi-head self-attention; encoder–decoder attention","feedforward":"Dense feedforward layers","why":"The 2017 design established the architecture this wager names. It was introduced for machine translation, not as a claim that attention alone solves every problem.","caveat":"The original also uses feedforward layers, residual connections and normalization. The title is not a literal inventory of its components.","source":"https://arxiv.org/abs/1706.03762","sourceTitle":"Attention Is All You Need"},{"id":"llama-3","name":"Llama 3 / 3.1","by":"Meta","year":2024,"family":"transformer","blocks":["A","A","A","A","A","A","A","A"],"summary":"A dense decoder-only Transformer with grouped-query attention.","mixing":"Causal grouped-query attention (GQA)","feedforward":"Dense feedforward layers","why":"Sharing key/value heads changes the cost of attention without replacing the sequence-mixing mechanism. It qualifies under all three of our rules.","caveat":"This entry covers the architecture documented in The Llama 3 Herd of Models, not every later model carrying the Llama name.","source":"https://arxiv.org/html/2407.21783v3#S3.SS2","sourceTitle":"The Llama 3 Herd of Models · §3.2"},{"id":"deepseek-v3","name":"DeepSeek-V3","by":"DeepSeek","year":2024,"family":"transformer","blocks":["A","A","A","A","A","A","A","A"],"summary":"Latent attention and sparse experts inside an attention-based backbone.","mixing":"Multi-head latent attention (MLA)","feedforward":"DeepSeekMoE","why":"MLA compresses the attention cache. Mixture-of-experts routes computation among feedforward experts. Neither change, on its own, makes this an attention-free model.","caveat":"“MoE” describes expert routing, not a competing sequence-mixing family. This card is for V3, not a blanket classification of later DeepSeek releases.","source":"https://arxiv.org/abs/2412.19437","sourceTitle":"DeepSeek-V3 Technical Report"},{"id":"mamba","name":"Mamba","by":"Gu & Dao","year":2023,"family":"attention-free","blocks":["S","S","S","S","S","S","S","S"],"summary":"An input-dependent state-space model without an attention module.","mixing":"Selective state-space updates","feedforward":"Integrated gated Mamba blocks","why":"The model carries a compact state forward and selectively updates it as tokens arrive. It is a clear example of replacing the conventional attention mechanism.","caveat":"Results at a particular model size do not establish a win across the field. Mamba layers can also be used inside hybrids; those are separate entries.","source":"https://arxiv.org/abs/2312.00752","sourceTitle":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces"},{"id":"mamba-2","name":"Mamba-2","by":"Dao & Gu","year":2024,"family":"attention-free","blocks":["S","S","S","S","S","S","S","S"],"summary":"A refined state-space model that exposes surprising links to attention.","mixing":"Structured state-space duality (SSD)","feedforward":"Gated Mamba-2 blocks","why":"Its paper connects structured state-space computations with variants of attention. Mathematical kinship does not mean the model contains conventional softmax attention layers.","caveat":"We classify the pure Mamba-2 backbone as state-space. A Mamba-2–attention hybrid would be classified separately.","source":"https://arxiv.org/abs/2405.21060","sourceTitle":"Transformers are SSMs: Structured State Space Duality"},{"id":"rwkv-6","name":"RWKV-6 / Finch","by":"Peng et al.","year":2024,"family":"attention-free","blocks":["R","R","R","R","R","R","R","R"],"summary":"Dynamic recurrence with a matrix-valued memory state.","mixing":"Recurrent time mixing","feedforward":"Channel mixing","why":"Finch updates a persistent state instead of retrieving from a growing softmax attention cache. We place it in the recurrent branch of this guide.","caveat":"Recurrent and linear-attention formulations can overlap mathematically. Our broad rule follows the explicitly documented architectural family, not every possible equivalence.","source":"https://arxiv.org/abs/2404.05892","sourceTitle":"Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence"},{"id":"hyena","name":"Hyena","by":"Poli et al.","year":2023,"family":"attention-free","blocks":["C","C","C","C","C","C","C","C"],"summary":"Long convolutions and input-controlled gates replace attention.","mixing":"Implicit long convolutions with gating","feedforward":"Feedforward projections","why":"Hyena demonstrates a different way to mix information across a sequence. It belongs on the challenger map even though architecture alone says nothing about a model’s rank.","caveat":"This is the original Hyena operator and language-model work. Later hybrids using Hyena and attention need their own classification.","source":"https://arxiv.org/abs/2302.10866","sourceTitle":"Hyena Hierarchy: Towards Larger Convolutional Language Models"},{"id":"jamba","name":"Jamba","by":"AI21 Labs","year":2024,"family":"hybrid","blocks":["S","S","S","A","S","S","S","S"],"summary":"A mix of attention and Mamba, with sparse experts.","mixing":"Attention + selective state-space layers","feedforward":"Dense and mixture-of-experts layers","why":"The published configuration has one attention layer for every seven Mamba layers. It is the concrete boundary case used in our verdict experiment.","caveat":"The historical comparison uses the original Jamba paper. Later Jamba releases are not interchangeable with those scores.","source":"https://arxiv.org/html/2403.19887v2","sourceTitle":"Jamba: A Hybrid Transformer-Mamba Language Model"},{"id":"nemotron-h","name":"Nemotron-H","by":"NVIDIA","year":2025,"family":"hybrid","blocks":["S","S","S","A","S","S","S","A"],"summary":"A Mamba–Transformer family built around inference efficiency.","mixing":"Mamba layers with selected attention layers","feedforward":"Feedforward layers","why":"Most self-attention layers are replaced by Mamba layers. Under our narrow rule it falls outside; once hybrids qualify, it moves inside.","caveat":"The diagram is schematic. Layer placement and model sizes vary across this family; speed depends on hardware and workload.","source":"https://arxiv.org/abs/2504.03624","sourceTitle":"Nemotron-H: Accurate and Efficient Hybrid Mamba-Transformer Models"},{"id":"qwen3-next","name":"Qwen3-Next","by":"Qwen","year":2025,"family":"hybrid","blocks":["L","L","L","A","L","L","L","A"],"summary":"Gated DeltaNet and gated attention, plus sparse experts.","mixing":"Linear attention + gated softmax attention","feedforward":"Sparse mixture-of-experts","why":"The 80B-A3B model interleaves three Gated DeltaNet layers with one gated-attention layer. It contains both mechanisms; calling it simply “linear attention” loses that distinction.","caveat":"Hybrid architecture is different from a model switching between thinking and non-thinking modes. Both uses of “hybrid” appear in AI discussions.","source":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct","sourceTitle":"Qwen3-Next-80B-A3B-Instruct · official model card"},{"id":"gated-deltanet","name":"Gated DeltaNet","by":"Yang, Kautz & Hatamizadeh","year":2024,"family":"linear","blocks":["L","L","L","L","L","L","L","L"],"summary":"A linear-attention model that combines forgetting with targeted memory updates.","mixing":"Gated delta-rule linear attention","feedforward":"Feedforward layers","why":"It is explicitly framed as a linear Transformer. Our broadest definition includes the pure model; the narrower definitions do not.","caveat":"The paper also studies hybrids with sliding-window attention and Mamba-2. This card labels the pure Gated DeltaNet variant, not every experiment in the paper.","source":"https://arxiv.org/abs/2412.06464","sourceTitle":"Gated Delta Networks: Improving Mamba2 with Delta Rule"},{"id":"qwen3-5","name":"Qwen3.5-397B-A17B","by":"Qwen","year":2026,"family":"hybrid","blocks":["L","L","L","A","L","L","L","A"],"summary":"A multimodal model with linear and full attention in its language backbone.","mixing":"Gated DeltaNet + gated attention","feedforward":"Sparse mixture-of-experts","why":"The official model description puts both mechanisms inside a large multimodal system. It extends the hybrid question beyond the earlier text-only comparisons.","caveat":"This classification concerns the language backbone. It does not classify every vision component or transfer this model’s results into the 2024 benchmark basket.","source":"https://huggingface.co/Qwen/Qwen3.5-397B-A17B","sourceTitle":"Qwen3.5-397B-A17B · official model card"},{"id":"nemotron-3-super","name":"Nemotron 3 Super","by":"NVIDIA","year":2026,"family":"hybrid","blocks":["S","S","A","S","S","S","A","S"],"summary":"A 2026 Mamba–Transformer hybrid with latent mixture-of-experts.","mixing":"Mamba + Transformer attention","feedforward":"LatentMoE","why":"NVIDIA describes a 120B-total, 12B-active model that combines these ingredients. It is another documented hybrid, not evidence that attention has disappeared.","caveat":"This entry describes Super. It is not a claim that all Nemotron models share one architecture, and the sketch does not encode exact layer counts.","source":"https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/","sourceTitle":"NVIDIA Nemotron 3 Super · research release"},{"id":"gated-deltanet-2","name":"Gated DeltaNet-2","by":"Hatamizadeh, Choi & Kautz","year":2026,"family":"linear","blocks":["L","L","L","L","L","L","L","L"],"summary":"Separate controls for erasing old memory and writing new information.","mixing":"Gated delta-rule linear attention","feedforward":"Feedforward layers","why":"This 2026 proposal further develops the linear-attention branch. It makes the broad definition useful without collapsing every recurrent architecture into one category.","caveat":"A research result at the reported scale is not a frontier-wide verdict. The original preprint and its experiment settings remain the source of record.","source":"https://arxiv.org/abs/2605.22791","sourceTitle":"Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention"}],"benchmarks":{"source":"https://arxiv.org/html/2403.19887v2","benchmarkModels":[{"id":"llama13","name":"Llama 2 13B","family":"transformer"},{"id":"llama70","name":"Llama 2 70B","family":"transformer"},{"id":"gemma","name":"Gemma 7B","family":"transformer"},{"id":"mixtral","name":"Mixtral 8×7B","family":"transformer"},{"id":"jamba","name":"Jamba","family":"hybrid"}],"baskets":{"language":{"title":"General language · 13 tasks","table":"Table 2","metric":"Reported score","description":"Five models in the Jamba paper’s 2024 academic comparison. Each reported task or aggregate receives one vote; this selection includes coding and math.","tasks":[{"name":"HellaSwag","description":"Sentence completion · 10-shot","scores":{"llama13":80.7,"llama70":85.3,"gemma":81.2,"mixtral":86.7,"jamba":87.1}},{"name":"WinoGrande","description":"Pronoun resolution · 5-shot","scores":{"llama13":72.8,"llama70":80.2,"gemma":72.3,"mixtral":81.2,"jamba":82.5}},{"name":"ARC-Easy","description":"Science questions · 0-shot","scores":{"llama13":77.3,"llama70":80.2,"gemma":81.5,"mixtral":77.6,"jamba":73.5}},{"name":"ARC-Challenge","description":"Science questions · 25-shot","scores":{"llama13":59.4,"llama70":67.3,"gemma":53.2,"mixtral":66,"jamba":64.4}},{"name":"PIQA","description":"Physical commonsense · 0-shot","scores":{"llama13":80.5,"llama70":82.8,"gemma":81.2,"mixtral":83,"jamba":83.2}},{"name":"Natural Questions","description":"Closed-book QA · 5-shot","scores":{"llama13":37.7,"llama70":46.9,"gemma":32.6,"mixtral":44.8,"jamba":45.9}},{"name":"TruthfulQA","description":"Truthfulness · 0-shot","scores":{"llama13":37.4,"llama70":44.9,"gemma":44.8,"mixtral":46.8,"jamba":46.4}},{"name":"BoolQ","description":"Yes/no questions · 10-shot","scores":{"llama13":81.7,"llama70":85,"gemma":87.2,"mixtral":88.4,"jamba":88.2}},{"name":"QuAC","description":"Conversational reading · 0-shot","scores":{"llama13":42.7,"llama70":42.4,"gemma":39.2,"mixtral":40.9,"jamba":40.9}},{"name":"GSM8K","description":"Grade-school math · 3-shot CoT","scores":{"llama13":34.7,"llama70":55.3,"gemma":54.5,"mixtral":60.4,"jamba":59.9}},{"name":"HumanEval","description":"Code generation · pass@1","scores":{"llama13":18.3,"llama70":29.9,"gemma":32.3,"mixtral":34.8,"jamba":29.3}},{"name":"MMLU","description":"Knowledge aggregate · 5-shot","scores":{"llama13":54.8,"llama70":69.8,"gemma":64.3,"mixtral":70.6,"jamba":67.4}},{"name":"BBH","description":"Reasoning aggregate · 3-shot","scores":{"llama13":39.4,"llama70":51.2,"gemma":55.1,"mixtral":50.3,"jamba":45.4}}]},"long-context":{"title":"Long-context QA · 5 tasks","table":"Table 3","metric":"F1","description":"Jamba and Mixtral in the same paper’s 3-shot, long-context question-answering setup. F1 ranges from 0 to 1; higher is better.","tasks":[{"name":"LongFQA","description":"Financial documents","scores":{"mixtral":0.42,"jamba":0.44}},{"name":"CUAD","description":"Contract questions","scores":{"mixtral":0.46,"jamba":0.44}},{"name":"NarrativeQA","description":"Long narratives","scores":{"mixtral":0.29,"jamba":0.3}},{"name":"Natural Questions","description":"Long-context Wikipedia QA","scores":{"mixtral":0.58,"jamba":0.6}},{"name":"SFiction","description":"Science fiction","scores":{"mixtral":0.42,"jamba":0.4}}]}}},"milestones":[{"year":2017,"date":"12 Jun 2017","title":"Attention gets the starring role.","text":"The Transformer paper introduces an attention-based encoder–decoder and reports machine-translation results.","implication":"The reference architecture is established. Later models can inherit the backbone without reproducing every detail.","source":"https://arxiv.org/abs/1706.03762","sourceLabel":"Attention Is All You Need"},{"year":2021,"date":"01 Nov 2021","title":"State spaces become a serious alternative.","text":"S4 develops an efficient structured state-space approach to long sequences.","implication":"A different sequence mechanism enters the conversation. A compelling direction is not yet a field-wide replacement.","source":"https://arxiv.org/abs/2111.00396","sourceLabel":"Efficiently Modeling Long Sequences with Structured State Spaces"},{"year":2023,"date":"21 Feb 2023","title":"Convolutions make their case.","text":"Hyena combines long convolutions with input-controlled gates as a replacement for attention.","implication":"There is more than one route away from the attention operator.","source":"https://arxiv.org/abs/2302.10866","sourceLabel":"Hyena Hierarchy"},{"year":2023,"date":"01 Dec 2023","title":"Mamba makes selection recurrent.","text":"Selective state-space updates let information affect what the model retains and forgets.","implication":"Attention-free language modeling becomes a concrete, testable contender.","source":"https://arxiv.org/abs/2312.00752","sourceLabel":"Mamba"},{"year":2024,"date":"28 Mar 2024","title":"The hybrid arrives.","text":"Jamba combines Mamba, attention and mixture-of-experts in one language model.","implication":"A win for a hybrid can support either side, depending on the definition. Our interactive comparison uses this paper.","source":"https://arxiv.org/abs/2403.19887","sourceLabel":"Jamba"},{"year":2024,"date":"31 May 2024","title":"The family tree gets less tidy.","text":"Mamba-2’s state-space duality connects SSMs and variants of attention.","implication":"Mathematical relationships complicate the taxonomy. We still distinguish the implemented mechanism from a possible reformulation.","source":"https://arxiv.org/abs/2405.21060","sourceLabel":"Transformers are SSMs"},{"year":2024,"date":"09 Dec 2024","title":"Linear attention learns to edit memory.","text":"Gated DeltaNet combines adaptive forgetting with the delta update rule.","implication":"Whether linear Transformers count deserves a separate rule from whether hybrids count.","source":"https://arxiv.org/abs/2412.06464","sourceLabel":"Gated Delta Networks"},{"year":2025,"date":"04 Apr 2025","title":"Hybrids scale into model families.","text":"Nemotron-H uses Mamba layers alongside selected Transformer attention layers.","implication":"The proportion of attention can shrink while the system retains some of its capabilities.","source":"https://arxiv.org/abs/2504.03624","sourceLabel":"Nemotron-H"},{"year":2025,"date":"Sep 2025","title":"Attention meets linear attention.","text":"Qwen3-Next pairs Gated DeltaNet with gated attention and sparse experts.","implication":"“Attention versus no attention” misses the mixture inside the model.","source":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct","sourceLabel":"Qwen3-Next official model card"},{"year":2026,"date":"Feb 2026","title":"The hybrid question becomes multimodal.","text":"Qwen3.5’s documented language backbone combines Gated DeltaNet and gated attention.","implication":"A system can contain several kinds of computation. The component being classified has to be explicit.","source":"https://huggingface.co/Qwen/Qwen3.5-397B-A17B","sourceLabel":"Qwen3.5 official model card"},{"year":2026,"date":"10 Mar 2026","title":"A new generation keeps mixing ingredients.","text":"NVIDIA’s research release describes Nemotron 3 Super as a Mamba–Transformer hybrid with latent experts.","implication":"The 2027 wager now sits alongside a practical engineering question: which mixture works best for the workload?","source":"https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/","sourceLabel":"Nemotron 3 Super research release"},{"year":2026,"date":"21 May 2026","title":"The alternatives keep evolving, too.","text":"Gated DeltaNet-2 separates erasing and writing within linear-attention memory updates.","implication":"This is an active research branch, not a fixed list of old challengers.","source":"https://arxiv.org/abs/2605.22791","sourceLabel":"Gated DeltaNet-2"}]}