The alternatives keep evolving, too.
Gated DeltaNet-2 separates erasing and writing within linear-attention memory updates.
This is an active research branch, not a fixed list of old challengers.
Gated DeltaNet-2 ↗2017—2026 / A short architectural history
Follow the ideas that made a straightforward bet harder to call. Move through the years to reveal the research behind the argument.
New hybrids and linear-attention proposals make the meaning of “Transformer-like” central to interpreting the bet.
A retrospective of publication milestones, not archived leaderboard standings. Selecting a year hides later milestones; the explanations are written with hindsight.
Gated DeltaNet-2 separates erasing and writing within linear-attention memory updates.
This is an active research branch, not a fixed list of old challengers.
Gated DeltaNet-2 ↗NVIDIA’s research release describes Nemotron 3 Super as a Mamba–Transformer hybrid with latent experts.
The 2027 wager now sits alongside a practical engineering question: which mixture works best for the workload?
Nemotron 3 Super research release ↗Qwen3.5’s documented language backbone combines Gated DeltaNet and gated attention.
A system can contain several kinds of computation. The component being classified has to be explicit.
Qwen3.5 official model card ↗Qwen3-Next pairs Gated DeltaNet with gated attention and sparse experts.
“Attention versus no attention” misses the mixture inside the model.
Qwen3-Next official model card ↗Nemotron-H uses Mamba layers alongside selected Transformer attention layers.
The proportion of attention can shrink while the system retains some of its capabilities.
Nemotron-H ↗Gated DeltaNet combines adaptive forgetting with the delta update rule.
Whether linear Transformers count deserves a separate rule from whether hybrids count.
Gated Delta Networks ↗Mamba-2’s state-space duality connects SSMs and variants of attention.
Mathematical relationships complicate the taxonomy. We still distinguish the implemented mechanism from a possible reformulation.
Transformers are SSMs ↗Jamba combines Mamba, attention and mixture-of-experts in one language model.
A win for a hybrid can support either side, depending on the definition. Our interactive comparison uses this paper.
Jamba ↗Selective state-space updates let information affect what the model retains and forgets.
Attention-free language modeling becomes a concrete, testable contender.
Mamba ↗Hyena combines long convolutions with input-controlled gates as a replacement for attention.
There is more than one route away from the attention operator.
Hyena Hierarchy ↗S4 develops an efficient structured state-space approach to long sequences.
A different sequence mechanism enters the conversation. A compelling direction is not yet a field-wide replacement.
Efficiently Modeling Long Sequences with Structured State Spaces ↗The Transformer paper introduces an attention-based encoder–decoder and reports machine-translation results.
The reference architecture is established. Later models can inherit the backbone without reproducing every detail.
Attention Is All You Need ↗01 Jan 2027 / The wager’s date
The timer reaches zero automatically. An official resolution must come from the people who made the bet. This guide will not turn a clock into a claim of victory.
How resolution will be recorded ↗