2017—2026 / A short architectural history

How we got
to “it depends.”

Follow the ideas that made a straightforward bet harder to call. Move through the years to reveal the research behind the argument.

12 milestones in view

The clock runs down. Research carries on.

New hybrids and linear-attention proposals make the meaning of “Transformer-like” central to interpreting the bet.

A retrospective of publication milestones, not archived leaderboard standings. Selecting a year hides later milestones; the explanations are written with hindsight.

21 May 2026

The alternatives keep evolving, too.

Gated DeltaNet-2 separates erasing and writing within linear-attention memory updates.

This is an active research branch, not a fixed list of old challengers.

Gated DeltaNet-2 ↗
10 Mar 2026

A new generation keeps mixing ingredients.

NVIDIA’s research release describes Nemotron 3 Super as a Mamba–Transformer hybrid with latent experts.

The 2027 wager now sits alongside a practical engineering question: which mixture works best for the workload?

Nemotron 3 Super research release ↗
Feb 2026

The hybrid question becomes multimodal.

Qwen3.5’s documented language backbone combines Gated DeltaNet and gated attention.

A system can contain several kinds of computation. The component being classified has to be explicit.

Qwen3.5 official model card ↗
Sep 2025

Attention meets linear attention.

Qwen3-Next pairs Gated DeltaNet with gated attention and sparse experts.

“Attention versus no attention” misses the mixture inside the model.

Qwen3-Next official model card ↗
04 Apr 2025

Hybrids scale into model families.

Nemotron-H uses Mamba layers alongside selected Transformer attention layers.

The proportion of attention can shrink while the system retains some of its capabilities.

Nemotron-H ↗
09 Dec 2024

Linear attention learns to edit memory.

Gated DeltaNet combines adaptive forgetting with the delta update rule.

Whether linear Transformers count deserves a separate rule from whether hybrids count.

Gated Delta Networks ↗
31 May 2024

The family tree gets less tidy.

Mamba-2’s state-space duality connects SSMs and variants of attention.

Mathematical relationships complicate the taxonomy. We still distinguish the implemented mechanism from a possible reformulation.

Transformers are SSMs ↗
28 Mar 2024

The hybrid arrives.

Jamba combines Mamba, attention and mixture-of-experts in one language model.

A win for a hybrid can support either side, depending on the definition. Our interactive comparison uses this paper.

Jamba ↗
01 Dec 2023

Mamba makes selection recurrent.

Selective state-space updates let information affect what the model retains and forgets.

Attention-free language modeling becomes a concrete, testable contender.

Mamba ↗
21 Feb 2023

Convolutions make their case.

Hyena combines long convolutions with input-controlled gates as a replacement for attention.

There is more than one route away from the attention operator.

Hyena Hierarchy ↗
12 Jun 2017

Attention gets the starring role.

The Transformer paper introduces an attention-based encoder–decoder and reports machine-translation results.

The reference architecture is established. Later models can inherit the backbone without reproducing every detail.

Attention Is All You Need ↗

01 Jan 2027 / The wager’s date

A deadline, not the end of the story.

The timer reaches zero automatically. An official resolution must come from the people who made the bet. This guide will not turn a clock into a claim of victory.

How resolution will be recorded ↗