The architecture field guide
What’s inside
the model?
Names change. Building blocks tell a better story. Explore the mechanisms, open the sources, and decide where you draw the line.
Sequence-mixing sketches only. They omit other components and are not exact layer maps.
2017 · Vaswani et al.The original Transformer
The reference point. Attention replaces recurrence and convolution in the sequence backbone.
AAAAAAAATransformerQualifies
The original Transformer
The reference point. Attention replaces recurrence and convolution in the sequence backbone.
Multi-head self-attention; encoder–decoder attention
Dense feedforward layers
Why it matters to the bet
The 2017 design established the architecture this wager names. It was introduced for machine translation, not as a claim that attention alone solves every problem.
The original also uses feedforward layers, residual connections and normalization. The title is not a literal inventory of its components.
2024 · MetaLlama 3 / 3.1
A dense decoder-only Transformer with grouped-query attention.
AAAAAAAATransformerQualifies
Llama 3 / 3.1
A dense decoder-only Transformer with grouped-query attention.
Causal grouped-query attention (GQA)
Dense feedforward layers
Why it matters to the bet
Sharing key/value heads changes the cost of attention without replacing the sequence-mixing mechanism. It qualifies under all three of our rules.
This entry covers the architecture documented in The Llama 3 Herd of Models, not every later model carrying the Llama name.
2024 · DeepSeekDeepSeek-V3
Latent attention and sparse experts inside an attention-based backbone.
AAAAAAAATransformerQualifies
DeepSeek-V3
Latent attention and sparse experts inside an attention-based backbone.
Multi-head latent attention (MLA)
DeepSeekMoE
Why it matters to the bet
MLA compresses the attention cache. Mixture-of-experts routes computation among feedforward experts. Neither change, on its own, makes this an attention-free model.
“MoE” describes expert routing, not a competing sequence-mixing family. This card is for V3, not a blanket classification of later DeepSeek releases.
2023 · Gu & DaoMamba
An input-dependent state-space model without an attention module.
SSSSSSSSAttention-freeOutside your rule
Mamba
An input-dependent state-space model without an attention module.
Selective state-space updates
Integrated gated Mamba blocks
Why it matters to the bet
The model carries a compact state forward and selectively updates it as tokens arrive. It is a clear example of replacing the conventional attention mechanism.
Results at a particular model size do not establish a win across the field. Mamba layers can also be used inside hybrids; those are separate entries.
2024 · Dao & GuMamba-2
A refined state-space model that exposes surprising links to attention.
SSSSSSSSAttention-freeOutside your rule
Mamba-2
A refined state-space model that exposes surprising links to attention.
Structured state-space duality (SSD)
Gated Mamba-2 blocks
Why it matters to the bet
Its paper connects structured state-space computations with variants of attention. Mathematical kinship does not mean the model contains conventional softmax attention layers.
We classify the pure Mamba-2 backbone as state-space. A Mamba-2–attention hybrid would be classified separately.
2024 · Peng et al.RWKV-6 / Finch
Dynamic recurrence with a matrix-valued memory state.
RRRRRRRRAttention-freeOutside your rule
RWKV-6 / Finch
Dynamic recurrence with a matrix-valued memory state.
Recurrent time mixing
Channel mixing
Why it matters to the bet
Finch updates a persistent state instead of retrieving from a growing softmax attention cache. We place it in the recurrent branch of this guide.
Recurrent and linear-attention formulations can overlap mathematically. Our broad rule follows the explicitly documented architectural family, not every possible equivalence.
2023 · Poli et al.Hyena
Long convolutions and input-controlled gates replace attention.
CCCCCCCCAttention-freeOutside your rule
Hyena
Long convolutions and input-controlled gates replace attention.
Implicit long convolutions with gating
Feedforward projections
Why it matters to the bet
Hyena demonstrates a different way to mix information across a sequence. It belongs on the challenger map even though architecture alone says nothing about a model’s rank.
This is the original Hyena operator and language-model work. Later hybrids using Hyena and attention need their own classification.
2024 · AI21 LabsJamba
A mix of attention and Mamba, with sparse experts.
SSSASSSSHybridQualifies
Jamba
A mix of attention and Mamba, with sparse experts.
Attention + selective state-space layers
Dense and mixture-of-experts layers
Why it matters to the bet
The published configuration has one attention layer for every seven Mamba layers. It is the concrete boundary case used in our verdict experiment.
The historical comparison uses the original Jamba paper. Later Jamba releases are not interchangeable with those scores.
2025 · NVIDIANemotron-H
A Mamba–Transformer family built around inference efficiency.
SSSASSSAHybridQualifies
Nemotron-H
A Mamba–Transformer family built around inference efficiency.
Mamba layers with selected attention layers
Feedforward layers
Why it matters to the bet
Most self-attention layers are replaced by Mamba layers. Under our narrow rule it falls outside; once hybrids qualify, it moves inside.
The diagram is schematic. Layer placement and model sizes vary across this family; speed depends on hardware and workload.
2025 · QwenQwen3-Next
Gated DeltaNet and gated attention, plus sparse experts.
LLLALLLAHybridQualifies
Qwen3-Next
Gated DeltaNet and gated attention, plus sparse experts.
Linear attention + gated softmax attention
Sparse mixture-of-experts
Why it matters to the bet
The 80B-A3B model interleaves three Gated DeltaNet layers with one gated-attention layer. It contains both mechanisms; calling it simply “linear attention” loses that distinction.
Hybrid architecture is different from a model switching between thinking and non-thinking modes. Both uses of “hybrid” appear in AI discussions.
2024 · Yang, Kautz & HatamizadehGated DeltaNet
A linear-attention model that combines forgetting with targeted memory updates.
LLLLLLLLLinear attentionOutside your rule
Gated DeltaNet
A linear-attention model that combines forgetting with targeted memory updates.
Gated delta-rule linear attention
Feedforward layers
Why it matters to the bet
It is explicitly framed as a linear Transformer. Our broadest definition includes the pure model; the narrower definitions do not.
The paper also studies hybrids with sliding-window attention and Mamba-2. This card labels the pure Gated DeltaNet variant, not every experiment in the paper.
2026 · QwenQwen3.5-397B-A17B
A multimodal model with linear and full attention in its language backbone.
LLLALLLAHybridQualifies
Qwen3.5-397B-A17B
A multimodal model with linear and full attention in its language backbone.
Gated DeltaNet + gated attention
Sparse mixture-of-experts
Why it matters to the bet
The official model description puts both mechanisms inside a large multimodal system. It extends the hybrid question beyond the earlier text-only comparisons.
This classification concerns the language backbone. It does not classify every vision component or transfer this model’s results into the 2024 benchmark basket.
2026 · NVIDIANemotron 3 Super
A 2026 Mamba–Transformer hybrid with latent mixture-of-experts.
SSASSSASHybridQualifies
Nemotron 3 Super
A 2026 Mamba–Transformer hybrid with latent mixture-of-experts.
Mamba + Transformer attention
LatentMoE
Why it matters to the bet
NVIDIA describes a 120B-total, 12B-active model that combines these ingredients. It is another documented hybrid, not evidence that attention has disappeared.
This entry describes Super. It is not a claim that all Nemotron models share one architecture, and the sketch does not encode exact layer counts.
2026 · Hatamizadeh, Choi & KautzGated DeltaNet-2
Separate controls for erasing old memory and writing new information.
LLLLLLLLLinear attentionOutside your rule
Gated DeltaNet-2
Separate controls for erasing old memory and writing new information.
Gated delta-rule linear attention
Feedforward layers
Why it matters to the bet
This 2026 proposal further develops the linear-attention branch. It makes the broad definition useful without collapsing every recurrent architecture into one category.
A research result at the reported scale is not a frontier-wide verdict. The original preprint and its experiment settings remain the source of record.
And the models we can’t inspect?
We do not infer an architecture from a product name, a benchmark score, or a rumor. An undisclosed model belongs in an “unknown” category until a reliable source establishes enough detail. This catalog focuses on documented examples; it is not a census of frontier systems.
Read the classification policy ↗