The architecture field guide

What’s inside
the model?

Names change. Building blocks tell a better story. Explore the mechanisms, open the sources, and decide where you draw the line.

14 entries · 8 qualify under your rules

A curated field guide, not a leaderboard.

Softmax attentionState-spaceLinear attentionRecurrenceConvolution

Sequence-mixing sketches only. They omit other components and are not exact layer maps.

2017 · Vaswani et al.

The original Transformer

The reference point. Attention replaces recurrence and convolution in the sequence backbone.

Transformer
Qualifies
Sequence mixing

Multi-head self-attention; encoder–decoder attention

Other building blocks

Dense feedforward layers

Why it matters to the bet

The 2017 design established the architecture this wager names. It was introduced for machine translation, not as a claim that attention alone solves every problem.

The original also uses feedforward layers, residual connections and normalization. The title is not a literal inventory of its components.

2024 · Meta

Llama 3 / 3.1

A dense decoder-only Transformer with grouped-query attention.

Transformer
Qualifies
Sequence mixing

Causal grouped-query attention (GQA)

Other building blocks

Dense feedforward layers

Why it matters to the bet

Sharing key/value heads changes the cost of attention without replacing the sequence-mixing mechanism. It qualifies under all three of our rules.

This entry covers the architecture documented in The Llama 3 Herd of Models, not every later model carrying the Llama name.

2024 · DeepSeek

DeepSeek-V3

Latent attention and sparse experts inside an attention-based backbone.

Transformer
Qualifies
Sequence mixing

Multi-head latent attention (MLA)

Other building blocks

DeepSeekMoE

Why it matters to the bet

MLA compresses the attention cache. Mixture-of-experts routes computation among feedforward experts. Neither change, on its own, makes this an attention-free model.

“MoE” describes expert routing, not a competing sequence-mixing family. This card is for V3, not a blanket classification of later DeepSeek releases.

2023 · Gu & Dao

Mamba

An input-dependent state-space model without an attention module.

Attention-free
Outside your rule
Sequence mixing

Selective state-space updates

Other building blocks

Integrated gated Mamba blocks

Why it matters to the bet

The model carries a compact state forward and selectively updates it as tokens arrive. It is a clear example of replacing the conventional attention mechanism.

Results at a particular model size do not establish a win across the field. Mamba layers can also be used inside hybrids; those are separate entries.

2024 · Dao & Gu

Mamba-2

A refined state-space model that exposes surprising links to attention.

Attention-free
Outside your rule
Sequence mixing

Structured state-space duality (SSD)

Other building blocks

Gated Mamba-2 blocks

Why it matters to the bet

Its paper connects structured state-space computations with variants of attention. Mathematical kinship does not mean the model contains conventional softmax attention layers.

We classify the pure Mamba-2 backbone as state-space. A Mamba-2–attention hybrid would be classified separately.

2024 · Peng et al.

RWKV-6 / Finch

Dynamic recurrence with a matrix-valued memory state.

Attention-free
Outside your rule
Sequence mixing

Recurrent time mixing

Other building blocks

Channel mixing

Why it matters to the bet

Finch updates a persistent state instead of retrieving from a growing softmax attention cache. We place it in the recurrent branch of this guide.

Recurrent and linear-attention formulations can overlap mathematically. Our broad rule follows the explicitly documented architectural family, not every possible equivalence.

2023 · Poli et al.

Hyena

Long convolutions and input-controlled gates replace attention.

Attention-free
Outside your rule
Sequence mixing

Implicit long convolutions with gating

Other building blocks

Feedforward projections

Why it matters to the bet

Hyena demonstrates a different way to mix information across a sequence. It belongs on the challenger map even though architecture alone says nothing about a model’s rank.

This is the original Hyena operator and language-model work. Later hybrids using Hyena and attention need their own classification.

2024 · AI21 Labs

Jamba

A mix of attention and Mamba, with sparse experts.

Hybrid
Qualifies
Sequence mixing

Attention + selective state-space layers

Other building blocks

Dense and mixture-of-experts layers

Why it matters to the bet

The published configuration has one attention layer for every seven Mamba layers. It is the concrete boundary case used in our verdict experiment.

The historical comparison uses the original Jamba paper. Later Jamba releases are not interchangeable with those scores.

2025 · NVIDIA

Nemotron-H

A Mamba–Transformer family built around inference efficiency.

Hybrid
Qualifies
Sequence mixing

Mamba layers with selected attention layers

Other building blocks

Feedforward layers

Why it matters to the bet

Most self-attention layers are replaced by Mamba layers. Under our narrow rule it falls outside; once hybrids qualify, it moves inside.

The diagram is schematic. Layer placement and model sizes vary across this family; speed depends on hardware and workload.

2025 · Qwen

Qwen3-Next

Gated DeltaNet and gated attention, plus sparse experts.

Hybrid
Qualifies
Sequence mixing

Linear attention + gated softmax attention

Other building blocks

Sparse mixture-of-experts

Why it matters to the bet

The 80B-A3B model interleaves three Gated DeltaNet layers with one gated-attention layer. It contains both mechanisms; calling it simply “linear attention” loses that distinction.

Hybrid architecture is different from a model switching between thinking and non-thinking modes. Both uses of “hybrid” appear in AI discussions.

2024 · Yang, Kautz & Hatamizadeh

Gated DeltaNet

A linear-attention model that combines forgetting with targeted memory updates.

Linear attention
Outside your rule
Sequence mixing

Gated delta-rule linear attention

Other building blocks

Feedforward layers

Why it matters to the bet

It is explicitly framed as a linear Transformer. Our broadest definition includes the pure model; the narrower definitions do not.

The paper also studies hybrids with sliding-window attention and Mamba-2. This card labels the pure Gated DeltaNet variant, not every experiment in the paper.

2026 · Qwen

Qwen3.5-397B-A17B

A multimodal model with linear and full attention in its language backbone.

Hybrid
Qualifies
Sequence mixing

Gated DeltaNet + gated attention

Other building blocks

Sparse mixture-of-experts

Why it matters to the bet

The official model description puts both mechanisms inside a large multimodal system. It extends the hybrid question beyond the earlier text-only comparisons.

This classification concerns the language backbone. It does not classify every vision component or transfer this model’s results into the 2024 benchmark basket.

2026 · NVIDIA

Nemotron 3 Super

A 2026 Mamba–Transformer hybrid with latent mixture-of-experts.

Hybrid
Qualifies
Sequence mixing

Mamba + Transformer attention

Other building blocks

LatentMoE

Why it matters to the bet

NVIDIA describes a 120B-total, 12B-active model that combines these ingredients. It is another documented hybrid, not evidence that attention has disappeared.

This entry describes Super. It is not a claim that all Nemotron models share one architecture, and the sketch does not encode exact layer counts.

2026 · Hatamizadeh, Choi & Kautz

Gated DeltaNet-2

Separate controls for erasing old memory and writing new information.

Linear attention
Outside your rule
Sequence mixing

Gated delta-rule linear attention

Other building blocks

Feedforward layers

Why it matters to the bet

This 2026 proposal further develops the linear-attention branch. It makes the broad definition useful without collapsing every recurrent architecture into one category.

A research result at the reported scale is not a frontier-wide verdict. The original preprint and its experiment settings remain the source of record.

And the models we can’t inspect?

We do not infer an architecture from a product name, a benchmark score, or a rumor. An undisclosed model belongs in an “unknown” category until a reliable source establishes enough detail. This catalog focuses on documented examples; it is not a census of frontier systems.

Read the classification policy ↗