Seven questions / A few useful distinctions

Still a
Transformer?

Read the ingredients before you see the name. Some answers are clear. Others depend on where you draw the line.

Question 1 of 7

A familiar backbone, shared keys.

A decoder-only model uses grouped-query attention. Multiple query heads share key and value heads, and dense feedforward layers do the rest.

Softmax attentionState-spaceLinear attentionRecurrenceConvolution

Schematic sequence-mixing blocks; not an exact layer map.

How would you classify it?

Progress is saved on this browser only. No account, public score, or vote is created.