Seven questions / A few useful distinctions
Still a
Transformer?
Read the ingredients before you see the name. Some answers are clear. Others depend on where you draw the line.
Question 1 of 7
A familiar backbone, shared keys.
A decoder-only model uses grouped-query attention. Multiple query heads share key and value heads, and dense feedforward layers do the rest.
AAAAAAAA
ASoftmax attentionSState-spaceLLinear attentionRRecurrenceCConvolution
Schematic sequence-mixing blocks; not an exact layer map.
Progress is saved on this browser only. No account, public score, or vote is created.