What are Queries, Keys and Values in an LLM?
Three learned projections of the same token embedding. The query is what this token is looking for; the key is what it advertises to others; the value is what it hands over when attended to. Scores come from query · key; the output is a weighted sum of values.
Overview
Why three and not one
If a token used the same vector to search with and to be found by, attention would collapse into plain similarity: tokens would attend to tokens like themselves. Separating query from key lets a token look for something different from itself — a verb seeking its subject, a pronoun seeking its referent.
Separating value from key matters just as much. What makes a token worth attending to is not necessarily what it should contribute once attended to. The key is the address; the value is the payload.
Parameters
Visualisation
—Readout
What to watch
- All three come from one embedding through three different matrices.
- The score uses K; the output uses V — they are deliberately separate.
- K and V for a past token are fixed forever, which is the cache's premise.
What are Queries, Keys and Values in an LLM?: A Practical Guide
What do Q, K and V actually mean in attention, and why are they three separate things?
The mechanism in one line
softmax(QKT / √d) V. The dot products score every query against every key; the division keeps the softmax out of its saturated region, where gradients vanish; softmax turns scores into weights summing to one; and those weights are applied to the values.
In self-attention all three come from the same sequence. In cross-attention the queries come from one sequence and the keys and values from another — which is how a decoder reads an encoder, and the same shape as retrieval: a query against a set of keys.
Why this is a serving question too
Generating token n needs the keys and values of all previous tokens — and those never change, because each depends only on tokens before it. So they are computed once and kept: the KV cache.
Without it every new token would recompute K and V for the whole prefix, making generation quadratic. With it, each step is linear in the context length. The cost is memory: the cache grows with sequence length × layers × heads, and it is usually what limits how many requests a GPU can serve at once.
Where the shapes come from
For a sequence of n tokens with model dimension d, the three projections are matrices of shape d×dk, so Q, K and V each come out n×dk. QKT is then n×n — every token scored against every token — and multiplying that by V returns to n×dk.
That n×n matrix is why attention is quadratic in sequence length, and why context windows were hard to extend. It is also why flash attention and its relatives matter: they compute the same result without ever materialising the full matrix in memory.
In multi-head attention the projections are split into h heads of dimension d/h, each computing its own attention, with the outputs concatenated. Same arithmetic, run in parallel over subspaces, so different heads can specialise — some track syntax, some track position, some appear to do very little.
Masking, and what it has to do with the cache
A decoder must not attend to tokens it has not generated yet, so the scores above the diagonal are set to negative infinity before the softmax, which drives their weights to zero. That is causal masking, and it is what makes the KV cache correct rather than merely convenient: since token i can only ever attend to tokens up to i, its keys and values can never be affected by anything generated later.
The cache's cost is memory, and it is the usual limit on how many requests a GPU can serve at once: roughly 2 × layers × heads × head dimension × sequence length × batch, per precision byte. Grouped-query and multi-query attention exist mostly to shrink it, by sharing K and V across heads while keeping separate queries — a large memory saving for a small quality cost.
A lookup that returns a blend
Attention is easiest to read as a dictionary lookup that returns a mixture rather than one entry.
An ordinary dictionary takes a key and returns exactly one value. Attention takes a query, compares it against every key, and returns a weighted mixture of all the values — weighted by how well each key matched.
- Query — what this position is looking for.
- Key — the label each position advertises.
- Value — what each position contributes when selected.
All three are linear projections of the same input, with separate learned matrices:
Q = XWQ K = XWK V = XWV
That is where every parameter in the mechanism lives. Attention itself has no weights — it is a fixed formula applied to those three projections.
Why three separate projections
If all three come from the same input, why not use the input directly for all of them?
Because the three roles want different information. What a token is looking for is not the same as how it wants to be found, and neither is the same as what it should contribute.
Take "it" in "The cat sat on the mat because it was warm". As a query, "it" needs to search for a nearby noun that could plausibly be warm. As a key, "mat" needs to advertise itself as a warmable surface. As a value, "mat" contributes its full semantic content.
The clean demonstration is that Q and K are used only to compute compatibility, and V only to produce output. A token can therefore be attended to for one reason and contribute something quite different — which a single shared projection could not express.
The computation, with shapes
Attention(Q, K, V) = softmax(QKᵀ / √dk) V
For 4 tokens with head dimension 64:
| Step | Operation | Shape |
|---|---|---|
| 1 | QKᵀ — every query against every key | 4×4 |
| 2 | Divide by √64 = 8 | 4×4 |
| 3 | Softmax over each row | 4×4, rows sum to 1 |
| 4 | Multiply by V | 4×64 |
Row i of the 4×4 matrix is token i's attention weights over all tokens. Multiplying by V replaces token i's representation with the weighted mixture.
The √dₖ division is not cosmetic. Dot products of d-dimensional vectors grow with d, so without scaling the scores reach the tens or hundreds, the softmax saturates into something nearly one-hot, and the gradient through it approaches zero. Dividing by √dₖ keeps the score variance near 1 regardless of dimension.
import torch, torch.nn.functional as F
def attention(x, Wq, Wk, Wv, mask=None):
Q, K, V = x @ Wq, x @ Wk, x @ Wv
scores = Q @ K.transpose(-2, -1) / (K.size(-1) ** 0.5)
if mask is not None:
scores = scores.masked_fill(mask == 0, float("-inf"))
return F.softmax(scores, dim=-1) @ VSetting a score to negative infinity makes its softmax weight exactly zero — which is how both causal masking and padding masks work.
Three projections, and why you need all three
Attention gives every token a query, a key and a value, and the usual explanation stops at a database analogy. This builds the mechanism from vectors you can read, then removes each projection in turn to show what breaks -- which is the only way to see why three are needed rather than one.
Things to try
- Attend from sat. It puts most of its weight on cat, its subject — which a token searching with its own embedding could never do.
- Raise the score scale to 4. The softmax collapses to nearly one-hot; that saturation is what dividing by √d exists to prevent.
- Switch the focus token and watch the whole distribution move. Each token's query is asking a different question of the same keys.
What to remember
Q, K and V are three learned projections of one embedding. The query is what a token looks for, the key is what it advertises, the value is what it contributes. Scores come from Q·K and the output is a weighted sum of V. Because K and V for a token never change once computed, they can be cached — which is what makes generation linear rather than quadratic.
Self-attention and cross-attention
In self-attention, Q, K and V all come from the same sequence — a sentence relating its own words to each other.
In cross-attention, the queries come from one sequence and the keys and values from another. A decoder queries an encoder's output; a text query attends to image patches in a multimodal model.
| Type | Q from | K, V from |
|---|---|---|
| Encoder self-attention | The sequence | The same sequence |
| Decoder self-attention (masked) | The sequence | The same sequence |
| Cross-attention | The decoder | The encoder |
That asymmetry has a practical consequence during generation, and it underpins the most important inference optimisation in language models.
Why the KV cache exists
When generating token by token, the keys and values of all previous tokens do not change — only the new token's query is new.
Without a cache, generating the 1,000th token means recomputing keys and values for all 999 previous tokens. Generation is quadratic in output length.
With a KV cache, each new token computes its own K and V, appends them to the stored ones, and attends against the cache. Generation becomes linear.
The cost is memory: two tensors per layer per token. The size is calculable:
bytes = 2 × layers × heads × head_dim × tokens × bytes_per_value
For a 32-layer model with 32 heads of dimension 128, at 16-bit, with 8,000 tokens: about 4.2GB for one request. Ten concurrent requests exceed most single GPUs before the weights are counted.
That arithmetic is why grouped-query attention exists: sharing key and value projections across groups of query heads shrinks the cache several-fold. With 32 query heads and 8 KV groups it is a quarter the size, with quality close to full multi-head. It is an architectural choice driven entirely by inference economics.
Multiple heads
Multi-head attention runs several of these computations in parallel with different projections, then concatenates and projects once more.
With a model dimension of 512 and 8 heads, each head works in 64 dimensions, so the total cost is roughly that of one 512-dimensional head — and the model gets 8 independent views of the sequence rather than 1.
Inspected in trained models, heads specialise: some track the previous token, some link verbs to subjects, some match brackets, some connect pronouns to referents, and some do very little. Pruning studies find many can be removed with minimal loss, which suggests the redundancy helps training rather than inference.
Questions people ask
Why the names query, key and value? By analogy with a database lookup. The analogy is genuinely useful; the mechanism is soft rather than exact.
Must Q and K be the same size? Yes — they are dotted together. V can differ in principle and is usually the same in practice.
Why divide by √dₖ? To stop the softmax saturating in high dimensions, which would kill the gradient.
Where are the parameters? In the three projection matrices per head, plus one output projection per layer. The attention formula itself has none.
Can K and V come from elsewhere? That is exactly cross-attention, and it is how decoders read encoders and how modalities are connected.
Why is the KV cache so large? Two tensors per layer per token, and it grows with context length — which is why long-context serving is a memory problem.
Recap in one screen
- Query is what a position seeks, key is how a position advertises itself, value is what it contributes.
- All three are learned linear projections; the attention formula itself has no parameters.
- Score with
QKᵀ, scale by√dₖ, softmax, then weight the values. - Masking by setting scores to negative infinity gives causal generation and correct padding.
- Keys and values never change during generation, which is what makes KV caching — and linear decoding — possible.