Skip to content
SHASHWAT // SYSTEM ARCHIVE
∞
SYSTEM.ARTICLE

Attention Mechanisms from First Principles: Why Q, K, V Even Exist

avatarShashwat Sharma
9 min read

Your favorite models run on a 2017 paper that most people who use it every day have never derived from scratch. Query, Key, and Value aren't arbitrary names bolted onto a matrix multiply — each one exists to solve a specific failure of earlier sequence models. This post builds attention up from that failure instead of handing you the formula and moving on.


The Problem With Naive Sequence Modeling

Before attention, the standard way to model a sequence was recurrence. An RNN reads a sentence one token at a time and carries a hidden state forward. Token 50 only "sees" the sentence through whatever survived being compressed and re-compressed 49 times.

This has two concrete problems.

The first is dependency decay. If token 50 needs information from token 2, that information has to pass through 48 intermediate compressions. Each one is lossy. By the time it reaches token 50, the signal is faint. This is the vanishing gradient problem, and LSTMs and GRUs were built specifically to slow it down — they didn't remove it.

The second is sequential computation. An RNN at step 50 needs the hidden state from step 49, which needs step 48, and so on. You cannot parallelize this across the sequence. On modern hardware, which is built to do enormous numbers of operations at once, this is a direct waste of the hardware you're paying for.

What you actually want is: every token should be able to look directly at every other token, in one step, with no decay and no forced order. That's O(n²) pairs of tokens for a sequence of length n — expensive, but each pair is now a direct connection instead of a chain of 48 lossy hops.

Attention is the mechanism that computes those O(n²) direct connections. The question is how to compute "how much should token A care about token B" for every pair, cheaply and learnably.

Why You Need Both a Query and a Key

Start with the simplest possible idea: represent every token as a single vector, and measure relevance between two tokens with a dot product between their vectors.

# naive: one vector per token, used for everything
relevance = token_vector[A] @ token_vector[B]

This has a subtle but fatal flaw. The vector for token A is being asked to do two jobs at once: represent what token A means, and represent what token A is looking for. Those aren't the same thing. A pronoun like "it" needs to look for a noun earlier in the sentence — its search criteria are different from its own meaning.

Splitting the single vector into two separate learned projections fixes this. The Query is what a token uses when it's searching — "what am I looking for?" The Key is what a token exposes when it's being searched against — "what do I offer, if someone is looking?"

Q = token_embedding @ W_query   # "what I'm looking for"
K = token_embedding @ W_key     # "what I offer to searchers"

Now the pronoun "it" can have a query vector that looks like "I need a singular noun, recently mentioned." A candidate noun earlier in the sentence has a key vector that looks like "I am a singular noun." These two vectors can point in a similar direction in the learned space even though the tokens themselves ("it" and "cat", say) have nothing in common as raw embeddings.

This is the first design decision: query and key are separate learned projections of the same input, because searching and being found are different roles.

The Dot Product: Selective Attention Through Q·K^T

Once every token has a query and every token has a key, computing relevance between all pairs is one matrix multiply.

# Q: [seq_len, d_k]   K: [seq_len, d_k]
scores = Q @ K.T   # [seq_len, seq_len]

Entry scores[i][j] is the dot product between token i's query and token j's key — a single number saying how strongly token i's search matches what token j offers.

The dot product is doing real work here, not just serving as a convenient similarity function. Two vectors that point in a similar direction and have large magnitude produce a large dot product. Two vectors that point in unrelated directions produce a dot product near zero. This is exactly the "selective" part of attention: most token pairs in a real sentence are irrelevant to each other, and the dot product naturally suppresses them toward zero without any explicit masking.

Before this can be used as a weighting, two more steps happen. First, scale by 1/sqrt(d_k). As the dimension of the query and key vectors grows, dot products grow with it — larger vectors, larger sums — and if the scores get too large, softmax saturates and gradients vanish. Dividing by the square root of the dimension keeps the scale roughly constant regardless of how wide the model is.

scaled_scores = scores / (d_k ** 0.5)
attention_weights = softmax(scaled_scores, axis=-1)

Second, softmax turns the raw scores into a proper probability distribution over "which other tokens should I attend to," per token. Every row sums to 1. This is what makes the mechanism selective rather than additive noise — a token can spend almost all of its attention on one or two other tokens if that's what the scores support, or spread it thin if nothing stands out.

Why Values Exist: The Information Bottleneck

At this point you have attention_weights, a [seq_len, seq_len] matrix telling you how much each token should attend to every other token. The obvious next move is to use these weights to combine something — but combine what?

You might guess: combine the keys. It's tempting because they're already computed. This is where the third projection, Value, earns its place.

Key and Query are optimized for one job only: producing a good matching score. Their job is comparison, not content. If you forced the same vectors to also carry the full semantic payload that gets passed forward, you'd be asking one representation to be good at two different things — sharp enough to match precisely, and rich enough to carry everything downstream layers need. Those goals pull in different directions during training.

The Value projection separates them. It's a third learned view of the same token, optimized purely to answer: "if you decide to attend to me, what should you actually receive?"

V = token_embedding @ W_value   # "what I pass along if attended to"
output = attention_weights @ V  # weighted mixture of values

This is the information bottleneck in the literal sense. The Key/Query pair decides how much information flows from token j to token i. The Value decides what that information actually is. Separating the routing decision from the payload lets both get optimized independently — the model can learn a matching function that's excellent at finding the right tokens without that same function also having to compress everything worth knowing about a token into one vector.

Empirically, this split is why models still hold up under mixed input distributions: the same architecture proposed for translation between languages generalized directly to image patches and video frames, where "relevance" and "content" have completely different structure, precisely because query/key and value were never coupled to a single notion of what a token's vector has to represent.

Putting It Together: The Full Equation

All three pieces combine into the equation everyone quotes and few derive:

def attention(Q, K, V, d_k):
    scores = Q @ K.T / (d_k ** 0.5)
    weights = softmax(scores, axis=-1)
    return weights @ V

Read as a sentence: for every token, compare what it's looking for (Q) against what every other token offers (K), turn those comparisons into a normalized weighting, and use that weighting to pull a mixture of what those tokens actually carry (V).

Three learned projections, one dot product, one softmax, one weighted sum. That's the entire mechanism. Everything else in a Transformer — multiple heads running this in parallel, residual connections, layer norm, positional encodings — is scaffolding around this one operation, not a replacement for it.

Why This Design Won

The Transformer paper that introduced this mechanism in 2017 was built and evaluated for machine translation. Nothing in the architecture is specific to language. Q, K, and V are just three learned linear projections of whatever tokens you feed in — words, image patches, audio frames, user-item interactions. That generality is why the same attention block was adopted for image processing, video processing, and recommender systems with comparatively small modifications, mostly around how the input gets tokenized in the first place.

The other reason it won is the parallelism from the first section. Q @ K.T for an entire sequence is one large matrix multiply — the exact kind of operation GPUs are built to do fast. Compare that to an RNN, where step 50 cannot start until step 49 finishes. Attention trades O(n²) pairwise comparisons for full parallelism across the sequence, and on modern hardware that trade is almost always worth it until sequences get very long.

Conclusion

Query, Key, and Value are not three arbitrary names for three matrices that happen to work. They're the answer to three separate design questions: how does a token search, how does a token get found, and what does a token actually hand over once it's found. Splitting these roles is what lets the same dot-product-and-softmax mechanism generalize from translating sentences to describing images without changing its shape.

Next time you see softmax(QK^T / sqrt(d_k))V, you should be able to read it as a sentence, not memorize it as a formula.