TERMSACROSSTHE BOOK.
A BACK COVER THAT DOUBLES AS AN INDEX: 77 TERMS, FOUR THEMES, EVERY PAGE ONE CLICK AWAY.
SYNTHESIS GLYPH
24 PAGES ON ONE RING. EACH CHORD JOINS TWO CHAPTERS THAT SHARE AN IDEA; PULSES RUN IN READING ORDER.
THE LANGUAGE-MODEL LOOP IN ONE LINE:
pages = 24terms = 77chords = 31loop = 8 s
HOVER OR FOCUS A PAGE NUMBERTO SEE ITS LINKS.
TAP A PAGE NUMBER TO OPEN IT.
Glossary
Terms Across the Book
77 KEY TERMS FROM PAGES 002 – 023, GROUPED BY THE BOOK'S FOUR THEMES. EACH ONE HAS A SHORT DEFINITION AND LINKS BACK TO THE PAGES WHERE THE IDEA IS DRAWN.
The book's four themes as a stack of layers — 04 generation: 27 terms, 03 transformation: 20 terms, 01 attention: 11 terms, 02 representation: 19 terms.
01ATTENTIONHOW TOKENS SHARE INFORMATION AND CONTEXT.11 TERMS
- Attention
- Each token builds its new vector as a weighted mix of other tokens' values.
- Pages: 005001007
- Query, key, value
- Three vectors per token: what it looks for, what it offers, and what it passes on.
- Pages: 005014020
- Attention head
- One attention computation; several run in parallel, each learning its own pattern.
- Pages: 005001
- Causal mask
- Hides later tokens, so each position can only attend to itself and earlier ones.
- Pages: 014011
- Positional encoding
- A signal that tells the model where each token sits in the sequence.
- Pages: 005019
- KV cache
- Keys and values of earlier tokens, saved so each new token needs less computation.
- Pages: 014
- Grouped-query attention (GQA)
- Several query heads share one set of keys and values, so the KV cache needs less memory.
- Pages: 014
- Prefill and decode
- The two phases of inference: read the whole prompt in parallel, then write one token per step.
- Pages: 014
- Cross-attention
- Attention where one sequence queries another, e.g. image features reading a text prompt.
- Pages: 020
02REPRESENTATIONHOW TEXT, IMAGES AND MEANING BECOME VECTORS.19 TERMS
- Vocabulary
- Every token a model knows, each with an integer ID (often 50,000 or more).
- Pages: 004006011
- Latent space
- The space of a model's internal vectors, where related things lie close together.
- Pages: 022001009020
- Retrieval-augmented generation (RAG)
- Search a document store, often by embedding similarity, and add the best passages to the prompt.
- Pages: 017021
- Vision Transformer (ViT)
- A transformer that reads an image as a sequence of patch embeddings, just as it reads text.
- Pages: 019
- CLIP
- An image encoder and a text encoder trained so each picture lands near its caption in one shared space.
- Pages: 019020
- Residual stream
- Each token's running vector; every layer reads it and adds its output back.
- Pages: 007022
- Superposition
- Hypothesis that models store more features than dimensions as nearly orthogonal directions.
- Pages: 022
- Sparse autoencoder (SAE)
- Rewrites an activation as a sum of a few features from a large learned dictionary.
- Pages: 022
03TRANSFORMATIONHOW LAYERS COMPUTE AND HOW MODELS LEARN.20 TERMS
- Transformer layer
- Attention plus a feed-forward network, joined by residual connections and normalization.
- Pages: 007010
- Feed-forward network (MLP)
- A small two-layer network applied to each token separately inside every layer.
- Pages: 007
- Cross-entropy
- The next-token loss: minus the log of the probability given to the true token.
- Pages: 011021
- Scaling laws
- Loss falls predictably, roughly as a power law, as parameters, data and compute grow.
- Pages: 012
- Compute-optimal training
- For a fixed compute budget, grow parameters and data together: about 20 training tokens per parameter.
- Pages: 012
- Emergent abilities
- Skills that seem to switch on suddenly at scale; partly an effect of all-or-nothing scoring.
- Pages: 012
- Reward model
- A model trained on people's choices between two answers, used to give any answer a score.
- Pages: 015
- KL penalty
- A cost for drifting from the reference model, so RLHF gains reward without forgetting how to write.
- Pages: 015
04GENERATIONHOW OUTPUTS ARE PRODUCED, USED AND CHECKED.27 TERMS
- Next-token prediction
- The core task: assign a probability to every possible next token.
- Pages: 011006010
- Greedy decoding
- Always picking the single most likely next token, so the same prompt gives the same text.
- Pages: 013006
- Top-k
- Keep only the k best candidates: the k likeliest tokens when sampling, or the k closest chunks in a search.
- Pages: 013017
- Top-p (nucleus) sampling
- Sampling only from the smallest set of tokens whose probabilities add up to p.
- Pages: 013
- Test-time compute
- Spending more computation while answering, by writing longer reasoning or sampling more attempts.
- Pages: 016
- Self-consistency
- Sampling several chains of reasoning and keeping the answer most of them reach.
- Pages: 016
- Model Context Protocol (MCP)
- An open standard for plugging tools and data sources into AI applications.
- Pages: 018
- Diffusion model
- Generates an image by removing noise step by step, starting from pure noise.
- Pages: 009020
- Classifier-free guidance
- Pushes each denoising step toward the prompt by comparing predictions with and without it.
- Pages: 020
- Knowledge cutoff
- The date the training data ends; the model has seen nothing that happened after it.
- Pages: 021
- LLM-as-a-judge
- Using a model to compare or grade answers when there is no single correct one.
- Pages: 023
LATENT
How AI Works
A MOTION DESIGN STYLE-FRAME BOOK
COLOPHON
- EDITION
- VOL. 01 · SEP 2026
- PAGES
- 001 — 024
- CREDIT
- Compiled & developed by Vu Dinh
- TYPE
- IBM Plex Mono · Inter Tight
