HOW AI WORKS - FRAME BOOK
VOL. 01·SEP 2026

FROM PREDICTOR TO ASSISTANT.

A PRE-TRAINED MODEL CONTINUES TEXT. FINE-TUNING AND HUMAN FEEDBACK SHIFT WHICH CONTINUATIONS IT PREFERS.

BASE MODELSFTRLHF

What is a black hole?

MORE QUESTIONS32%4%0%UNRELATED TEXT31%7%1%WEAK ANSWER23%45%17%HELPFUL ANSWER14%44%82%
E[r]=−0.93\mathbb{E}[r] = -0.93E[r]=0.10\mathbb{E}[r] = 0.10E[r]=0.92\mathbb{E}[r] = 0.92KL=—\mathrm{KL} = \text{—}KL=0\mathrm{KL} = 0KL=0.36\mathrm{KL} = 0.36
REWARD r(y)r(y)

SAME KNOWLEDGE. NEW HABITS.

MOST KNOWLEDGE COMES FROM PRE-TRAINING. POST-TRAINING MAINLY RESHAPES BEHAVIOUR: WHICH REPLIES BECOME LIKELY.

π∗(y∣x)∝πref(y∣x) e r(x,y)/β\pi^*(y \mid x) \propto \pi_{\text{ref}}(y \mid x)\, e^{\,r(x,y)/\beta}
πθ\pi_\theta
Policy (the model)
πref\pi_{\text{ref}}
Reference (SFT)
x, yx,\ y
Prompt, response
rr
Reward
β\beta
KL strength

01BASE MODEL

y∼πbase( ⋅∣x)y \sim \pi_{\text{base}}(\,\cdot \mid x)

It continues the page instead of answering.

PRE-TRAINED MODEL · NEXT-TOKEN ONLYWhat is a black hole?What is a white hole?What is a wormhole?What is dark matter?

02SUPERVISED FINE-TUNING

LSFT=−∑tlog⁡πθ(yt∣x,y<t)\mathcal{L}_{\text{SFT}} = -\sum_{t} \log \pi_\theta(y_t \mid x, y_{<t})

Imitate example answers written by people.

USERWhat is a black hole?ASSISTANTA black hole is a region of spacetime wheregravity is so strong that nothing, not evenlight, can escape its pull.LOSS ON RESPONSE TOKENS ONLY

03HUMAN PREFERENCES

yw≻yly_w \succ y_l

People compare two answers and pick one.

What is a black hole?
AA region of spacetime where gravity is so strong that not even light can escape.
BA hole in space that sucks in everything around it, forever.
✓ PREFERRED

04REWARD MODEL

P(yw≻yl)=σ(rϕ(x,yw)−rϕ(x,yl))P(y_w \succ y_l) = \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)

Learn a score that agrees with people.

P(yw≻yl)P(y_w \succ y_l)0.51Δr\Delta r00.92

05PPO WITH KL

max⁡θ E[rϕ]−β KL(πθ ∥ πref)\max_{\theta}\ \mathbb{E}\big[r_\phi\big] - \beta\,\mathrm{KL}\big(\pi_\theta \,\|\, \pi_{\text{ref}}\big)

Chase reward, but stay near the original.

POLICY πθ\pi_\thetaREWARD MODEL rϕr_\phiPPO UPDATEπref\pi_{\text{ref}}πθ\pi_\theta← KLREWARD →

06ALTERNATIVES

Newer recipes replace parts of the pipeline.

DPOLearns from preference pairs directly: no reward model, no RL loop.RLAIFAn AI model, not a person, supplies the preference labels.CONSTITUTIONAI feedback guided by a written list of principles.

Fine-tuning & RLHF

From Predictor to Assistant

PRE-TRAINING TEACHES A MODEL TO CONTINUE TEXT. FINE-TUNING ON WRITTEN EXAMPLES TEACHES IT TO ANSWER; A REWARD MODEL LEARNED FROM HUMAN COMPARISONS THEN STEERS IT TOWARD REPLIES PEOPLE PREFER, WHILE A KL PENALTY KEEPS IT CLOSE TO WHERE IT STARTED.

ASSISTANTRLHF~104–105 PROMPTSREWARD MODEL~105 COMPARISONSSFT~104 DEMOSPRE-TRAINING~1013 TOKENSKLNOT TO SCALE: THE REAL CAP IS FAR THINNER

A thick pre-training slab (about 10 trillion tokens) carries the model's knowledge. Three thin layers sit on top — supervised fine-tuning (about ten thousand demonstrations), a reward model (about a hundred thousand comparisons) and RLHF (tens of thousands of prompts) — tied to the SFT layer by a KL leash, leading up to the assistant. Not to scale: the real cap is far thinner.

SFT
LSFT=−∑tlog⁡πθ(yt∣x,y<t)\mathcal{L}_{\text{SFT}} = -\sum_{t} \log \pi_\theta(y_t \mid x, y_{<t})
REWARD MODEL
LRM=−log⁡σ(rϕ(x,yw)−rϕ(x,yl))\mathcal{L}_{\text{RM}} = -\log \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)
RL (PPO)
max⁡θ E[rϕ(x,y)]−β KL(πθ ∥ πref)\max_{\theta}\ \mathbb{E}\big[r_\phi(x, y)\big] - \beta\,\mathrm{KL}\big(\pi_\theta \,\|\, \pi_{\text{ref}}\big)
DPO
LDPO=−log⁡σ(β Δw−β Δl)\mathcal{L}_{\text{DPO}} = -\log \sigma\big(\beta\,\Delta_w - \beta\,\Delta_l\big)
Δ=log⁡πθ(y∣x)πref(y∣x)\Delta = \log \dfrac{\pi_\theta(y \mid x)}{\pi_{\text{ref}}(y \mid x)}
ALSO
RLAIF: preference labels from an AI model
Constitutional AI: feedback guided by written principles
IN PRACTICE
InstructGPT 1.3B was preferred over GPT-3 175B (2022).