HOW AI WORKS - FRAME BOOK
VOL. 01·SEP 2026

THINKING IN TOKENS.

A MODEL CANNOT PAUSE TO THINK. IT CAN ONLY WRITE — AND EVERY TOKEN OF SCRATCH WORK IS ONE MORE PASS THROUGH THE NETWORK.

17×24=?17 \times 24 = ?24=20+424 = 20 + 417×20=34017 \times 20 = 34017×4=6817 \times 4 = 68340+68=408340 + 68 = 40817×4=5817 \times 4 = 58340+58=398340 + 58 = 39824=25−124 = 25 - 117×25=42517 \times 25 = 425425−17=408425 - 17 = 408425−7=418425 - 7 = 418
398×1
408×1×2×3MAJORITY
418×1
418NO STEPS

ROUTES

  • A · B24=20+4→17×20=340→17×4=68→40824 = 20 + 4 \to 17 \times 20 = 340 \to 17 \times 4 = 68 \to 408
  • C⋯→17×4=58→398\dots \to 17 \times 4 = 58 \to 398
  • E24=25−1→17×25=425→40824 = 25 - 1 \to 17 \times 25 = 425 \to 408
  • D⋯→425−7=418\dots \to 425 - 7 = 418
  • DIRECT418418

MANY PATHS, ONE ANSWER.

SAMPLE SEVERAL CHAINS OF THOUGHT. SOME SLIP, SOME TAKE ANOTHER ROUTE. THE ANSWER MOST CHAINS REACH WINS THE VOTE.

VOTE

408 ×3 · 398 ×1 · 418 ×1

DIRECT ANSWER: 418

01CHAIN OF THOUGHT

QWhat is 17 × 24?
DIRECTSTEP BY STEPA: 41817 × 20 = 34017 × 4 = 68340 + 68 = 408A: 4081 TOKEN≈\approx 25 TOKENS
q  →  c1 c2⋯ck  →  aq \;\rightarrow\; c_1\, c_2 \cdots c_k \;\rightarrow\; a

Write the steps first, then the answer.

02ONE PASS PER TOKEN

LL1 TOKEN25 TOKENS≈\approx 14 GFLOP≈\approx 350 GFLOP7B MODEL
C≈2N⋅ntokensC \approx 2N \cdot n_{\text{tokens}}
N=7×109⇒14 GFLOP / tokenN = 7 \times 10^{9} \Rightarrow 14\ \text{GFLOP / token}

Same cost per token. More tokens, more compute.

03TEST-TIME COMPUTE

ACCURACY050%100%101010210^{2}10310^{3}10410^{4}10510^{5}THINKING TOKENS (LOG)ILLUSTRATIVE
accuracy↑ as log⁡nthink↑\text{accuracy} \uparrow \ \text{as}\ \log n_{\text{think}} \uparrow

Longer thinking helps, until it plateaus.

04SELF-CONSISTENCY

K = 5 · T = 0.7398×1408×1×2×3418×1MAJORITY
a^=arg⁡max⁡a∑k=1K1[ak=a]\hat a = \arg\max_{a} \sum_{k=1}^{K} \mathbb{1}[a_k = a]

Sample several chains. Keep the majority answer.

05TRAINED TO THINK

STANDARDREASONINGTHINKINGANSWERANSWERPROBLEMTHINKANSWERCHECKREWARD
r=1[ a=a∗ ]r = \mathbb{1}[\,a = a^{*}\,]

Rewarded for correct answers, it thinks longer.

06LIMITS

17 × 20 = 34017 × 4 = 58340 + 58 = 398A: 398P(ALL RIGHT)12345678910STEPS kk0.60
P(all k steps right)=pkP(\text{all } k \text{ steps right}) = p^{k}
0.9510≈0.600.95^{10} \approx 0.60

Plausible steps can still be wrong.

Reasoning

Thinking in Tokens

A LANGUAGE MODEL REASONS BY WRITING. EACH INTERMEDIATE TOKEN HOLDS SCRATCH WORK AND BUYS ONE MORE FIXED SLICE OF COMPUTE; SAMPLING SEVERAL CHAINS LETS THE ANSWERS VOTE. MORE TOKENS, MORE THOUGHT, MORE COST.

34068408
QUESTIONqqTHOUGHTSc1⋯ckc_1 \cdots c_kANSWERaa17×24=?17 \times 24 = ?408
Three stacked plates. On the top plate the question q. A beam carries it down to the start of a snake of 27 token beads on the middle plate, the thoughts c1 to ck, three rows of nine; the last bead of each row holds a result: 340, 68, 408. A second beam carries the last bead down to the answer plate, where rings spread around the answer 408.
QUESTIONqq
THOUGHTSc1 c2⋯ckc_1\, c_2 \cdots c_k

EACH TOKEN = ONE FORWARD PASS

ANSWERaa
COMPUTEC≈2N⋅ntokensC \approx 2N \cdot n_{\text{tokens}}
VOTEa^=majority⁡(a1,…,aK)\hat a = \operatorname{majority}(a_1, \dots, a_K)
  • TOKEN
  • MAJORITY ANSWER
  • SLIPPED STEP