Chapter 1 · Frame and baseline

Text to tensors: tokenization, embeddings, position

Use text to tensors: tokenization, embeddings, position to move the transformer from scratch production brief toward a defensible release.

65–90 min4 key conceptsReviewed 26 Aug 2026
01 · Production proposition

Domain autocomplete needs a model whose every component the team can explain.

This lesson isolates text to tensors: tokenization, embeddings, position as one decision inside that system. The people affected are product users and operators; the learning data must carry event time, availability time, ownership, and version; and the operating envelope is one-device educational build.

Decision

Choose whether and how to use tokenization at a declared prediction cutoff.

Metric

Measure decision utility alongside calibration, slice reliability, and system latency—not model score alone.

Failure consequence

Domain autocomplete needs a model whose every component the team can explain. An unsafe release must degrade to a named baseline or the last known-good version.

02 · Intuition & prerequisites

Build the mental model before the machinery.

The core move is to treat text to tensors: tokenization, embeddings, position as a contract between data, a computation, and an action. Version tokenizer, vocabulary, special tokens, architecture, weights, prompt template, and generation policy together. The implementation becomes easier to debug once you can state which inputs exist, which state is learned, what output means, and what must remain invariant after serialization.

01

tokenization

Define it in a hand-checkable form and name the prediction-time inputs.

02

embeddings

Connect it to the production metric and identify what it cannot guarantee.

03

RoPE

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

04

QKV projections

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

Bring forward

Neural networks, Sequence models, GPU training

03 · Formal treatment

Name every symbol. Check every shape.

Scaled dot-product attention is the central invariant for this lesson. The formula is useful only when its inputs match the production cutoff and its output maps to an action.

Formal treatment
Attention(Q,K,V)=softmax ⁣(QKdk+M)V\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}} + M\right)V

Scaled dot-product attention

Symbol, shape or unit contract
SymbolMeaning / shape / unit
Q, K, Vquery, key, value matrices
d_khead width
Mcausal or padding mask
Open derivation and numerical substitution

Start from the production quantity being optimized, substitute the observed values with their declared units, then isolate the model-controlled term. Preserve shape annotations at each step so broadcasting or aggregation cannot silently change the result.

  1. Write the named inputs: Q, K, V, d_k, M.
  2. Substitute one small, hand-checkable batch before vectorizing.
  3. Calculate an independent reference value and compare within a declared tolerance.
# equation → code contract
inputs = validate_shapes_and_units(batch)
value = compute_c14(inputs)
assert is_finite(value)
04 · Three views of the idea

Calculate it small. Shape it realistically. Break it on purpose.

HAND-CALCULATED TOY

A result you can reproduce on paper

Compute one attention head for four tokens and alter one key to watch a softmax row move.

  1. Write every input and unit.
  2. Substitute values into the scaled dot-product attention equation above.
  3. Compare the result to one simple baseline and explain the direction of the difference.
PRODUCTION-SHAPED

The same reasoning under real constraints

Train a small decoder on versioned support text and serve bounded, KV-cached completions.

The production record includes the data snapshot, transformation state, artifact identity, cutoff, score, decision, and the version of the policy that consumed it.

FAILURE / COUNTEREXAMPLE

The attractive result you should reject

Without the causal mask, each position sees its target and validation loss becomes impossibly good.

Diagnostic: replay the smallest failing slice from immutable inputs, then compare each boundary rather than retuning the model.

05 · Deterministic lab

Change one assumption and make the tradeoff visible.

This lab runs predefined TypeScript only. It never executes learner code. Use the slider, numeric input, reset, live text, or table—the computation is the same.

Exact computation

Attention temperature lab

Change score temperature and inspect the exact softmax distribution.

Primary1.61 bits
Secondary53.1%
DiagnosisUsable spread

Assumption: Fixed scores [1.2, 0.3, −0.4, 1.8].

Open nonvisual data table
ItemComputed stateInterpretation
token 11.20.2915
token 20.30.1185
token 3-0.40.0589
token 41.80.5311
06 · Production implications

Trace the complete operating path.

  1. 01

    Validate and version tokenization.

  2. 02

    Compute text to tensors: tokenization, embeddings, position from prediction-time-safe inputs.

  3. 03

    Persist model, feature, and configuration identities together.

  4. 04

    Serve or materialize behind explicit one-device educational build.

  5. 05

    Join telemetry to mature outcomes and retain a rollback path.

Observability

Join service health, input quality, prediction distributions, slice behavior, and mature outcomes by exact version.

Cost

Measure storage, preprocessing, compute, queueing, and human review under a representative arrival pattern.

Failure modes

Without the causal mask, each position sees its target and validation loss becomes impossibly good. Add a detector, owner, mitigation, and stop condition for this class of failure.

Alternatives

Compare a rule, a simpler statistical baseline, and a different system boundary before adding model complexity.

07 · Check understanding

Explain the contract, not just the vocabulary.

Browser-graded checkpointPass ≥ 80%
01Why divide logits by √d_k?
02What does the causal mask guarantee?
03What does KV caching remove?
08 · Apply in production

Decoder-only transformer from scratch

Specify tokenizer, attention blocks, causal training, unit tests, sampling, and KV-cache parity without a transformer module.

  • Shape/mask tests
  • Tiny-batch overfit
  • Held-out evaluation
  • Cached/uncached parity
Open assignment and rubric
09 · Sources & next depth

Read primary material with a purpose.

10 · Production resolution

Return to the opening failure.

Version tokenizer, vocabulary, special tokens, architecture, weights, prompt template, and generation policy together.

For this lesson, the release evidence is a hand-checked formal result, deterministic simulation output, a ≥80% checkpoint, the production rubric, and a named fallback. The course resolves when the system can produce tokenizer, attention, rope, residual stack, training loop, generation, and kv cache.