● The AI Primer, illustrated
How modern AI works, from the ground up.
54 lessons, in order, from the math notation to production AI systems. Every idea starts with an everyday picture, is worked through with numbers you can check by hand, drawn, and built in plain Python you can run.
Choose a path
Every lesson in reading order, one line per part. Pick who you are to see the lessons meant for you, in order.
Every lesson
Each one stands on the ones before it.
Before you begin
The notation every formula in this primer uses, decoded as short loops.
Part 1: how the model works inside
From a single neuron to a working transformer, and how models are trained and served.
- 01The big pictureWhat happens, end to end, when you send a promptFully illustrated
- 02Neural networksNeurons, activations, the forward pass, backprop by handFully illustrated
- 03OptimizersSGD, momentum, Adam/AdamW, learning-rate warmup and decayFully illustrated
- 04Training deep networksVanishing/exploding gradients, residuals, normalization, initializationFully illustrated
- 05AttentionQueries, keys, values, softmax, masking, multi-head, GQA, O(n²)Fully illustrated
- 06Positional informationWhy order must be added, sinusoids and RoPEFully illustrated
- 07The transformerThe block, a tiny GPT, parameter counts, mixture of expertsFully illustrated
- 08TokenizationBPE from scratch, byte-level tokens, why models miscount lettersFully illustrated
- 09Training stagesPretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAGFully illustrated
- 10Pretraining at scaleData curation and deduplication, parallelism across GPUs, mixed precisionFully illustrated
- 11Fine-tuning in practicePreparing data, forgetting old skills, merging modelsFully illustrated
- 12Reinforcement learningPolicy gradients from scratch, PPO, GRPO, reward hackingFully illustrated
- 13Reasoning modelsChain of thought, test-time compute, verifiers, learning to reason with RLFully illustrated
- 14Alignment and safetyConstitutional AI, red-teaming, sycophancy, refusalsFully illustrated
- 15The hardware underneathGPUs, the memory hierarchy, FLOPs vs. bandwidth, number formats
- 16InferencePrefill vs. decode, the KV cache, sampling, speculative decoding, memory mathFully illustrated
- 17Structured outputConstrained decoding: grammars and JSON schemas that guarantee valid output
- 18Long context and efficient architecturesSliding-window and sparse attention, state-space models, KV-cache compression
- 19Loss functionsCross-entropy, perplexity, MSE/MAE, contrastive losses
- 20MetricsPrecision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE
- 21Reading benchmarksWhat benchmarks measure, contamination, leaderboards and arenas
- 22Overfitting and regularizationOverfitting, early stopping, dropout, L1/L2, leakage
- 23Trees and boostingDecision trees, random forests, gradient boosting, and when they still win
- 24CNNs and RNNsHow convolutions see and recurrent nets remember, and why transformers won
- 25Looking inside the modelProbes, the logit lens, activation patching, superposition, sparse autoencoders
Embeddings, the centerpiece
Vectors that capture meaning, and the search systems built on them.
- 26Word embeddingsWhere embeddings came from, analogies, the "bank" problem
- 27SimilarityCosine vs. dot vs. distance, normalization, anisotropy, thresholds
- 28Training embedding modelsContrastive learning, hard negatives, CLIP
- 29Dimensions and compressionStorage math, Matryoshka truncation, int8 and binary quantization
- 30Vector indexesFlat, IVF, PQ and HNSW from scratch, recall vs. latency
- 31RetrievalBM25, hybrid search with RRF, rerankers, ColBERT, chunking
- 32Clustering and matchingk-means, density clustering, dedup, routing, semantic caching
- 33Embeddings in productionModel migrations, domain mismatch, measuring retrieval on its own
Generating images, audio and video
Autoencoders, GANs, diffusion, and the multimodal models that connect them to language.
- 34Autoencoders and VAEsSqueezing data into a code and back, and sampling new data from it
- 35GANsA forger against a detective: adversarial training, and why it is unstable
- 36Diffusion and flow matchingTurning noise into images one small step at a time
- 37Multimodal modelsImages, audio and video into a language model
Part 2: building systems people rely on
Agents, tools, retrieval, memory, evaluation, safety, cost and deployment.
- 38Talking to a modelThe message format, and what tool calling really is
- 39OrchestrationWorkflows vs. agents, and the named patterns
- 40The agent loopA production agent loop: budgets, loop detection, recovery
- 41ToolsTool design, validation, idempotency, approvals, least privilege
- 42Coding and computer-use agentsEdit, run, test, repeat; sandboxes; driving a screen
- 43Model Context ProtocolMCP on the wire, and its security risks
- 44Retrieval-augmented generationRAG end to end, with citations and access controlFully illustrated
- 45Context engineeringWhat goes in the window, compression, cache-friendly layout
- 46MemoryShort- and long-term memory, tenant isolation, forgetting
- 47PlanningPlan-and-execute, decomposition, reflection, compounding error
- 48EvaluationGolden sets, graders, LLM-as-judge calibration
- 49GuardrailsPrompt injection and privilege separation, PII, output checks
- 50Cost and latencyRouting, caching, batching, budgets, cost per successful task
- 51ObservabilityTraces, OpenTelemetry GenAI attributes, the improvement loop
- 52Safe deploymentShadow mode, graduated autonomy, canaries, kill switches, audit logs
- 53Why the hard ones failThe common failure modes, and the fix for each