rumblr Work in progressWIP

● The AI Primer, illustrated

How modern AI works, from the ground up.

54 lessons, in order, from the math notation to production AI systems. Every idea starts with an everyday picture, is worked through with numbers you can check by hand, drawn, and built in plain Python you can run.

Choose a path

Every lesson in reading order, one line per part. Pick who you are to see the lessons meant for you, in order.

Fully illustrated

Every lesson

Each one stands on the ones before it.

Before you begin

The notation every formula in this primer uses, decoded as short loops.

  1. 00Math notation, from zeroEvery symbol in an ML formula, as a short loopFully illustrated

Part 1: how the model works inside

From a single neuron to a working transformer, and how models are trained and served.

  1. 01The big pictureWhat happens, end to end, when you send a promptFully illustrated
  2. 02Neural networksNeurons, activations, the forward pass, backprop by handFully illustrated
  3. 03OptimizersSGD, momentum, Adam/AdamW, learning-rate warmup and decayFully illustrated
  4. 04Training deep networksVanishing/exploding gradients, residuals, normalization, initializationFully illustrated
  5. 05AttentionQueries, keys, values, softmax, masking, multi-head, GQA, O(n²)Fully illustrated
  6. 06Positional informationWhy order must be added, sinusoids and RoPEFully illustrated
  7. 07The transformerThe block, a tiny GPT, parameter counts, mixture of expertsFully illustrated
  8. 08TokenizationBPE from scratch, byte-level tokens, why models miscount lettersFully illustrated
  9. 09Training stagesPretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAGFully illustrated
  10. 10Pretraining at scaleData curation and deduplication, parallelism across GPUs, mixed precisionFully illustrated
  11. 11Fine-tuning in practicePreparing data, forgetting old skills, merging modelsFully illustrated
  12. 12Reinforcement learningPolicy gradients from scratch, PPO, GRPO, reward hackingFully illustrated
  13. 13Reasoning modelsChain of thought, test-time compute, verifiers, learning to reason with RLFully illustrated
  14. 14Alignment and safetyConstitutional AI, red-teaming, sycophancy, refusalsFully illustrated
  15. 15The hardware underneathGPUs, the memory hierarchy, FLOPs vs. bandwidth, number formats
  16. 16InferencePrefill vs. decode, the KV cache, sampling, speculative decoding, memory mathFully illustrated
  17. 17Structured outputConstrained decoding: grammars and JSON schemas that guarantee valid output
  18. 18Long context and efficient architecturesSliding-window and sparse attention, state-space models, KV-cache compression
  19. 19Loss functionsCross-entropy, perplexity, MSE/MAE, contrastive losses
  20. 20MetricsPrecision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE
  21. 21Reading benchmarksWhat benchmarks measure, contamination, leaderboards and arenas
  22. 22Overfitting and regularizationOverfitting, early stopping, dropout, L1/L2, leakage
  23. 23Trees and boostingDecision trees, random forests, gradient boosting, and when they still win
  24. 24CNNs and RNNsHow convolutions see and recurrent nets remember, and why transformers won
  25. 25Looking inside the modelProbes, the logit lens, activation patching, superposition, sparse autoencoders

Embeddings, the centerpiece

Vectors that capture meaning, and the search systems built on them.

  1. 26Word embeddingsWhere embeddings came from, analogies, the "bank" problem
  2. 27SimilarityCosine vs. dot vs. distance, normalization, anisotropy, thresholds
  3. 28Training embedding modelsContrastive learning, hard negatives, CLIP
  4. 29Dimensions and compressionStorage math, Matryoshka truncation, int8 and binary quantization
  5. 30Vector indexesFlat, IVF, PQ and HNSW from scratch, recall vs. latency
  6. 31RetrievalBM25, hybrid search with RRF, rerankers, ColBERT, chunking
  7. 32Clustering and matchingk-means, density clustering, dedup, routing, semantic caching
  8. 33Embeddings in productionModel migrations, domain mismatch, measuring retrieval on its own

Generating images, audio and video

Autoencoders, GANs, diffusion, and the multimodal models that connect them to language.

  1. 34Autoencoders and VAEsSqueezing data into a code and back, and sampling new data from it
  2. 35GANsA forger against a detective: adversarial training, and why it is unstable
  3. 36Diffusion and flow matchingTurning noise into images one small step at a time
  4. 37Multimodal modelsImages, audio and video into a language model

Part 2: building systems people rely on

Agents, tools, retrieval, memory, evaluation, safety, cost and deployment.

  1. 38Talking to a modelThe message format, and what tool calling really is
  2. 39OrchestrationWorkflows vs. agents, and the named patterns
  3. 40The agent loopA production agent loop: budgets, loop detection, recovery
  4. 41ToolsTool design, validation, idempotency, approvals, least privilege
  5. 42Coding and computer-use agentsEdit, run, test, repeat; sandboxes; driving a screen
  6. 43Model Context ProtocolMCP on the wire, and its security risks
  7. 44Retrieval-augmented generationRAG end to end, with citations and access controlFully illustrated
  8. 45Context engineeringWhat goes in the window, compression, cache-friendly layout
  9. 46MemoryShort- and long-term memory, tenant isolation, forgetting
  10. 47PlanningPlan-and-execute, decomposition, reflection, compounding error
  11. 48EvaluationGolden sets, graders, LLM-as-judge calibration
  12. 49GuardrailsPrompt injection and privilege separation, PII, output checks
  13. 50Cost and latencyRouting, caching, batching, budgets, cost per successful task
  14. 51ObservabilityTraces, OpenTelemetry GenAI attributes, the improvement loop
  15. 52Safe deploymentShadow mode, graduated autonomy, canaries, kill switches, audit logs
  16. 53Why the hard ones failThe common failure modes, and the fix for each