Articles
7indexed
49min read
#TitleTagsDateReadDepth
01
Backpropagation From First PrinciplesFEATURED

Derive the chain rule through a 2-layer net by hand before you ever touch autograd.

calculusgradients
2024-118 minDeep
02
Why Attention Is Just Weighted Memory

Strip away the jargon and understand self-attention as a differentiable key-value lookup.

transformersattention
2024-106 minMedium
03
The Vanishing Gradient Problem, Visualized

See exactly where gradients die in deep nets and why residual connections are the antidote.

deep learningtraining
2024-105 minMedium
04
Loss Surfaces Are Weirder Than You Think

A tour of saddle points, flat minima, and why SGD accidentally finds generalizable solutions.

optimizationtheory
2024-0910 minDeep
05
Normalization Layers: BN vs LN vs RMS

When does batch normalization break, and why did transformers abandon it for layer norm?

architecturetraining
2024-097 minMedium
06
Embeddings Are Just Learned Coordinates

Demystify word2vec and positional encodings by thinking in metric spaces.

embeddingsnlp
2024-084 minLight
07
The Bias–Variance Tradeoff Is Not a Tradeoff

Modern overparameterized models break the classical U-curve. Here's the new picture.

theorygeneralization
2024-089 minDeep