Articles
7indexed
49min read
| # | Title | Tags | Date | Read | Depth | |
|---|---|---|---|---|---|---|
| 01 | Backpropagation From First PrinciplesFEATURED Derive the chain rule through a 2-layer net by hand before you ever touch autograd. | calculusgradients | 2024-11 | 8 min | Deep | |
| 02 | Why Attention Is Just Weighted Memory Strip away the jargon and understand self-attention as a differentiable key-value lookup. | transformersattention | 2024-10 | 6 min | Medium | |
| 03 | The Vanishing Gradient Problem, Visualized See exactly where gradients die in deep nets and why residual connections are the antidote. | deep learningtraining | 2024-10 | 5 min | Medium | |
| 04 | Loss Surfaces Are Weirder Than You Think A tour of saddle points, flat minima, and why SGD accidentally finds generalizable solutions. | optimizationtheory | 2024-09 | 10 min | Deep | |
| 05 | Normalization Layers: BN vs LN vs RMS When does batch normalization break, and why did transformers abandon it for layer norm? | architecturetraining | 2024-09 | 7 min | Medium | |
| 06 | Embeddings Are Just Learned Coordinates Demystify word2vec and positional encodings by thinking in metric spaces. | embeddingsnlp | 2024-08 | 4 min | Light | |
| 07 | The Bias–Variance Tradeoff Is Not a Tradeoff Modern overparameterized models break the classical U-curve. Here's the new picture. | theorygeneralization | 2024-08 | 9 min | Deep |