← Back to the site

Reading list / ongoing

What I've
been reading.

Papers and technical write-ups, each with a one-line takeaway and my fuller notes.

04 entries

Most recent first

  1. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    DeepSeek-AI · 2025

    Long chain-of-thought behaviour emerges from RL with rule-based rewards alone — no supervised reasoning traces required. The cold-start SFT stage exists for readability, not capability.

    RL post-training · Reasoning · GRPO

    Read the paper ↗

    What it does

    DeepSeek-R1-Zero applies large-scale reinforcement learning directly to a base model using rule-based accuracy and format rewards, with GRPO in place of PPO (no learned value model). Extended reasoning chains, self-verification, and backtracking emerge without any supervised reasoning data.

    R1 then adds a small cold-start SFT stage plus a second RL round, mainly to fix language mixing and readability problems in R1-Zero’s output — the reasoning capability itself came from RL.

  2. FP8 Formats for Deep Learning

    Micikevicius et al. · 2022

    Two FP8 encodings rather than one: E4M3 for forward activations and weights where precision matters, E5M2 for gradients where dynamic range matters. The split is the whole point.

    FP8 · Numerics · Training

    Read the paper ↗

    What it does

    Proposes a two-format FP8 specification. E4M3 (4 exponent, 3 mantissa bits) buys precision at the cost of range and suits weights and activations; E5M2 matches FP16’s exponent range and suits gradients, whose distributions are wide and long-tailed.

    The paper shows training and inference results matching FP16/BF16 baselines across a range of model sizes, provided per-tensor scaling is used to keep values inside the representable range.

  3. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

    Dao, Fu, Ermon, Rudra, Ré · 2022

    Attention was never compute-bound — it was bound by HBM traffic. Tiling the computation so the softmax never materialises the full N×N matrix makes it exact and faster at once.

    Attention · GPU kernels · Memory

    Read the paper ↗

    What it does

    Standard attention writes the full N×N score matrix to high-bandwidth memory, reads it back for the softmax, and writes it again for the value multiply. On modern GPUs that memory traffic — not the matrix multiplies — is the bottleneck.

    FlashAttention tiles the computation so that a block of queries, keys, and values is loaded into SRAM once, and the softmax is computed incrementally using the online-softmax trick (running max and running normaliser). The N×N matrix is never materialised. The result is numerically exact, not an approximation, and both faster and dramatically lower in memory.

  4. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

    Frantar, Ashkboos, Hoefler, Alistarh · 2022

    One-shot weight quantization to 3–4 bits with negligible perplexity loss, by quantizing column by column and immediately correcting the remaining weights against the error just introduced.

    Quantization · Inference · Compression

    Read the paper ↗

    What it does

    GPTQ frames post-training quantization as layer-wise reconstruction: pick quantized weights that minimise the output error on a small calibration set. It solves this greedily — quantize one column, then update all not-yet-quantized columns to absorb the error, using a Hessian inverse computed once via Cholesky decomposition rather than recomputed per step.

    The practical payoff is 3–4 bit weights on very large models in a single pass, with no retraining, and perplexity close to FP16.