DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · 2025
Long chain-of-thought behaviour emerges from RL with rule-based rewards alone — no supervised reasoning traces required. The cold-start SFT stage exists for readability, not capability.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · 2025
Long chain-of-thought behaviour emerges from RL with rule-based rewards alone — no supervised reasoning traces required. The cold-start SFT stage exists for readability, not capability.
What it does
DeepSeek-R1-Zero applies large-scale reinforcement learning directly to a base model using rule-based accuracy and format rewards, with GRPO in place of PPO (no learned value model). Extended reasoning chains, self-verification, and backtracking emerge without any supervised reasoning data.
R1 then adds a small cold-start SFT stage plus a second RL round, mainly to fix language mixing and readability problems in R1-Zero’s output — the reasoning capability itself came from RL.