Just now, the DeepSeek-R1 paper appeared as a cover article in the authoritative scientific journal Nature, with DeepSeek founder and CEO Liang Wenfeng serving as the corresponding author.

The research team hypothesized that human-defined reasoning patterns may limit the exploration of AI models, whereas unrestricted reinforcement learning (RL) training could better stimulate the emergence of novel reasoning abilities in large language models (LLMs).
Through experiments, they demonstrated that an LLM’s reasoning capabilities can be enhanced purely via RL, reducing the need for human input to boost performance, and outperforming traditionally trained LLMs in tasks such as mathematics, programming competitions, and graduate-level STEM questions.
Since its release, DeepSeek-R1 has received widespread acclaim from developers worldwide, and as of this writing, it has **91.1k stars on GitHub.
In a contemporaneously published perspective article, **Assistant Professor Daphne Ippolito from Carnegie Mellon University** and her PhD student Zhang Yiming (now an LLM safety and alignment researcher at Anthropic) commented:
“DeepSeek-R1 has evolved from a powerful but opaque problem solver into a system capable of human-like dialogue. This journey reflects the need for AI systems that not only solve problems accurately but also become tools that humans can understand, trust, and collaborate with meaningfully.”
Moreover, in its editorial, Nature praised the work:
“DeepSeek-R1 is the first mainstream LLM published after peer review — a welcome step toward transparency.”
They pointedly noted that peer-reviewed publication helps clarify how LLMs work and assess whether they are **“genuine”** (whether they do what they purport to do).

The Science Behind DeepSeek-R1
Human-defined reasoning patterns may restrict model exploration, while unrestricted RL training can better encourage the emergence of novel reasoning abilities in LLMs.
Achieving general reasoning in machines akin to humans has long been a core challenge in AI. While methods like Chain-of-Thought (CoT) improve LLM reasoning, they rely heavily on human annotation, which is not scalable and may impose human cognitive biases that prevent models from exploring superior, non-human reasoning pathways.
The significance of DeepSeek-R1 lies in its demonstration that pure RL alone can stimulate LLM reasoning abilities, without relying on annotated reasoning data. Unlike earlier methods based on prompting or supervised learning, the team proposed a new paradigm: within an RL framework, minimize dependence on human annotations while exploring the model’s potential to develop reasoning capabilities through self-evolution.
Prompting vs. Supervised Learning vs. RL
As Ippolito et al. analogize:
RL works like a human learning to play a video game: the player experiments in the game world, discovering which actions earn rewards — e.g., “collect coins” increases the score, while “hit an enemy” resets it.
Prompting methods are akin to learning by reading a manual.
Supervised learning is like observing hundreds of others play and trying to imitate them.
The team found that when an LLM is trained via RL trial-and-error to produce correct answers, it naturally learns to output its reasoning process.
For tasks like math and programming with verifiable answers, they implemented a scoring system to help DeepSeek-R1 improve during training — correct answers earn high scores, incorrect answers low scores.
They introduced an RL algorithm called Group Relative Policy Optimization (GRPO) and trained models such as DeepSeek-R1-Zero and DeepSeek-R1 based on the DeepSeek-V3 Base foundation model.
RL Framework
Starting with DeepSeek-V3 Base, the research team trained DeepSeek-R1-Zero, DeepSeek-R1 Dev1, Dev2, Dev3, and finally DeepSeek-R1 through a multi-stage pipeline involving rejection sampling, RL, and supervised fine-tuning (SFT).
The Multi-Stage Pipeline of DeepSeek-R1
DeepSeek-R1-Zero naturally evolved diverse and complex reasoning behaviors. When solving reasoning problems, the model tended to generate longer responses that included verification, reflection, and exploration of alternatives, demonstrating that RL enabled it to learn superior reasoning strategies.
However, DeepSeek-R1-Zero had limitations, such as poor readability and mixed-language output. Since its rule-based RL training focused only on reasoning tasks, performance in writing or open-domain question-answering was weaker.
Subsequent stages improved its capabilities:
The models were evaluated across 21 mainstream benchmarks including MMLU, MMLU-Pro, C-Eval, GPQA Diamond, SimpleQA, SWE-bench Verified, LiveCodeBench, and AIME 2024. DeepSeek-R1 achieved superior results on nearly all, validating the RL framework’s effectiveness.

Evaluation Results of Each Training Stage of DeepSeek-R1
The RL framework also facilitated emergent high-level reasoning behaviors, such as self-reflection, verification, and dynamic strategy adaptation, which can systematically guide and enhance reasoning in smaller models.
Lessons: Curbing AI Hype
Beyond the scientific significance, Nature highlighted an under-discussed issue: most widely used LLMs, rapidly transforming human knowledge acquisition, have not undergone independent peer review — a notable gap.
The publication of DeepSeek-R1 is “a welcome step toward transparency.” Its originality, methodology, and robustness were reviewed by eight human experts, with reviews and author responses published alongside.
Peer review helps:
Since DeepSeek-R1 is open-weight, anyone can download, test, and build upon it, making safety crucial. Reviewers noted the paper initially lacked information on safety testing; the team added a dedicated section describing safety evaluations compared to competing models.
Nature’s editorial concluded: peer review should become more common in AI, allowing companies to **justify claims with evidence** while ensuring validation.
Given the fierce global AI competition, some companies ignore data bias and model safety, or exaggerate model capabilities, posing real risks to society. As Nature notes, independent peer review is a means to mitigate AI hype.
Editor: Zhongxiaowen



