DeepSeek Founder’s Latest Research Crushes the China–U.S. Hardware Gap

cs_opinion_img
The latest paper by DeepSeek founder Liang Wenfeng once again highlights the latent potential of China’s tech sector when it breaks through U.S. technological constraints. The model training technique described in the paper bypasses what is currently China’s biggest hardware shortfall compared with the United States and may once again significantly reduce the cost of AI training.
January 15, 2026
The China Academy Picks
Top picks selected by the China Academy's editorial team from Chinese media, translated and edited to provide better insights into contemporary China.
Click Register
Register
Try Premium Member
for Free with a 7-Day Trial
Click Register
Register
Try Premium Member for Free with a 7-Day Trial

On the evening of January 12, Liang Wenfeng, founder of Chinese AI start-up DeepSeek, co-authored a technical paper with researchers from Peking University proposing a new model training method. They said the technique enables “aggressive parameter scaling” by bypassing GPU memory limitations.

A January 13 report by the South China Morning Post said the move underscores DeepSeek’s continued focus on maximizing cost efficiency despite its relative disadvantage in computing power compared with leading U.S. players. It also noted market speculation that the company may release a major new model before the Lunar New Year.

The highly technical paper is expected to draw wide attention from industry insiders in both China and the United States eager to learn about DeepSeek’s latest progress.

In the paper, titled “Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models,” the authors introduce a “conditional memory” technique called Engram. The method is designed to address a key bottleneck in scaling AI models — the limited capacity of high-bandwidth memory (HBM) on GPUs.

Existing large language models retrieve basic information through computation, a process that consumes vast computing resources. The researchers argue that this wastes valuable “sequential depth,” which could otherwise be allocated to higher-level reasoning tasks.

The SCMP noted that HBM is one of the largest gaps between China and the U.S. in AI hardware. Ray Wang, an analyst at South Korea-based SemiAnalysis, said that although China has made steady progress, its memory-chip champion ChangXin Memory Technologies (CXMT) still trails industry leaders such as Samsung Electronics, SK Hynix and U.S. firm Micron Technology by several years.

The paper explains that by “decoupling” computation from storage, Engram allows models to “look up” this foundational information far more efficiently. The technique also improves efficiency in handling long-context inputs — one of the biggest hurdles in turning AI chatbots into practical real-world agents.

The researchers validated the method on a 27-billion-parameter model, finding that it boosted performance on major industry benchmarks by several percentage points. Crucially, it also preserves more capacity for complex, compute-intensive reasoning.

DeepSeek Founder LiangWenfeng

“We believe conditional memory will become an indispensable modeling primitive in the next generation of sparse models,” they wrote, likening Engram’s potential impact to their previously developed Mixture-of-Experts (MoE) approach, which enables model scaling without proportional increases in computation and has since been adopted by other Chinese competitors.

Today’s largest models contain trillions of parameters. Elie Bakouch, a research engineer at open-source platform Hugging Face, praised the paper on social media, saying it had “validated the technique on real hardware during both inference and training.”

The paper lists 14 co-authors, including Zhang Huishuai, an assistant professor at Peking University’s Wangxuan Institute of Computer Technology and former principal researcher at Microsoft Research Asia.

Early last year, DeepSeek released its DeepSeek-R1 model, trained on a data center powered by Nvidia H800 GPUs. It completed training in just two months at a cost of US$5.5 million — only a fraction of what U.S. firms such as OpenAI reportedly spend — while achieving performance comparable to top American models, drawing global attention, particularly in the United States.

On January 12, the Financial Times reported that Microsoft president Brad Smith warned that U.S. AI companies are being overtaken by Chinese competitors in the race for users outside the West, citing China’s low-cost open-source models as a key advantage.

Smith said DeepSeek’s technology is spreading rapidly in emerging markets such as Africa, highlighting intensifying global competition. “We must recognize that, unlike a year ago, China now has — and increasingly has more than one — competitive open-source model,” he said.

The report added that a new Microsoft study found DeepSeek’s R1 model, released a year ago, helped accelerate global AI adoption thanks to its “ease of use and low cost,” particularly in countries across the Global South. This has enabled China to surpass the United States in global market share for open-source AI models, which are typically free for developers to use, modify and integrate.

The SCMP noted that as the first anniversary of R1 approaches, expectations are rising for DeepSeek to unveil another major model. Silicon Valley tech outlet The Information reported on January 9 that the company is expected to release a powerful new V4 model with strong programming capabilities in mid-February.

Editor: Zhongxiaowen

$10 MONTHLY
VIEWS BY

author_image
Top picks selected by the China Academy's editorial team from Chinese media, translated and edited to provide better insights into contemporary China.
Share This Post

Leave a Reply