Half of DeepSeek's top researchers are said to be labeling data

A widely-shared — but officially unconfirmed — recording of DeepSeek's Liang Wenfeng argues the real bottleneck to better AI right now isn't compute or genius. It's patiently cleaning data.

Unverified Source Circulated recording transcript (unconfirmed by DeepSeek), via 36Kr / Jiemian / social ⚑ AI strategy

The 60-second version

A viral, officially-unconfirmed transcript of a DeepSeek investor meeting quotes Liang Wenfeng saying half the company's core researchers label data, because data and taste — not compute — are the current bottleneck to better AI.

Key points

  • The headline quote: ~half of DeepSeek's core researchers are labeling data; 'solving AI now comes down to data.'
  • Bottleneck reframed: not chips or geniuses, but high-quality data + the taste to know what 'good' is — and time.
  • AGI as a staircase: last year chain-of-thought, this year agents, next continual learning (which today's AI lacks).
  • Philosophy: restraint as strategy, open source as shared upside, 'reasonable profit' pricing, half of work-time free to explore.
  • Backdrop: first external raise, reportedly >¥50B at a ~¥360B+ pre-money valuation — figures vary and are unconfirmed.

Verdict. Unverified, so read it as a widely-reported claim — but it lands because it's true in spirit: the quiet, human work of data and judgment may matter more than raw compute.

The line"Half our core researchers are labeling data"

The sentence that went viral this week: according to the transcript, DeepSeek founder Liang Wenfeng said you could think of half the company — and half of its core researchers, the most important people — as labeling data. His reasoning, as quoted: at this stage, solving AI comes down to data.

That's a deliberately provocative framing. It cuts against the usual AI story of ever-bigger models and ever-more GPUs, and puts the spotlight on something far less glamorous: patiently deciding what a good answer looks like, and teaching it to the model.

The real bottleneckData and taste, not just compute

In the transcript's telling, the constraint right now isn't chips or a shortage of brilliant people — it's high-quality data and the taste to recognize what "good" is. The bottleneck behind *that* is time: rivals like OpenAI and Anthropic started earlier, with more capital and more compute, so the stated approach is to be deliberate — label the cheaper data first, and keep grinding.

On stage, everyone talks about intelligence emerging. Backstage, the scarce thing is still human patience — cleaning dirty data one row at a time.

The roadmapAGI as a staircase: CoT → agents → continual learning

The transcript frames the path to AGI as climbing stairs. Last year's step was chain-of-thought — models reasoning step by step. This year's step is agents — models that take actions. The next step is *continual learning*: the ability to keep learning on the job.

The claim: today's AI doesn't lack taste or intuition — it lacks continual learning. You'd have to feed it all the context every time, which is nearly impossible. That, he argues, is why AI can't simply replace employees yet.

The philosophyRestraint, open source, and "reasonable profit"

Restraint (克制)A strategy — give some things up to gain more of what matters
Open sourceSharing the upside: staff pride, company cohesion, community benefit
AmbitionNot trying to be the next super-app / ByteDance / Tencent
PricingOnly a reasonable profit — ~the margin of recouping equipment in ~10 months
How people workAssigned work ≤ half your time; the other half is free to explore

The backdropA very expensive quiet

The meeting was tied to DeepSeek's first external funding round — reportedly raising more than ¥50B (~$7B) on a pre-money valuation around ¥360B+ (roughly ~$50B). Exact figures vary between reports, so hold them loosely.

Why it mattersCompute gets the headlines; data may decide the race

Even treated with healthy skepticism, the recording resonated for a real reason. Behind every headline about *emergent intelligence* sits an unglamorous, deeply human layer: deciding what a good answer is, and cleaning the data that teaches it. Compute gets the attention — data and taste may quietly decide who wins.