# Half of DeepSeek's top researchers are said to be labeling data

> A widely-shared — but officially unconfirmed — recording of DeepSeek's Liang Wenfeng argues the real bottleneck to better AI right now isn't compute or genius. It's patiently cleaning data.

_Source: Circulated recording transcript (unconfirmed by DeepSeek), via 36Kr / Jiemian / social · 2026-07-23 · 6 min read_

Canonical: https://iyu.app/e/deepseek-half-team-labels-data

## The 60-second version

A viral, officially-unconfirmed transcript of a DeepSeek investor meeting quotes Liang Wenfeng saying half the company's core researchers label data, because data and taste — not compute — are the current bottleneck to better AI.

**Key points**

- The headline quote: ~half of DeepSeek's core researchers are labeling data; 'solving AI now comes down to data.'
- Bottleneck reframed: not chips or geniuses, but high-quality data + the taste to know what 'good' is — and time.
- AGI as a staircase: last year chain-of-thought, this year agents, next continual learning (which today's AI lacks).
- Philosophy: restraint as strategy, open source as shared upside, 'reasonable profit' pricing, half of work-time free to explore.
- Backdrop: first external raise, reportedly >¥50B at a ~¥360B+ pre-money valuation — figures vary and are unconfirmed.

**Verdict.** Unverified, so read it as a widely-reported claim — but it lands because it's true in spirit: the quiet, human work of data and judgment may matter more than raw compute.

## Full explainer

> **⚑ Caveat:** **Unverified source.** This is based on a *circulated transcript* of a DeepSeek investor meeting attributed to Liang Wenfeng, republished by outlets like 36Kr and Jiemian. **DeepSeek has not confirmed the recording is authentic.** Treat every quote and figure below as a widely-reported claim, not an official statement.


### The line — "Half our core researchers are labeling data"

The sentence that went viral this week: according to the transcript, DeepSeek founder **Liang Wenfeng** said you could think of half the company — and half of its **core researchers, the most important people** — as labeling data. His reasoning, as quoted: at this stage, solving AI **comes down to data**.

That's a deliberately provocative framing. It cuts against the usual AI story of ever-bigger models and ever-more GPUs, and puts the spotlight on something far less glamorous: patiently deciding what a good answer looks like, and teaching it to the model.


### The real bottleneck — Data and taste, not just compute

In the transcript's telling, the constraint right now isn't chips or a shortage of brilliant people — it's **high-quality data** and the **taste** to recognize what "good" is. The bottleneck behind *that* is time: rivals like OpenAI and Anthropic started earlier, with more capital and more compute, so the stated approach is to be deliberate — **label the cheaper data first**, and keep grinding.

> On stage, everyone talks about intelligence emerging. Backstage, the scarce thing is still human patience — cleaning dirty data one row at a time.

> **◆** A tension worth noting: DeepSeek's own R1 research (a 2025 Nature cover paper) emphasized using reinforcement learning to *reduce* reliance on manual annotation. "Labeling data" here likely means crafting high-quality examples and reward signals — curation and judgment — rather than old-style bulk tagging.


### The roadmap — AGI as a staircase: CoT → agents → continual learning

The transcript frames the path to AGI as climbing stairs. **Last year's step** was chain-of-thought — models reasoning step by step. **This year's step** is agents — models that take actions. **The next step** is *continual learning*: the ability to keep learning on the job.

The claim: today's AI doesn't lack taste or intuition — it lacks continual learning. You'd have to feed it all the context every time, which is nearly impossible. That, he argues, is why AI **can't simply replace employees** yet.


### The philosophy — Restraint, open source, and "reasonable profit"

- **Restraint (克制):** A strategy — give some things up to gain more of what matters
- **Open source:** Sharing the upside: staff pride, company cohesion, community benefit
- **Ambition:** Not trying to be the next super-app / ByteDance / Tencent
- **Pricing:** Only a reasonable profit — ~the margin of recouping equipment in ~10 months
- **How people work:** Assigned work ≤ half your time; the other half is free to explore


### The backdrop — A very expensive quiet

The meeting was tied to DeepSeek's **first external funding round** — reportedly raising **more than ¥50B** (~$7B) on a **pre-money valuation around ¥360B+** (roughly ~$50B). Exact figures vary between reports, so hold them loosely.

> **⚑ Caveat:** Valuation and fundraise numbers differ across republished versions (some cite ¥367.5B pre-money, others higher post-money figures). None are officially confirmed by DeepSeek.


### Why it matters — Compute gets the headlines; data may decide the race

Even treated with healthy skepticism, the recording resonated for a real reason. Behind every headline about *emergent intelligence* sits an unglamorous, deeply human layer: deciding what a good answer is, and cleaning the data that teaches it. Compute gets the attention — **data and taste may quietly decide who wins.**


## Primary sources

- [@vista8 on X — the viral post](https://x.com/vista8/status/2080127337790927060)
- [36Kr — Liang Wenfeng's 4-hour investor Q&A transcript](https://www.36kr.com/p/3907578194417028)
- [Jiemian — DeepSeek fundraising meeting transcript](https://www.jiemian.com/article/14816174.html)

---
_Published by iyu (https://iyu.app) — the day's AI news, checked against primary sources and rewritten in plain language. Free to quote with attribution and a link to the canonical URL._
