← Back to all posts
News

Beijing's Answer to 17.9% Youth Joblessness: Go Train the AI

September 11, 2026 · 03:12 UTC · News
Beijing's Answer to 17.9% Youth Joblessness: Go Train the AI

TL;DR

Rest of World reported on September 9 that underemployed Chinese professionals are taking gig work on platforms run by ByteDance, Alibaba, Tencent and Moonshot AI, writing training data that teaches models to perform their own trades. Pay is 100 to 500 yuan ($15 to $74) for a task that usually takes hours, and rejected work pays nothing. The market this feeds is forecast by IDC at 7.8 billion yuan ($1.1 billion) in 2026, up 25% on the year, and Beijing is openly recruiting graduates into it while youth unemployment sits at 17.9%.


The job is not labeling. It is dictation.

Forget bounding boxes and sentiment tags. The work described here is a working professional sitting down with documents from their actual job, inventing a task the model cannot currently do, and then writing out their own reasoning so the model can copy it.

Architects write building proposals and profitability analyses. Lawyers give advice as a virtual lawyer, then draft judicial decisions from a judge's perspective. Software engineers upload text files that probe text processing and build assignments aimed at agentic capabilities. Trainers are explicitly barred from using AI to do any of it, which is a rule with a certain poetry to it.

"It is just like teaching a child. Every task has to come from my actual work." Cuicui, a Shenzhen architect working through the platform TalentsAI

This is the part worth internalizing if you build with models. The bottleneck in workplace and agentic AI is not tokens or parameters. It is what China's National Data Administration called, in a June policy push, the knowledge density of training data.

Think of it as the difference between a group chat log and a textbook. Scrape enough group chats and the model learns what people sound like. It does not learn what a structural engineer checks before signing off on a load calculation, because nobody types that out. Somebody has to be paid to.

What it pays, and what happens when it does not

Rest of World puts the platform gig tier at 100 to 500 yuan per task, with tasks typically running several hours. Quality control can reject the submission, and ByteDance's Xpert says in its terms that trainers are not paid if work fails review. Alibaba's Siriser, launched in January, lets trainers revise or appeal.

There is a higher tier. A May report from 36Kr found hourly rates of 500 to 800 yuan for vertical specialists in finance, law and medicine, with 300 to 500 yuan common in finance alone, and outsourced monthly roles paying 8,000 to 10,000 yuan against a traditional annotation wage of 3,000 to 4,000. The entry requirement in that tier has risen to a master's degree or above.

monthly pay, outsourced annotation roles (yuan, 36Kr May 2026) traditional low3,000 traditional high4,000 expert tier low8,000 expert tier high10,000 entry bar for the expert tier: master's degree or above
The expert tier pays roughly 2.5x the old annotation wage, and asks for a graduate degree to get in.

Note the shape of that. Annotation did not get more lucrative. A different, harder job took its name.

Why there are people to hire

China's urban unemployment rate for 16 to 24 year olds excluding students hit 17.9% in July, up from 14.9% in June and the worst reading since August 2025, per National Bureau of Statistics data reported by Caixin. The 25 to 29 bracket sat at 7.2%, and 30 to 59 year olds at 3.9%. Part of the July jump is the annual graduation cycle, but it is the highest July print in three years.

Julie Yujie Chen, a University of Toronto researcher quoted in the piece, framed the supply side bluntly: China's ordinary white-collar middle class faces "this urgency to find a supplementary income." The platforms found a workforce of credentialed professionals with spare hours and a mortgage.

The state is not a bystander here

In June the National Data Administration backed the construction of high-quality datasets, called for industry experts to be pulled deeper into annotation to raise knowledge density, and encouraged college graduates to treat annotation as flexible employment.

The scale is already public. Per a July 19 government release, China had built roughly 120,000 high-quality datasets totaling over 1,565 petabytes as of the end of June, up more than 60% from the end of Q1. Seven data annotation pilot cities, including Chengdu, Shenyang, Hefei and Changsha, had passed 119 petabytes of annotation work and employed about 140,000 people.

Industrial policy usually means fabs and subsidies. This is industrial policy for a labor input, and the input is other people's professional judgment.

The same work, priced two ways

The Western version of this market is not smaller. It is enormously larger, and it is one company. Mercor, three years old, hit a $2 billion gross annualized revenue run rate in June, at least double where it sat entering 2026, per Sacra and Dealroom, both drawing on The Information and Mercor's own disclosures. It pays out more than $2 million a day to over 30,000 weekly active contractors, who take 60% to 70% of topline, and Sacra puts average expert earnings above $85 an hour. It was in talks in July at a $20 billion valuation.

annualized scale, us dollars (billions) Mercor, Jun 20262.0 gross China market 20261.1 forecast not like for like: customer spend vs a whole-market forecast
One three-year-old marketplace moves more gross spend than China's entire forecast training-data market.

Set against that, IDC's August forecast of 7.8 billion yuan ($1.1 billion) for China's whole training-data market in 2026 is a modest number, and it follows 6.26 billion yuan in 2025. IDC also flagged the direction of travel: demand is rotating out of consumer and entertainment data and into workplace productivity.

That rotation is the story underneath the story. Tiezhen Wang, formerly Hugging Face's Asia-Pacific ecosystem head, told Rest of World that Chinese labs have largely closed the gap on general reasoning and coding. The next contested ground is software that does your job, and that is exactly the data these platforms are buying.

How a task becomes a weight

your real workfile + reasoning a task the modelcannot yet solve platform qualitycontrol review 100-500 yuan,or nothing
Hours of expert work, graded pass or fail, with the downside entirely on the trainer.

The workers in the piece are clear-eyed about what they are building. A Shanghai software engineer: "You have to keep evolving yourself to avoid getting completely replaced." Liu, who holds a law degree and trains on Moonshot and Tencent platforms, put the game theory in one line.

"If I don't pass my knowledge to AI, others will do that."

Jessica, a software engineer with 20 years of experience, earned hundreds of yuan for her time and came away worried about where her uploaded material ends up. That concern has no obvious remedy in any of the platform terms described.

What a builder should take from this

  • Expert data has a spot price now, and it is regional. The same category of reasoning trace is bought at 100 to 500 yuan a task in Shenzhen and through a $2 billion gross run rate marketplace in San Francisco. If you are budgeting a domain eval set or a fine-tune corpus, that spread is your negotiating range.
  • The moat in workplace AI is sourcing, not architecture. Both countries are converging on the same answer: pay credentialed practitioners to narrate their judgment. Nobody is scraping their way to a virtual tax accountant.
  • Rejection risk sits on the supplier. Unpaid rejected work is a structural feature of these marketplaces, which means published market sizes count what buyers spend, not what workers earn.

Caveats

The dollar comparison in the second chart is directional, not apples to apples. Mercor's $2 billion is gross customer spend annualized from a point in time, with 60% to 70% flowing to contractors; IDC's 7.8 billion yuan is a full-year forecast for a market category whose boundaries IDC defines and which may exclude in-house labeling done by the Chinese labs themselves. The Rest of World pay figures come from named and anonymous workers on specific platforms, not a survey, and the workers quoted use partial names. The 140,000 annotation employment figure covers seven pilot cities, not the country. And the 17.9% youth rate excludes students and carries a known seasonal graduation spike, so it is a real signal about graduate absorption rather than a clean measure of white-collar distress.

Key Takeaways

  • Chinese professionals are paid 100 to 500 yuan ($15 to $74) per task, taking hours each, to write training data that teaches models their own trades, on platforms from ByteDance, Alibaba, Tencent and Moonshot AI.
  • Rejected submissions pay nothing under Xpert's terms; Siriser allows revision or appeal. The quality risk sits with the worker.
  • IDC forecasts China's training-data market at 7.8 billion yuan ($1.1 billion) in 2026, up 25% from 6.26 billion in 2025, with demand shifting from consumer data to workplace productivity.
  • Beijing's National Data Administration backed the buildout in June and encouraged graduates into annotation as flexible employment, against 17.9% youth unemployment in July.
  • Mercor alone hit a $2 billion gross annualized run rate in June, paying over $2 million a day to more than 30,000 weekly active contractors at an average above $85 an hour, which dwarfs China's entire forecast market on a gross basis.
  • Expert reasoning traces, not scraped text, are the contested input for agentic and workplace AI in both markets.

Sources: Rest of World, "Now it's China's experts who are gig workers training AI data", 36Kr on China's AI data experts, Caixin Global on July youth unemployment, The State Council of the PRC on high-quality datasets, Communications Today on the 2025 IDC market figure, Global Times on the data annotation industry, Sacra on Mercor, Dealroom on Mercor's run rate, TechCrunch on Mercor's valuation talks

AITrendsChinaTraining DataRLHFGig EconomyFuture of WorkMercor
CONSOLE
$