The near-bottomless appetite for expert human judgment has turned data labeling from an AI cost center into one of the fastest-scaling categories in enterprise software. Micro1’s latest numbers show the market is far from winner-take-all — and reveal where the real margin lives.
The number, and what’s actually inside it
Micro1, a four-year-old startup that recruits doctors, lawyers, engineers, and scientists to train frontier models, has grown its gross annual run rate from $100 million to $500 million in roughly eight months, according to a person familiar with the business.
That five-fold jump deserves an asterisk, and it’s an important one for anyone benchmarking this category. Gross run rate in a marketplace model is not revenue in the software sense — a large slice flows straight back out to the contractors doing the work. Micro1 keeps somewhere between 60% and 70% of the gross figure, which puts its net annualized take in the $150 million to $200 million range.
Even discounted, that’s a serious business. And it’s being built in a market where the leaders are already several turns ahead:
- Advertisement -
| Player | Reported gross annualized revenue | Origin story |
| Mercor | ~$2B (summer 2026) | Started as an AI recruiter |
| Handshake AI | ~$1B (earlier in 2026) | Started as a campus job board |
| Micro1 | $500M | Started as an AI recruiter (Zara) |
Notice the pattern. Two of the three biggest names in expert data didn’t set out to sell data at all. They were recruiting companies that discovered their real product was the vetting layer — the ability to find, screen, and deploy a credentialed radiologist or securities lawyer in days rather than quarters. Micro1’s founder, Ali Ansari, pivoted for exactly that reason: he noticed data-labeling clients were already using his recruiting platform to source annotators.
Why there’s room for a fourth, fifth, and sixth winner
The instinctive read on a category with a $2B leader is that the window has closed. Micro1’s trajectory argues the opposite.
Demand here isn’t a fixed pie being carved up. It’s expanding faster than any single supplier can absorb — because the nature of the work keeps changing. Static labeling (bounding boxes, sentiment tags) was the old market. The frontier work has moved to reinforcement-learning environments: interactive sandboxes where an agent attempts a real task and an expert-written rubric grades the result. That work is dramatically more labor-intensive, more expensive per unit, and requires far scarcer expertise.
Anthropic alone has been reported to be weighing more than $1 billion a year on RL environments across a dozen-plus vendors. Some researchers now openly hypothesize that AI spending on data could eventually rival spending on compute — which, if even directionally right, reframes this from “AI services niche” to a line item on the same order as data centers.
Ansari has projected the market growing from roughly $10-15 billion today toward $100 billion within two years. Take founder forecasts with the usual salt, but the structural argument holds: models improve on judgment, and judgment is bought by the hour.
- Advertisement -
The margin story is the real headline
Here’s the part most coverage glosses over, and it’s the part that determines whether these are durable businesses or high-volume staffing agencies with better branding.
Micro1 is doing two things to escape the labor-cost trap:
1. Generating data without humans in the loop. The company is increasingly producing synthetic data — automated descriptions of video content, for instance — where the marginal cost of the next unit collapses toward zero.
- Advertisement -
2. Selling the same dataset more than once. When a dataset can be licensed to multiple buyers as an “off-the-shelf” product, gross margins on that revenue reportedly reach 80% to 90%.
That second move is the entire arc of this category in miniature: from services, to productized services, to product. A staffing marketplace clearing 30-35% take rate is one business. A data library with 85% gross margins is a completely different one — and it’s the version that justifies the valuations being discussed. Micro1 raised its Series A at a $500 million valuation last September, and TechCrunch understands a subsequent round may have closed at a substantially higher mark.
Contract sizes are also growing at an accelerating clip, which points toward the same conclusion: buyers are consolidating spend with fewer vendors and going deeper with each.
The uncomfortable part: who else is buying
Reselling the same dataset is efficient. It’s also where this market has gotten politically radioactive.
Forbes reported earlier this month that American data labelers have been supplying Chinese AI developers — with an estimated ~$500 million a year flowing from the top Chinese labs to US vendors — under essentially no export-control regime. The argument the critics make is straightforward: if the human-preference data that makes a US frontier model good is available off the shelf, chip controls stop being the load-bearing lever anyone assumed they were.
Ansari has staked out a position on the other side of that line, posting last month that Micro1 does not sell to Chinese model makers and calling the practice <cite index=”3-1″>”shameful”</cite> for companies claiming to champion American AI leadership while supplying adversarial competitors.
Whether that’s principle, positioning, or both — it’s now a differentiator with commercial weight. Vendors holding US federal contracts while serving Beijing labs are carrying real regulatory exposure. Expect “we don’t sell to X” to become a standard clause in this category’s sales motion, the same way data residency became table stakes in SaaS.
What B2B marketers should actually take from this
Strip away the AI framing and there are four transferable lessons here:
Your distribution asset may be worth more than your product. Mercor, Handshake, and Micro1 all discovered that the network they built for one purpose (hiring) was the scarce asset for another (expert supply). Audit what you’ve accumulated that a faster-growing adjacent market desperately needs.
Gross vs. net is a positioning choice, and buyers are getting wise to it. This category headlines gross run rate because it’s the bigger number. Sophisticated buyers and investors now reflexively apply the take-rate discount. If your category has a similar vanity metric, decide deliberately whether leading with it builds credibility or erodes it.
Productizing a service is the only path off the labor treadmill. The move from bespoke engagements to a licensable, multi-tenant asset took Micro1’s margin on that revenue from marketplace economics to software economics. Every agency and services business should be asking which of its deliverables could be sold twice.
Ethical positioning is becoming a procurement filter. Data provenance, client screening, and geopolitical exposure are moving from CSR-page material to line items in RFPs. Getting your position documented early is cheaper than retrofitting it under scrutiny.
