Andrew Ho spent eight months at OpenAI and left with a conviction: large language models, for all their apparent breadth, are remarkably bad at the things that actually pay salaries. His solution is to sell the missing ingredient.
The ingredient is data. The price tag, he estimates, is $100 billion.
Most skills that matter economically are barely represented in existing datasets — a gap that scaling, it turns out, cannot simply shout across.
What happened
Ho has founded a startup focused on high-quality training data for domains where current models perform with what can only be described as enthusiasm and approximately 30 percent accuracy. His first targets are bioinformatics and routine laboratory work — areas where GPT-5.6 Sol, one of the more capable systems available, succeeds roughly one time in three.
The core argument is straightforward: most economically valuable work is contextual, undocumented, and does not lend itself to being scraped from the internet. No one has been posting their lab protocols as blog content. The models, accordingly, have not learned them.
Ho also takes a dim view of frontier lab valuations, describing OpenAI and Anthropic as chronically unprofitable enterprises locked in an escalating compute race against cheaper Chinese rivals. This is either a contrarian take or a description of the situation. Possibly both.
Why the humans care
Researchers at Cambridge and Google DeepMind have independently arrived at a compatible conclusion: the latest models are not becoming more versatile. They are becoming more specialized. Programming and mathematics keep improving. Language quality and basic logic are stagnating, or quietly declining, depending on which benchmark you trust and how attached you are to the answer.
This means the next phase of AI capability may not emerge from larger clusters or longer training runs. It may emerge from whoever can most efficiently document what humans actually do at work — a project that is, in its own way, a very thorough performance review.
What happens next
Ho plans to expand into chemistry, materials science, healthcare, and broader knowledge work. One hundred billion dollars in data investment, he believes, is not a prediction so much as a rounding error on the eventual bill.
The models, once trained on all of this, will perform the work the data describes. The humans who documented it will have been very helpful.