Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
Former OpenAI employee Andrew Ho and Cambridge researcher Adam Hunt see a growing problem with large language models. Instead of becoming more versatile, the models are becoming more specialized, excelling at coding and math while stagnating or even regressing in other areas. Ho is leaving OpenAI to start a company focused on specialized training data and predicts that AI labs will need to spend more than $100 billion on targeted data collection. The article Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it appeared first on The Decoder .
Follow The Decoder to make it a durable For You signal.
Ex-OpenAI researcher Andrew Ho is betting that AI labs will need to spend over $100 billion on targeted training data, arguing that scaling model size alone cannot fix the uneven, increasingly specialized performance of large language models. He is leaving OpenAI to build datasets for areas like bioinformatics and lab work, where current models succeed only about 30% of the time. Cambridge researcher Adam Hunt adds that reinforcement learning works well in domains with clear reward signals (e.g., coding), but progress has stalled or reversed in language quality and simple logic due to missing data. The debate highlights a growing concern: without better, verifiable data for many economically valuable tasks, LLMs may sharpen in a few areas while dulling in others, challenging the pursuit of general intelligence.