
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

Most improvements to AI systems happen before deployment: better training data, better fine-tuning, better RLHF. Once the model is out in the world, you generally get the performance you trained for and nothing better. A paper from early 2026 takes a different approach. Instead of trying to bake everything into
