
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

The “more data wins” era is mostly over. What’s replacing it is messier and more interesting: a stack of techniques — fine-tuning, preference optimization, retrieval — each with its own cost curve, its own failure modes, and its own demands on data quality. Pick the wrong lever and you can
