
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

If you’ve worked on NLP for low-resource languages — Yoruba, Quechua, Tigrinya, Hmong — you know the math is brutal. Native speakers with domain expertise are scarce. Their time is expensive. The data needs are enormous. And you can’t bootstrap from a “close enough” language without inheriting serious bias. A
