
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

A 34-author survey led by Yu Xia at UC San Diego — with collaborators across Adobe Research and nine other labs — was accepted to ACL 2025, and it’s worth your time if you build data pipelines. The title carries the whole argument: “From Selection to Generation.” For most of
