
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

Here’s a finding that genuinely surprised me when I sat with it. Shaoyang Xu and Wenxuan Zhang at the Singapore University of Technology and Design (arXiv:2601.11227, January 2026) showed that the language an LLM uses for its internal reasoning measurably changes the diversity of its English outputs — even when
