
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

There’s a quiet result from Cheng et al. (Renmin University and Microsoft Research, 2026) that I keep coming back to. They gave large language models access to a minimal sandbox — essentially a code interpreter with a file system — and watched what happened. No additional training. No new data.
