
More Data Won’t Save Your LLM. Better Data Will.
There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

There was a point, not long ago, when the dominant strategy for improving large language models was simple: feed them more. More tokens, more compute, more parameters. The scaling laws

A practical post. The GitHub repo LLM4Annotation (Zhen-Tan-dmml/LLM4Annotation) is a curated, updated list of papers and tools at the intersection of LLMs and data annotation. If you work in this space, it’s worth bookmarking. What’s in it: The reason it’s worth following — beyond the obvious “saves you a few
