
The 42× number that should change how you scope annotation projects
If you’ve worked on NLP for low-resource languages — Yoruba, Quechua, Tigrinya, Hmong — you know the math is brutal. Native speakers with domain expertise are scarce. Their time is
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

If you’ve worked on NLP for low-resource languages — Yoruba, Quechua, Tigrinya, Hmong — you know the math is brutal. Native speakers with domain expertise are scarce. Their time is

There’s a quiet result from Cheng et al. (Renmin University and Microsoft Research, 2026) that I keep coming back to. They gave large language models access to a minimal sandbox — essentially a code interpreter with a file system — and watched what happened. No additional training. No new data.
