
The Simplest Step That Makes LLMs Actually Useful
Before you reach for RLHF, before you design a reward model, before you start thinking about reinforcement learning from verifiable rewards — there’s a more fundamental question worth asking: has
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Before you reach for RLHF, before you design a reward model, before you start thinking about reinforcement learning from verifiable rewards — there’s a more fundamental question worth asking: has

There’s a quiet result from Cheng et al. (Renmin University and Microsoft Research, 2026) that I keep coming back to. They gave large language models access to a minimal sandbox — essentially a code interpreter with a file system — and watched what happened. No additional training. No new data.
