
AI Text Detectors Fail in the Wild. Now We Know Why.
Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they

There’s a quiet result from Cheng et al. (Renmin University and Microsoft Research, 2026) that I keep coming back to. They gave large language models access to a minimal sandbox — essentially a code interpreter with a file system — and watched what happened. No additional training. No new data.
