
AI Text Detectors Fail in the Wild. Now We Know Why.
Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they

Here’s a finding that genuinely surprised me when I sat with it. Shaoyang Xu and Wenxuan Zhang at the Singapore University of Technology and Design (arXiv:2601.11227, January 2026) showed that the language an LLM uses for its internal reasoning measurably changes the diversity of its English outputs — even when
