
AI Text Detectors Fail in the Wild. Now We Know Why.
Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they

A new paper from Christopher Burger, Karmece Talley, and Christina Trotter (arXiv:2512.23587), accepted to the 59th Hawaii International Conference on System Sciences, asks a deceptively simple question: can today’s frontier LLMs reliably identify AI-generated text? They tested GPT-4, Claude, and Gemini in a computing-education setting, where the stakes — academic
