
AI Text Detectors Fail in the Wild. Now We Know Why.
Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they

Most improvements to AI systems happen before deployment: better training data, better fine-tuning, better RLHF. Once the model is out in the world, you generally get the performance you trained for and nothing better. A paper from early 2026 takes a different approach. Instead of trying to bake everything into
