
AI Text Detectors Fail in the Wild. Now We Know Why.
Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Most AI text detectors look good in testing. They’re trained and evaluated on specific models, specific prompt styles, specific domains — and they perform well within that distribution. Then they

Researchers built a study to answer a question that sounds simple: can people tell whether a scientific abstract was written by an LLM or a human? The answer, for people with machine-learning expertise, was mostly: no. This result is interesting on its own. The implications for how we build annotation
