
What If AI Agents Could Catch Their Own Mistakes?
Most improvements to AI systems happen before deployment: better training data, better fine-tuning, better RLHF. Once the model is out in the world, you generally get the performance you trained
Booth 21-25 | AI Data Management Zone | Tokyo Big Sight

Most improvements to AI systems happen before deployment: better training data, better fine-tuning, better RLHF. Once the model is out in the world, you generally get the performance you trained

A new paper from Christopher Burger, Karmece Talley, and Christina Trotter (arXiv:2512.23587), accepted to the 59th Hawaii International Conference on System Sciences, asks a deceptively simple question: can today’s frontier LLMs reliably identify AI-generated text? They tested GPT-4, Claude, and Gemini in a computing-education setting, where the stakes — academic
