AI can now hide its own writing from detection software, Duke-IISc researchers find
AI-summarised brief · reviewed before publication
Researchers from Duke University and the Indian Institute of Science have demonstrated that AI‑generated text can evade detection tools by sourcing most of its words from a “raw” base model rather than a fine‑tuned chatbot. The method exploits a weakness identified by Carnegie Mellon researchers: detectors are calibrated to spot patterns introduced during the second‑stage instruction‑following training of models like ChatGPT or Claude. By feeding a detector‑evading base model into Anthropic’s Claude Opus 5, the team produced output that Pangram’s widely used detector classifies as human‑written. The findings, posted online on 25 September and not yet peer‑reviewed, arrive shortly after the European Union mandated labeling of machine‑generated content and after a major AI conference rejected papers based on detector verdicts.
💡 Why It Matters
- · The technique reveals a loophole that could undermine current safeguards against undisclosed AI authorship, threatening the credibility of academic, journalistic, and educational publications.