Last updated on August 7, 2026
Centuries ago, people accused of witchcraft were bound and thrown into bodies of water, a test known as “ducking.” If the accused witch floated, they were guilty of being a witch and sent to trial, which could end in imprisonment or being burned at the stake. If they sank, they were innocent of being a witch, but lost their life.
Today there’s a new, more modern type of witch hunt going on: the AI witch hunt. It has to do with whether or not artists, writers, students, and companies are using artificial intelligence in their products.
Instead of throwing the accused in the nearest river, we have AI detectors that can tell us how much of something appears to be generated by artificial intelligence. A document or image is scanned and a number appears on the screen: 87% AI. 52% AI. 26% AI.
Behold, the modern witch’s mark.
AI detectors are a form of automated trust: the practice of substituting an algorithmic score for human judgment about credibility and authorship. Rather than evaluating evidence, context, and process, a decision-maker accepts a probability output as a finding of fact.
A person, group, or platform then turns that probability into a verdict (“this is AI”) then acts on the verdict as truth. And the witch hunt begins.
But do AI detectors really prove the use of generative AI? Or do they make errors, punish polished work, and potentially ruin reputations? Can we really automate trust using the same tools we’re told to distrust everywhere else?
The good news is that there are concrete ways to protect yourself before an accusation ever happens, and just as importantly, ways to respond if one does.
In this post, we’ll discuss how AI detection works, the risks of AI detection false positives, and how to protect your creative and intellectual property from the AI witch hunt.
Is AI-Assisted Work the New Norm?
Not every accusation is an AI witch hunt. Some people really are passing off completely unedited AI output as their own work, educational, or creative content.
Many people have been vocal in recent years about wanting to consume human-only content in their feeds and to have some reassurance that they are supporting human creators instead of AI bots. That preference has value and deserves to be taken seriously.
But the reality is that these days, almost no one’s creative, educational, or professional life lives at 100% or 0% AI.
Generative AI tools are now built directly into the software people already use to create things. Consider these common scenarios:
- A writer who uses Grammarly as an editor to catch typos and tighten phrasing on an essay they wrote entirely themselves.
- A photographer who uses AI-assisted tools like background and spot removal on their own photos in Adobe Photoshop.
- An audio or video editor who uses AI to save time removing filler words, silence, and background noise during post-production in Descript.
- A blogger who designed an infographic template by hand in Canva, then uses AI within the same tool to auto-place their own written copy and resize elements to fit each time.
All of them touched a workflow that has AI somewhere inside it, and could be marked with the witch’s mark of “AI-generated.” But at what point does “AI-assisted” become “AI-generated”? And who decides how much AI is acceptable?
Schools, publishers, and social media have started to deploy AI detection software (like Turnitin, GPTZero, or Pangram) to address the problem of AI slop and misrepresented authorship.
Institutions and platforms aren’t deploying AI detection because it’s 100% accurate. They’re deploying it because doing something, even something unreliable, looks better than doing nothing.
That’s how the AI witch hunt develops: an anxious population, a shortage of certainty, and a test that arrives promising to settle the question.
The irony is that AI is being used to police AI. The thing we are told to be suspicious of is also the judge of whether or not someone’s work is human.
What Is AI Detection, and What Does It Actually Measure?
AI detection is the practice of using statistical classifiers to estimate whether text was produced by a human or a large language model. The tools score signals like perplexity, meaning how predictable each word choice is, and burstiness, meaning how much sentence length and structure vary. Predictable, evenly paced writing reads as machine.
The detector’s output is a probability, not a finding of fact. A probability is a guess. A finding of fact is something you can check and validate.
When a detector says “82% AI,” it’s not telling you anything about the human process that was going on while the piece was written. It’s telling you how much the finished text resembles other text the tool has already labeled as AI.
The tools learn by comparing large amounts of human writing to large amounts of AI writing, looking for patterns that separate the two. But since both piles come from human-written text originally, the “AI” pattern it flags is really just a pattern that already existed in human writing. The detector has no idea who wrote something or how much effort went into it. It only checks whether the text crosses a statistical threshold.
AI learned to sound human by reading human work. Now it’s the authority on which humans sound too much like AI.
Picture two people writing the same paper. One spends twenty hours researching, outlining, drafting, and re-writing. They use an AI tool as an editor to help improve their writing, but rewrite any generative AI sections to make them their own.
The other person types “write me a paper about this topic” into a chatbot and passes it off as their own.
A detector cannot tell these two processes apart. It never saw either one happen. It only sees the finished text, and it collapses twenty hours of genuine effort and five seconds of prompting into the same binary label of “AI” or “Not AI”.
Who’s Deploying AI Detection, and Why It’s Spreading Fast
AI detection didn’t stay confined to one industry for long. Once schools started using it to catch cheating, the same dream of detecting AI spread quickly to publishing, hiring, and social platforms within a couple of years. Each of these institutions is making the same bet: that an algorithmic score is better than no check at all, even when that score is unreliable.
The following are examples of how AI detection is used in creative, professional, and educational fields:
- Schools trusting AI detectors to flag academic dishonesty in both written and coding assignments, including AI-assisted or “vibe-coded” work, often built directly into learning management systems. Australian Catholic University logged nearly 6,000 misconduct referrals in 2024, roughly 90% AI-related, before discontinuing its detector in March 2025 after acknowledging a quarter of referrals were dismissed on investigation.
- Hiring managers trusting resume-screening software to shortlist or reject candidates. US Chamber of Commerce 2025 hiring data found nearly 20% of talent acquisition professionals would reject a candidate over an AI-generated resume or cover letter, with 14.5% saying AI shouldn’t be used at any stage of an application.
- Video and image platforms using AI detection to flag or demonetize content suspected of being AI-generated, often without disclosing what triggered the flag or offering a clear appeal.
- Social media and self-publishing platforms trusting AI to decide whether a piece of writing counts as human, as Substack did in July 2026 with its Pangram integration, letting readers scan posts, notes, and comments for an estimate of AI involvement.
- Traditional book publishers using AI detection to screen manuscripts before signing or after backlash, as Hachette did with Mia Ballard’s novel Shy Girl in March 2026, withdrawing the book from both UK and US release following an AI-slop controversy.
Every one of these institutions is treating a probability score as settled fact. Few have stopped to ask how often that score is wrong, or who pays the price when it is.
Why AI Detectors Make Mistakes
Are AI detectors accurate? The honest answer is: more than they used to be, but not accurate enough to be treated as proof.
An AI detection score is spectral evidence: invisible, unfalsifiable, and authoritative only because someone with power says it counts.
Detectors were trained on a mix of human and AI text, and the patterns they learned to call “machine-like” often overlap with patterns that already existed in careful human writing, long before generative AI existed.
The newest commercial detectors, and Pangram in particular, are a real, measurable improvement over the tools that produced the false-positive scandals of 2023 and 2024. But better does not mean totally accurate. The improved numbers describe a best-case scenario.
The moment the question shifts to the messier, more common scenario, a human draft that was edited or partially rewritten with AI assistance: even Pangram’s own published accuracy drops to 73.0%, meaning the best detector on the market gets the call wrong roughly one time in four in the exact situation most people are actually in.
But even a false-positive rate as low as 0.5% turns into a large number of wrongly flagged people the moment it is applied across a big enough population.
Comparison: Reported AI Detector False-Positive Rates
| Tool | Reported false-positive rate | Source and date | Known limitation |
|---|---|---|---|
| Pangram | 0.19% on raw AI text; at or below 0.5% on medium and long text | Vendor figures; Jabarian and Imas, Chicago Booth BFI WP 2025-116, August 2025 | Accuracy drops to 73.0% on three-way classification, distinguishing human from AI-edited from fully AI-generated text |
| Copyleaks | 0.02% claimed | Vendor claim | No independent replication at that figure |
| Turnitin | Approximately 1% claimed | Vendor claim, ongoing | Turnitin’s own guidance states the score should not be the sole basis for adverse action against a student |
| GPTZero | At or below 1% on medium-to-long passages | Jabarian and Imas, Chicago Booth BFI WP 2025-116, August 2025 | Higher error rates on short passages |
| OpenAI classifier | 9% false positive at 26% true positive | OpenAI, retired 2023 | Withdrawn by its own maker |
The following are examples of what AI detection can get wrong:
- Non-native English speakers and neurodivergent writers. Detectors are more likely to flag writing that follows a second language learner’s formal, rule-based style, or that reflects the systematic structure, consistent formatting, or precise repetition common among autistic and ADHD writers.
- Professional and academic writing habits. Detectors flag low variation in sentence length, consistent grammar, and formal structure as signs of AI, the exact traits that get rewarded in academic papers, business reports, and technical documentation. A lawyer, a scientist, or a student who was taught to write cleanly is, statistically, writing in a way that resembles what the model was trained to catch.
- Overuse of certain punctuation and phrasing patterns. Writers who favor em dashes, semicolons, or tightly structured paragraphs get flagged more often, not because those habits are unusual, but because AI-generated text also leans on them heavily, since the models learned them from the same well-edited human writing in the first place.
- Editors and proofreading tools. Running your own writing through a grammar or style checker like Grammarly can shift sentence structure and word choice in ways that nudge a detector’s score higher, even when every idea and word originated with you.
- Genre and formatting conventions. Structured formats like FAQs, numbered steps, or bullet-heavy content, the exact things search engines and readers reward, resemble the templated, low-burstiness structure AI models tend to output, making SEO-optimized writing an unintentional target.
None of this is theoretical. Every category above describes a real person who did nothing wrong and got flagged anyway.
What Are the Risks of AI Detection False Positives?
Not every AI accusation can be proven, and not every accuser cares if they’re correct. That’s what makes false positives dangerous: there’s no visibility into the evidence behind a score, no way to examine or challenge how it was calculated, and frequently no appeal at all. The accused has to disprove a number they were never shown the reasoning for, argued in front of someone who already believes the machine, with no formal process guaranteeing anyone reviews the case a second time.
Here’s what that cost actually looks like in real life.
- Academic Consequences: Students face misconduct investigations, failed assignments, delayed graduations, and permanent notations on their records. The investigation itself is often the punishment, dragging on for months while transcripts sit frozen and applications go out with holds attached. Even a case that gets dismissed can cost a semester.
- Career and Hiring Damage: Job seekers get filtered out before a human reads a word of their application. There’s no notification, no appeals process, and no way to know it happened. A cover letter written entirely by hand can be rejected for sounding too polished, and the applicant never learns why the callback never came.
- Professional Credibility and Reputational Damage: A public accusation moves faster than any defense. Comment sections, review platforms, and social feeds pick it up within hours, and the burden lands on the accused to prove a negative in front of an audience that has already decided. Clients cancel, work gets pulled, and the accusation follows the person to the next project long after any retraction the original accuser bothers to issue.
- Self-Censorship and the Authenticity Tax: The fear of a false flag pushes people to write worse on purpose, stripping out punctuation they like, roughening clean syntax, and adding errors just to seem more human. The authenticity tax is the unpaid, uncompensated extra steps of proving you’re human.
- Erosion of Trust in Both Directions: Readers start scanning for AI tells instead of reading. Teachers approach students as suspects rather than learners. Editors second-guess contributors they have worked with for years. The relationship changes even when nobody gets formally accused of anything.
The consequences above aren’t hypothetical, and they don’t require malice to happen. A tired teacher, an overwhelmed hiring manager, a platform running a scan at scale, none of them need to be acting in bad faith for the damage to land.
How to Protect Yourself From an AI Accusation
You can’t control whether a detector flags your work. You can control how prepared you are when it does. The steps below aren’t about beating the algorithm, they’re about building a record that speaks for you before anyone else gets to decide what your work means. It’s the same principle behind personal AI governance: deciding your own terms with these tools before someone else decides them for you.
1. Decide Your Boundaries Ahead of Time
Some accusations, especially in personal contexts or on social media, don’t deserve a forensic defense. Deciding in advance where you’ll actually engage and where you’ll simply walk away makes the boundary much easier to hold in the moment. That includes the tools themselves: on platforms that let you opt out of a scan, declining isn’t hiding anything, it’s just refusing to hand a probability score authority over your credibility.
2. Keep Your Drafts and Version History
Write in a tool that timestamps automatically, like Google Docs, Word Online, or even a version-controlled folder, so the record builds itself without any extra effort on your part. It’s the single most effective thing you can do, and it’s the evidence that has actually cleared people in real disputes.
3. Disclose Your Workflow Before You’re Asked
A short “how I make this” note, whether that’s an about-page statement, a footer line, or something you say upfront in a submission, converts a future accusation into something you already addressed. Disclosure offered voluntarily reads very differently than an explanation given under pressure, since it shows you were transparent.
4. Don’t Let One Platform Hold Your Whole Reputation
Diversifying where your work lives means one company’s scanning policy, or one bad call, doesn’t define you everywhere at once. Spread the risk the way you’d spread any other exposure you can’t fully control, so a single flag never becomes a single point of failure.
5. Investigate the Process
If falsely accused of AI usage, determine who reviewed it, using what tool, and if there is an appeals process. A single, direct written question puts the burden back where it belongs, and often exposes that there wasn’t a real investigation behind the accusation at all, just a score someone treated as a conclusion without ever checking it against anything else.
Closing Spell: Receipts Before Reckoning
Nobody sane believes floating proves witchcraft anymore. But a percentage on a screen still gets treated like evidence, dressed up in the language of statistics instead of superstition.
You don’t have to win a fight with an algorithm. You just have to make sure that if anyone ever asks, the record already answers for you.
Keep your drafts. Name your process. Decide in advance what deserves your defense and what doesn’t. That’s the difference between spending your energy on your work or spending it proving your work is human.
The witch hunt only works on people who show up without receipts.
Want a framework for handling risks like this one? Start with the Personal Risk Management Framework.
If you want to be notified of new Cyber Risk Witch posts directly, join the mailing list below.



