Are AI Detectors Accurate? Why They Disagree, and Who Gets Flagged by Mistake

Short answer: not accurate enough to treat as proof. An AI detector returns a probability, not a fact. Different tools give different scores for the same text because they are trained on different data and set their cutoffs differently. They also make more mistakes on plain, formulaic or second-language writing, so a human can be flagged for work they wrote themselves. Use a score as a reason to look closer. Never use it as a verdict on a person.

Why detectors disagree with each other

Paste one email into three detectors and you may get three different answers. That is not a bug in one of them. It follows from how the tools are built.

  • Different training data. Each tool learns what "AI-like" looks like from its own collection of human and machine text. A tool trained mostly on older model output may miss newer output. A tool trained on mostly formal essays may misjudge a casual LinkedIn post.
  • Different thresholds. One vendor may call anything above a modest score "likely AI." Another may wait until it is very sure. Same underlying signal, different line in the sand.
  • Different signals. Many detectors lean on how predictable the word choices are. Others use a classifier trained to separate the two kinds of text. Some combine both. Predictability is a weak proxy for authorship, because plenty of people write predictably.
  • Text length. Two sentences give a detector almost nothing to work with. A short cold email is exactly the kind of text where scores swing the most.
  • Editing. Light edits to an AI draft, or heavy smoothing of a human draft with a grammar tool, move the text toward the middle where detectors are least sure.

The practical upshot: if two detectors disagree, neither has told you much. If they agree, you have slightly more to go on, but still not proof.

Why non-native writers get flagged more often

This is the part of the problem that does real harm, so it is worth being precise about the mechanism.

A predictability-based detector asks, roughly, "would a language model have picked these words?" Writers working in a second language often choose common words and safe sentence patterns. That is a sensible strategy. It keeps meaning clear and avoids awkward idioms. But it also produces text with low surprise, and low surprise is the same pattern detectors associate with machines.

Several habits push in the same direction:

  • Template phrases learned from textbooks and business-English courses, such as "I am writing to inform you that."
  • Translation tools and grammar checkers that normalise phrasing toward the most common version.
  • Short, even sentences used to stay safely grammatical.
  • Few personal details or odd specifics, because those are harder to express precisely.

The result is that a careful, honest writer can look more "machine-like" to a detector than a native speaker who writes loosely. If you hire, screen applications or review outreach, this matters. A flag on a resume or cover letter may say more about the writer's language background than about their tools. We cover the broader pattern in why AI detectors flag human writing.

The two kinds of error

ErrorWhat happensWho pays
False positiveHuman writing is labelled as AIThe writer: a rejected application, a lost deal, a damaged reputation
False negativeAI writing is labelled as humanThe reader or reviewer, who trusts a clean score

Vendors can tune a tool to reduce one error, but usually at the cost of the other. A stricter setting catches more AI text and wrongly accuses more humans. A looser one does the reverse. Ask any vendor which error they optimised for, and what their false positive rate is on writers like yours.

How to use a detector score sensibly

  1. Test a long enough sample. Run the whole email, post or page, not one paragraph. Short snippets give unstable results.
  2. Run it twice, on two tools. Large disagreement tells you the result is noise.
  3. Look at which passages get flagged. Highlighted sections are more useful than a single percentage. Flagged passages are usually the generic ones.
  4. Check the draft history. Version history in Google Docs or Word, notes and earlier drafts are better evidence of authorship than any score.
  5. Revise for substance, not for the score. Add the specific detail only you know: a number, a client name you can share, a decision you made and why.
  6. Stop chasing zero. A zero score is not a goal, and rewording until you hit it often makes writing worse.

An example: same meaning, different risk

These two openings are ones I wrote for illustration. They are not test results from any tool.

Generic: "I am writing to express my interest in the Operations Manager position. I have extensive experience in managing teams and improving processes."

Specific: "I ran warehouse operations for a 40-person team in Lagos for six years. When we moved to a new inventory system, I rewrote the pick-and-pack steps myself and trained every shift."

The first is grammatical and polite, and it is also what thousands of applicants and chatbots produce. A detector has every reason to be suspicious of it. The second is harder to mistake for template text, mostly because it contains things only this person could say. That is a better defence than any tool setting, and it is a better resume too. See our notes on words that make writing sound like AI.

Common mistakes

  • Treating a percentage as a measurement. "72% AI" does not mean 72% of the words were machine-written. It is usually a confidence score.
  • Using one detector as a gate. Any single tool will be wrong some of the time, and you will not know when.
  • Running polished text through more grammar tools to "fix" a flag. Extra smoothing often makes text more predictable, not less.
  • Accusing someone on a score alone. Ask for drafts and talk to them first.
  • Swapping synonyms to dodge detection. It tends to produce odd phrasing that readers notice even if the detector does not.

If you write in English as a second language

  • Keep your drafts and timestamps. Save a copy before you run any grammar tool.
  • Vary sentence length on purpose. Follow a long sentence with a short one.
  • Replace stock phrases with a concrete detail from your own work.
  • Write the first draft in your own words, then fix grammar after, not during.
  • If a flag causes a problem, explain how you write and offer the draft history.

We have a page for this situation: Plainspoke for non-native writers.

Where Plainspoke fits, and where it does not

Plainspoke is an AI writing checker and humanizer for professional writing: cold email, LinkedIn, resumes and product pages. It points out the generic, predictable passages in a draft and helps you rewrite them so the text sounds like a person with something to say. It does not read minds, and it cannot guarantee any outside detector will give a particular score. No tool can honestly promise that, given everything above. It is also not meant for schoolwork.

If you want to compare approaches, see how we differ from GPTZero and Originality.ai. For a rewriting workflow, start with how to make AI writing sound human.

Check a draft free

Frequently asked questions

Are AI detectors accurate?

Not reliably enough to prove anything. They estimate probability from patterns in the text. Results vary by tool, text length and writing style, and every tool produces both false positives and false negatives.

Why do two AI detectors give different scores on the same text?

They use different training data, different signals and different cutoffs for calling text AI. Short text and lightly edited text make the disagreement larger.

Do AI detectors flag non-native English speakers more often?

It is a real risk. Detectors that rely on predictability can read simple vocabulary and template phrasing as machine-like, and second-language writers often use both. Treat a flag on their work with extra caution and ask for draft history.

Can a detector tell if I used ChatGPT?

No, not with certainty. It can only say the text resembles patterns it has learned. Heavy editing lowers that resemblance, and plain human writing can raise it.

What should I do if my own writing is flagged as AI?

Keep your version history and earlier drafts, run the text through a second tool, and add specific details only you would know. If someone is accusing you, offer the drafts as evidence.

Is a 0% AI score a good target?

No. Chasing zero tends to make writing stiffer. Aim for clear, specific writing a reader would trust, and treat the score as a rough hint.

Paste a draft and see what reads as generated

Free account, no card. Checks and rewrites included.