Detector scores are the wrong first screen
A detector score is useful evidence only after an editor has read the prose. As the first screen, it asks the weakest question: “Can a classifier guess the author?” The editorial question is better: “What shape does the draft have?”
Run the math on a small queue. If 100 machine drafts arrive, a detector that tops out at 80% correct catches 80 and misses 20. If 100 human drafts from non-native English writers arrive, a false-positive rate as high as 70% wrongly marks 70 people as suspect. That is not a workable gate for hiring, grading, or rejecting a freelancer’s copy.
Human instinct is not a safe substitute. In one peer-reviewed evaluation, human raters separating human and AI-generated text reached 19% overall accuracy. Even a confident editor can miss it.
The better first pass is structural. Mark the draft, not the person. Count paragraph shapes. Look for the explanation that names its own lesson after already proving it. Check whether the piece moves in a neat straight line from definition to list to conclusion. Circle abstract claims made through bodily metaphor. Ask whether named evidence appears at all: the 30-article sample, the measured feature, the reported accuracy ceiling, the false-positive rate. Then scan sentence rhythm. Detector vendors often call this “burstiness”; the plain version is simpler. If sentence after sentence sits in the 12-to-18-word band, the prose has been planed flat.
Version history is better evidence than a single score: Google Docs saves revisions after every small edit, so a draft with visible iteration tells you more than a detector readout. Start where the pattern shows on the page.
The five structural signs worth checking in a live draft
Use the five signs as an editing pass, not as a verdict. Take a live paragraph, mark what it does, and then decide whether the pattern is strong enough to matter.
Example draft:
AI-written text can affect content quality in several ways. It often explains ideas too clearly and makes the writing feel less human. AI writing usually follows a predictable structure, starting with a definition, moving into benefits, and ending with a summary. Strong content needs specific examples, varied rhythm, and clear evidence. The result is writing that readers trust more.
That paragraph has no obvious grammar problem. It is also easy to mark up.
| Sign | What the editor marks | What it looks like in the example |
|---|---|---|
| Over-explanation | A sentence names the lesson after the paragraph has already made it | “The result is writing that readers trust more.” |
| Straight-line organization | The paragraph moves from general claim to abstract list to tidy close | “starting with a definition, moving into benefits, and ending with a summary” |
| Bodily metaphor | An abstract claim is explained through a human-body image | Not present here. Do not force the mark. |
| Missing named evidence | The paragraph gives advice without naming the study, sample, measure, tool, or observed feature | “specific examples” and “clear evidence” are requested but not supplied |
| Paragraph-level uniformity | Sentences sit in the same explanatory mode and do similar jobs | Claim, explanation, explanation, prescription, wrap-up |
Change the shape.
A better version would start where the draft has evidence, not where it has a theme:
The signal the detectors can actually use is structural, including sentence-length variance rather than a favourite phrase or giveaway adjective. The study compared human and machine-produced article introductions drawn from 30 open-access journal papers published before 2022. For an editor, that points to a practical first pass: count the paragraph moves before arguing about style. Does the draft explain its own lesson twice? Does every section proceed from definition to list to summary? Does it use body-based imagery for abstract systems? Does it name evidence precisely, or gesture at “studies” and “experts”? Are the sentences different lengths because the thought changes, or merely shuffled within the same template?
That revision still has a list, but the list is doing work. It is tied to a source, a sample, and an editing action. The body-metaphor check is also deliberately narrow: one metaphor is not a conviction. A cluster is worth cutting.

What the study is actually saying, and what blog summaries keep flattening
Do not outsource that call to a detector. Even the better tools top out at 80 percent accuracy, and one evaluation found false positives for non-native English speakers as high as 70 percent. Do not use it as a courtroom test for authorship.
A clean rule works better than a suspicion. If a piece has one odd phrase, edit the phrase. If it has repeated explanatory endings, a straight march from definition to list to recap, abstract body metaphors, unnamed evidence, and paragraphs built to the same length, review it as AI-written text until the writer can show the reporting, reasoning, or drafting history behind it.
The common overclaim is that the research gives editors a detector. It does not. It gives editors a better inspection method. A detector promises identity: human or machine. Structural review asks a narrower question: does the text have the shape associated with machine drafting, and does that shape weaken the work?
There is a real objection here. Disciplined institutional writing can look regular because house style demands it. Fair. That is why the test should trigger revision, not accusation. Ask for named evidence, cut the self-explaining sentences, vary paragraph function, and see whether the draft improves.
A detector can still be useful, just not as the verdict
The strongest objection is practical: editors do not always have time to perform close structural review on 80 submitted drafts, 300 archived pages, or 12,000 words from a contractor. A detector gives a fast signal. Dismissing that signal because it is imperfect is lazy in the other direction.
ZeroGPT, PhraslyAI, and Grammarly AI Detector separated machine-written, machine-edited, and human-written text consistently enough to show they were picking up real structural differences rather than noise. Agreement across the detectors ranged from merely moderate to extremely strong, broad enough to matter and far too consistent to dismiss as noise. AI detectors commonly score text from structural signals such as burstiness and sentence-length variance, then return a probability estimate. AI detectors top out at about 80 percent accuracy. These tools are measuring something real enough to help sort a queue.
Use a detector the way an editor uses a plagiarism match percentage: as a reason to inspect, not as the finding itself. A 78 percent AI-likelihood score should put the draft in the review pile. It should not produce an accusation, a rejected invoice, or a note to legal. The next step is structural: mark the explanatory endings, count the unnamed claims, look for a definition-list-recap march, flag abstract body metaphors, and compare paragraph function across the piece.
A workable rule:
| Detector result | Editorial action |
|---|---|
| Low score, clean structure | Edit normally |
| High score, clean structure | Ask for drafting notes or sources before escalating |
| Low score, machine-shaped structure | Revise the structure anyway |
| High score, machine-shaped structure | Treat as failed copy until repaired |
The third row matters most. A detector can miss a polished AI-assisted draft, and it can punish a cautious institutional writer whose prose has been flattened by approvals. The reader does not care which failure produced the page. It still fails without named evidence, and it still fails if every paragraph lands in the same measured stride.
So keep the detector. Put it at intake. Do not put it in the judge’s chair.
A review process an editor can run in ten minutes
Use detectors at intake for triage, but make structural review the first real editorial screen, because detection tops out around 80 percent accuracy and can misfire on non-native English writing at rates as high as 70 percent. A careful human can write bland sentences; a machine-drafted page usually leaves a more visible pattern.
| Minute | Mark | Editorial test |
|---|---|---|
| 0-2 | Claims without names | Circle every assertion that lacks a named source, study, person, product, place, or document. “Research shows” fails. “Harvard University Information Technology requires review of AI-generated content before publishing” passes. |
| 2-4 | Explanatory overhang | Cut the sentence after the example if it only tells the reader what the example already proved. |
| 4-6 | Shape | Draw the section path in the margin: definition, list, expansion, recap. If every section follows that path, send it back. |
| 6-8 | Metaphor and evidence | Flag abstract body metaphors. Replace them with the actual mechanism, source, or observed failure. |
| 8-10 | Revision history | Open Google Docs Version History. A document with one pasted block and no sourcing trail is not disqualified, but it needs drafting notes before it gets trust. |
Ignore detector drama during this pass. Also ignore single words that have become fashionable AI tells. Detection tools top out well short of certainty, so structural cues matter more than a verdict from software.
For AI-assisted work, require a short production note: who drafted it, what tool was used, what sources were checked, and what the human editor changed. That is why detector scores are a weak standard: human raters reached 19% accuracy, automated tools top out at 80%, and false positives for non-native English writers can hit 70%.
Publishable copy has named evidence, uneven human decisions, and revisions a reader would benefit from. Anything else is still raw material.
Key Takeaways
A detector score is weak evidence on its own. In controlled testing, these tools topped out at 80% accuracy, and false positives for non-native English writers ran as high as 70%. Treat them as triage, then check drafts against named sources, revision history, and citation integrity before you make the call.
| Structural sign | Weak review habit | Better editorial rule |
|---|---|---|
| Over-explanation | Flagging polished summaries | Cut the sentence that explains the example after the example already works. |
| Straight-line organization | Rewarding tidy order | Ask whether the piece ever enters through a case, exception, failure, or unresolved tension. |
| Bodily metaphor | Letting abstract metaphors pass | Replace vague physical imagery with the actual mechanism, source, or editorial decision. |
| Missing named evidence | Accepting “experts say” | Require named studies, products, documents, dates, or measurements where the claim depends on proof. |
| Paragraph uniformity | Counting grammar errors | Scan rhythm, section length, dissent, and revision marks. Perfectly even prose deserves suspicion. |
Frequently Asked Questions
How can I tell if text was written by AI?
Look at the structure before you look at the vocabulary. AI-written text often explains its own point after making it, moves in a tidy sequence, uses vague physical metaphors for abstract work, avoids named evidence, and keeps paragraphs unnaturally even. One odd phrase proves little. A repeated structural pattern is harder to dismiss.
Is AI-written text always wrong?
No, AI-written text is not always wrong, but it is often under-verified. The risky parts are citations, dates, named sources, recent events, and claims that depend on exact documents. A fluent paragraph can still contain a source that does not exist.
Can AI tools make up citations?
Yes, AI text-generation tools can produce non-existent citations and source references. Treat every book title, paper, author name, journal, and URL in machine-drafted copy as untrusted until checked. The failure is worse than a typo because it can make invented support look academically neat.
Can AI-written text include current information?
Only if the tool has access to current retrieval or the writer supplies current material. A model without real-time internet access cannot reliably include events after its training cutoff. For live topics, the review step should include date checks, source opening, and confirmation that the cited page says what the draft claims.
Do marketers use AI-written text?
Yes. AI text detectors top out at about 80 percent accuracy. That figure matters because the practical editorial question is no longer whether AI touched the draft. The question is whether a person added evidence, judgment, and revision strong enough to make the piece publishable.
What is the best way to edit AI-written text?
The best edit is to rebuild the evidence and shape, not to swap words for less common synonyms. Ask for the claim in one sentence, attach a named source or measurable example, cut the explanatory after-sentence, and vary the paragraph rhythm where the prose has gone flat. If the draft cannot support its claims without vague attribution, do not polish it. Rewrite the claim or remove it.



