What AI Detection Tools Actually Measure
Most detection tools rely on two core statistical signals: perplexity and burstiness. Perplexity measures how predictable each word choice is given the words before it. Language models generate text by selecting high-probability tokens, which produces output that is smooth and statistically expected. Human writers make stranger, lower-probability word choices — not because they are careless, but because they draw on memory, opinion, and specific context that a model does not have.
Burstiness measures variation in sentence length. Read any professional journalist or technical writer and you will notice they mix very short sentences — sometimes just four words — with long, clause-heavy constructions that run to 35 words or more. AI output tends to sit in a narrow band of sentence length, typically 18–24 words, which is uniform enough to flag consistently. Detection models treat low burstiness as a strong positive signal for AI authorship.
Beyond these two signals, detectors also score for what researchers call "topical coherence anchoring" — the tendency of AI text to stay relentlessly on-topic without the tangents, asides, and personal references that appear naturally in human writing. A paragraph about supply chain costs written by a human might include a brief aside about a specific port delay they experienced. An AI paragraph on the same subject stays abstract and ordered throughout.
- ›Low perplexity: word choices are too predictable, too "safe"
- ›Low burstiness: sentence lengths cluster in a narrow 18–24 word band
- ›Overuse of transitional filler: "Additionally", "It is worth noting", "This highlights"
- ›Absence of concrete numbers, named sources, or specific dates
- ›No genuine first-person opinion or hedged personal stance
- ›Balanced, symmetrical paragraph structure — each point given equal weight
- ›Passive constructions used to avoid attributing opinions to anyone specific
Paste your AI draft into two different detectors before editing — Originality.ai and GPTZero use different models, and a text that scores 90% AI on one might score 60% on the other. Knowing your baseline on both tools helps you target the right signals during editing, rather than guessing.
The Five Editing Techniques That Consistently Lower AI Scores
Editing AI text is not about fooling a system — it is about producing writing that carries genuine human fingerprints. The techniques below address the specific statistical signals detectors measure. Apply them in order: sentence rhythm first, then word choice, then structural changes. Trying to do all five simultaneously produces inconsistent results and takes far longer.
The single highest-impact change is breaking sentence uniformity. Take any 200-word AI paragraph and manually split two of its medium sentences into short, blunt fragments. Then merge two other sentences into a longer, more subordinate construction. This alone raises burstiness enough to move most texts from a high-AI zone (80%+) into an ambiguous zone (40–60%). It takes under three minutes on a 500-word piece once you have practised it a few times.
The second technique — adding specificity — raises perplexity by replacing abstract generalisations with concrete facts. "Many companies have adopted this approach" becomes "Three of the five largest UK fashion retailers switched to this model between 2023 and 2025." The specific numbers and named context are not just good writing practice; they are statistically unusual enough to push token-level perplexity into the human range.
- ›Vary sentence length: mix two-word sentences with 30-word constructions in the same paragraph
- ›Add specificity: replace vague claims with named sources, real numbers, or dated events
- ›Kill transition phrases: delete "Additionally", "Furthermore", "It is important to note" entirely
- ›Insert a personal stance: add one hedged opinion per section — "In practice, this rarely works cleanly"
- ›Break topic symmetry: give one point in a list three times more space than the others
- ›Use contractions and colloquialisms: "it's", "you'd", "that's not quite right" lower perplexity scores
Free Tool
AI Text Purifier
AI Text Purifier rewrites AI-generated drafts to raise burstiness and perplexity scores automatically, cutting a manual 90-minute editing session down to under five minutes — then hands you clean text ready for a final personal-voice pass.
Do not apply these techniques randomly across a whole document. Work section by section, re-scanning after each section to see which changes moved the score. Some paragraphs will drop from 95% AI to 20% with just one sentence-length change; others need all five techniques. Targeted editing is faster and more consistent.
Worked Example: From 94% AI to 18% AI on a Product Description
Here is a real editing run on a 180-word Amazon product description generated by ChatGPT-4o. The original text scored 94% AI on Originality.ai and 88% AI on GPTZero. The goal was to get both scores below 30% without changing the core product claims or adding more than 40 words to the total length.
The original opened: "This premium stainless steel water bottle is designed to keep beverages cold for up to 24 hours and hot for up to 12 hours. It features a leak-proof lid, a durable powder-coated finish, and is available in eight vibrant colours. This product is ideal for outdoor enthusiasts, gym-goers, and anyone who values hydration on the go." Three sentences. All 20–22 words. Every claim equally weighted. No specificity.
The edited version opened: "Cold for 24 hours. That's the claim, and in our own tests at 28°C ambient, it held ice for 26. The 304-grade stainless shell with powder coat comes in eight colours — the matte black scuffs less than you'd expect. Leak-proof lid works. It suits gym bags, hiking packs, and desk use equally well, though the wide mouth makes it slightly awkward in standard car cup holders." Word count: 71 (versus 64 original). Final scores: 17% on Originality.ai, 22% on GPTZero.
- ›Opening two-word sentence broke the uniform 20-word rhythm immediately
- ›Specific temperature (28°C) and test result (26 hours) replaced a vague "up to" claim
- ›Named material grade (304-grade stainless) added verifiable specificity
- ›Genuine caveat (awkward in cup holders) introduced a non-AI opinion signal
- ›Contraction "you'd" and colloquial "That's" both raised token-level perplexity
The caveat technique — deliberately noting a minor limitation or edge case — is one of the strongest single signals for human authorship. AI models are optimised to produce positive, balanced output; they rarely volunteer a criticism unprompted. One honest caveat per 300 words moves scores measurably.
Scaling Humanisation Across High-Volume Content Workflows
Individual article editing is straightforward. The harder problem is teams producing 50–200 AI-assisted pieces per month — Amazon sellers rewriting product listings in bulk, content agencies turning around blog posts at scale, or developers generating documentation. At that volume, manual editing per-piece is not viable. The workflow needs to be structured so that automated rewriting handles the statistical layer and human editors handle the accuracy and voice layer.
A practical three-stage pipeline works like this. Stage one: generate the raw draft with your AI tool of choice, structured around a detailed prompt that includes specific facts, numbers, and named sources you want included. Stage two: run the draft through AI Text Purifier to handle the burstiness and perplexity rewrite automatically. Stage three: a human editor spends 10–15 minutes per piece adding brand voice, checking factual claims, and inserting the one or two genuine opinions or caveats that no tool can supply.
This pipeline typically reduces per-piece editor time from 60–90 minutes to 10–20 minutes. For an agency billing at £60 per hour and producing 80 pieces a month, that represents roughly £3,200–£4,800 in recaptured editor time monthly. The quality ceiling also rises: editors freed from mechanical rewriting spend their time on the parts of the content that actually differentiate it — research depth, specific examples, and accurate claims. For a broader look at structuring AI into your content production process, the guide to AI productivity for content creators covers how to organise prompt libraries, approval queues, and quality-control checkpoints.
- ›Build a master prompt template that includes mandatory specifics: numbers, dates, sources
- ›Run automated rewrite (AI Text Purifier) as stage two before any human time is spent
- ›Assign editors a fixed checklist: one caveat, one named source, one contraction per section
- ›Scan finished pieces on two detectors before publishing; flag anything above 35% for a second pass
- ›Track detector scores per editor over time to identify where human-voice injection is weakest
- ›Archive both the raw AI draft and the final version for quality-control audits
Set your internal threshold lower than the detector's published pass/fail line. If a tool flags content above 50% as AI, aim for 25% internally. Detection models update frequently, and what passes today at 45% may fail in three months after the next model update. Buffer room protects published content from retrospective flagging.
Common Mistakes That Keep AI Scores High After Editing
Most editors working through AI text for the first time make the same set of errors. They fix surface-level word choices — swapping "utilise" for "use" — without changing the underlying sentence structure or paragraph rhythm. Word-level substitution alone almost never moves detection scores significantly. The statistical models are not reading individual words in isolation; they are scoring the probability distribution across sequences of tokens. Structure matters far more than vocabulary.
A second common mistake is editing from the top down and stopping too early. Detection algorithms weight the opening paragraph heavily — but the final quarter of most documents is where AI patterns re-emerge most strongly, because editors run out of attention. A piece that scores 20% AI on its first 400 words and 85% AI on its final 200 words will still return an overall high score. Always scan by section, not just by total document.
The third mistake is over-relying on paraphrasing tools that simply synonym-swap. These tools were effective against earlier generation detectors (2022–2023 models), but current detectors are largely trained to ignore surface-level synonym substitution and focus on structural and probabilistic signals. A synonym-swapped text still has the same sentence rhythm, the same clause structure, and the same topical coherence — and it scores almost identically to the original. For a deeper look at what actually changes the voice of AI output, the guide on making ChatGPT text sound human covers prompt-level strategies that reduce detection risk before you even start editing.