A British services company, twenty-two employees. Every evening the managing director answers around fifteen messages: a client disputing a line on an invoice, a supplier pushing a delivery back, a project manager waiting on a decision. For the past year he has been using an AI to rough out his replies. Until now, whether that showed was a matter of intuition.
That has changed. In August 2026, the journal Nature devoted an article to the leap in reliability of tools that detect AI-generated writing, and an independent lab, Epoch AI, published its own measurements on three of them. The debate is no longer whether these tools work. It is what a business owner does with that fact.
Quick answer: an AI text detector applied to a professional email proves nothing, it estimates a probability. The independent measurements from Epoch AI (June 2026) show 0 false positives across 495 texts written by humans for two detectors out of three, but 10 to 18 per cent of AI texts going unspotted depending on the detector when the generation imitates the style of a specific author. The conclusion for a business owner: the only tenable answer is not to escape detection, it is to send writing that is genuinely yours.
Can an AI text detector spot a professional email?
An AI text detector applied to a professional email does not say whether you used an AI: it estimates a probability. The independent measurements from 2026 show that two detectors out of three wrongly accused no human text, but that they let through 10 to 18 per cent of texts imitating a style.
Two caveats matter before we go further. First, these tools were evaluated on long texts, in the region of five hundred words. An eight-line email offers far less statistical material, and none of the measurements cited here relate to emails. Second, no client runs their messages through a detector. What makes the subject relevant is what these measurements reveal about the way machine writing differs from human writing.
Vendor figures and independent measurement: two sets to read side by side
Vendors publish results from their own in-house benchmarks. Pangram claims, on its internal tests, a false positive rate of roughly 1 in 10,000, and 0.004 per cent (1 in 25,000) on academic writing. For Pangram 4, released in July 2026, the vendor announces 0.34 per cent false negatives, and 2.9 per cent when the AI imitates an author's style, figures reported by Nature. GPTZero, for its part, claims 99.6 per cent accuracy and 0.13 per cent false positives on its own benchmark. Pangram itself acknowledges that it cannot run the same protocol on its competitors' products.
Epoch AI carried out an independent evaluation in June 2026 on three detectors. Across 495 texts written by humans and published before 2022 (blog posts, fiction, scientific writing, around 500 words each), Pangram and GPTZero produced no false positives; Originality.ai produced 19, or 3.84 per cent. Across 297 passages generated by asking the AI to imitate the style of a specific author from five samples, Pangram misses 30 (10.10 per cent), GPTZero 32 (10.77 per cent) and Originality.ai 53 (17.85 per cent).
| Detector | What the vendor claims | Epoch AI measurement, June 2026 | Version tested |
|---|---|---|---|
| Pangram | Around 1 false positive in 10,000 on internal tests, 0.004 per cent on academic writing. Pangram 4: 0.34 per cent false negatives, 2.9 per cent under style imitation | 0 false positives across 495 human texts. 30 AI texts unspotted out of 297 under style imitation (10.10 per cent) | 3.3.2 |
| GPTZero | 99.6 per cent accuracy and 0.13 per cent false positives on its own benchmark | 0 false positives across 495 human texts. 32 AI texts unspotted out of 297 (10.77 per cent) | Model 2026-05-11-base |
| Originality.ai | No figure reported in the sources cited here | 19 false positives across 495 human texts (3.84 per cent). 53 AI texts unspotted out of 297 (17.85 per cent) | Turbo 3.0.2 |
The nuance that rules out picking a winner. The vendor figures relate to the latest version of the detector; the independent measurement relates to the previous one. Pangram 4 came out in July 2026, Epoch AI tested version 3.3.2 in June. That gap is not a scheduling accident: a third-party evaluation will always be one release behind the product it evaluates. It is a reason to read both sets with care, not to crown a winner.
The blind spot: writing that imitates a style
The most interesting result from Epoch AI is not the false positives, it is the texts that go unspotted. The average across the three detectors sits at around 13 per cent of misses as soon as the AI has been given samples of an author's style before writing. On scientific writing in particular, around 26 per cent of style-imitating passages slip through. That register is highly codified, and one explanation can be read into this: the more a genre rests on shared conventions, the fewer individual markers a text carries. An executive's email is the opposite, it carries only yours.
Put differently: what trips a detector up is not a writing trick, it is the presence of a real personal style in the text. That finding is often read as a weakness of the tools. For a business owner, it says something far more concrete about their own emails. Writing that carries strong individual markers looks like what a human writes, because that is precisely what separates a human from a statistical average.
What this changes for an SME owner
Three practical consequences, and none of them is technical.
- The subject is not detection, it is recognition. Nobody is going to run your payment reminder through a detector. A client who has been reading you for three years, on the other hand, recognises your way of writing, and notices when it changes overnight.
- Generic is the real problem. What detectors spot best is also what a human reader finds most lukewarm: smooth, polite text that could apply to any case. That kind of message is not just detectable, it is of little use.
- Reading it over is the act that commits you. An email sent under your name commits you, whatever way the first draft was produced. That is why approval before sending is not a formality.
None of this sets the use of AI against the quality of the client relationship. What does set them against each other is use without review. Sorting the volume upstream is a separate subject, that of email productivity method.
The only tenable answer: send writing that is genuinely yours
If detection works, trying to get around it is a dead end: it means chasing models that improve faster than the tricks. The tenable answer lies elsewhere, and it is simpler. An email that carries your wording, your level of detail, your way of announcing bad news or setting a deadline does not have to defend the claim that it is yours. It is.
That is the logic Neston is built on: the assistant learns its user's writing style from around 300 sent emails and 500 received emails, then proposes a reply that the human always approves before it is sent. This is not a countermeasure to detection, and it does not claim to be one. It is the direct consequence of the same finding: the only writing that holds up over time in front of a client is the writing that sounds like you.
What triggers suspicion without any tool at all
In the life of a small company, judgement forms without a detector, on reading. Three signals keep coming back.
- The shift in register. You usually write in five lines, with no opening formula, and suddenly a carefully crafted introductory paragraph arrives. A regular correspondent notices the contrast immediately.
- Length out of proportion. A closed question that receives four paragraphs of context gives the impression that nobody read the question.
- The absence of case details. No order number, no date of exchange, no reference to what was said the week before. These are exactly the elements a machine cannot invent in your place, and the ones that make a message credible.
These three signals do not measure the use of AI. They measure care. That is also what writing scoring applied to emails tries to make objective.
Neston is in early access.
The assistant plugs into Outlook, learns your style from your emails and prepares a reply that you read and approve before it is sent. Early access is free, on a waiting list.
Join the waiting list →Windows 10/11 · Outlook · Optional Mistral EU
Further reading
- Email and AI: how to write with an assistant, the method on the day-to-day usage side
- Email productivity: the 2026 method, to handle the volume before talking about writing
- How an AI learns your email writing style, what happens before a reply is proposed to you
- Writing scoring: rating your emails from 0 to 100, the criteria that make a message inspire trust
- Do you have to disclose that an email was written with AI?, what Article 50 of the EU AI Act does, and does not, say
FAQ: AI text detectors and professional emails
🔬 Sources
- Nature, 25 August 2026, "AI-detection tools have made huge leaps forward, how good are they?", Nature 656, pages 808 to 811: progress of detectors, figures announced for Pangram 4 (0.34 per cent false negatives, 2.9 per cent under style imitation)
- Epoch AI, "AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors": June 2026 tests on Pangram 3.3.2, GPTZero 2026-05-11-base and Originality.ai Turbo 3.0.2, 495 human texts and 297 style-imitation passages
- Pangram, "All About False Positives in AI Detectors": false positive rates claimed by the vendor and the limits of its own comparison protocol
- GPTZero, comparison published by the vendor: 99.6 per cent accuracy and 0.13 per cent false positives claimed
Published 31 August 2026 · Reading time: 6 minutes · approx. 1,400 words