🔎 AI and email

AI text detectors work. What that changes for your professional emails.

A British services company, twenty-two employees. Every evening the managing director answers around fifteen messages: a client disputing a line on an invoice, a supplier pushing a delivery back, a project manager waiting on a decision. For the past year he has been using an AI to rough out his replies. Until now, whether that showed was a matter of intuition.

That has changed. In August 2026, the journal Nature devoted an article to the leap in reliability of tools that detect AI-generated writing, and an independent lab, Epoch AI, published its own measurements on three of them. The debate is no longer whether these tools work. It is what a business owner does with that fact.

Quick answer: an AI text detector applied to a professional email proves nothing, it estimates a probability. The independent measurements from Epoch AI (June 2026) show 0 false positives across 495 texts written by humans for two detectors out of three, but 10 to 18 per cent of AI texts going unspotted depending on the detector when the generation imitates the style of a specific author. The conclusion for a business owner: the only tenable answer is not to escape detection, it is to send writing that is genuinely yours.

Can an AI text detector spot a professional email?

An AI text detector applied to a professional email does not say whether you used an AI: it estimates a probability. The independent measurements from 2026 show that two detectors out of three wrongly accused no human text, but that they let through 10 to 18 per cent of texts imitating a style.

Two caveats matter before we go further. First, these tools were evaluated on long texts, in the region of five hundred words. An eight-line email offers far less statistical material, and none of the measurements cited here relate to emails. Second, no client runs their messages through a detector. What makes the subject relevant is what these measurements reveal about the way machine writing differs from human writing.

Vendor figures and independent measurement: two sets to read side by side

Vendors publish results from their own in-house benchmarks. Pangram claims, on its internal tests, a false positive rate of roughly 1 in 10,000, and 0.004 per cent (1 in 25,000) on academic writing. For Pangram 4, released in July 2026, the vendor announces 0.34 per cent false negatives, and 2.9 per cent when the AI imitates an author's style, figures reported by Nature. GPTZero, for its part, claims 99.6 per cent accuracy and 0.13 per cent false positives on its own benchmark. Pangram itself acknowledges that it cannot run the same protocol on its competitors' products.

Epoch AI carried out an independent evaluation in June 2026 on three detectors. Across 495 texts written by humans and published before 2022 (blog posts, fiction, scientific writing, around 500 words each), Pangram and GPTZero produced no false positives; Originality.ai produced 19, or 3.84 per cent. Across 297 passages generated by asking the AI to imitate the style of a specific author from five samples, Pangram misses 30 (10.10 per cent), GPTZero 32 (10.77 per cent) and Originality.ai 53 (17.85 per cent).

DetectorWhat the vendor claimsEpoch AI measurement, June 2026Version tested
Pangram Around 1 false positive in 10,000 on internal tests, 0.004 per cent on academic writing. Pangram 4: 0.34 per cent false negatives, 2.9 per cent under style imitation 0 false positives across 495 human texts. 30 AI texts unspotted out of 297 under style imitation (10.10 per cent) 3.3.2
GPTZero 99.6 per cent accuracy and 0.13 per cent false positives on its own benchmark 0 false positives across 495 human texts. 32 AI texts unspotted out of 297 (10.77 per cent) Model 2026-05-11-base
Originality.ai No figure reported in the sources cited here 19 false positives across 495 human texts (3.84 per cent). 53 AI texts unspotted out of 297 (17.85 per cent) Turbo 3.0.2

The nuance that rules out picking a winner. The vendor figures relate to the latest version of the detector; the independent measurement relates to the previous one. Pangram 4 came out in July 2026, Epoch AI tested version 3.3.2 in June. That gap is not a scheduling accident: a third-party evaluation will always be one release behind the product it evaluates. It is a reason to read both sets with care, not to crown a winner.

The blind spot: writing that imitates a style

The most interesting result from Epoch AI is not the false positives, it is the texts that go unspotted. The average across the three detectors sits at around 13 per cent of misses as soon as the AI has been given samples of an author's style before writing. On scientific writing in particular, around 26 per cent of style-imitating passages slip through. That register is highly codified, and one explanation can be read into this: the more a genre rests on shared conventions, the fewer individual markers a text carries. An executive's email is the opposite, it carries only yours.

Put differently: what trips a detector up is not a writing trick, it is the presence of a real personal style in the text. That finding is often read as a weakness of the tools. For a business owner, it says something far more concrete about their own emails. Writing that carries strong individual markers looks like what a human writes, because that is precisely what separates a human from a statistical average.

What this changes for an SME owner

Three practical consequences, and none of them is technical.

None of this sets the use of AI against the quality of the client relationship. What does set them against each other is use without review. Sorting the volume upstream is a separate subject, that of email productivity method.

The only tenable answer: send writing that is genuinely yours

If detection works, trying to get around it is a dead end: it means chasing models that improve faster than the tricks. The tenable answer lies elsewhere, and it is simpler. An email that carries your wording, your level of detail, your way of announcing bad news or setting a deadline does not have to defend the claim that it is yours. It is.

That is the logic Neston is built on: the assistant learns its user's writing style from around 300 sent emails and 500 received emails, then proposes a reply that the human always approves before it is sent. This is not a countermeasure to detection, and it does not claim to be one. It is the direct consequence of the same finding: the only writing that holds up over time in front of a client is the writing that sounds like you.

What triggers suspicion without any tool at all

In the life of a small company, judgement forms without a detector, on reading. Three signals keep coming back.

  1. The shift in register. You usually write in five lines, with no opening formula, and suddenly a carefully crafted introductory paragraph arrives. A regular correspondent notices the contrast immediately.
  2. Length out of proportion. A closed question that receives four paragraphs of context gives the impression that nobody read the question.
  3. The absence of case details. No order number, no date of exchange, no reference to what was said the week before. These are exactly the elements a machine cannot invent in your place, and the ones that make a message credible.

These three signals do not measure the use of AI. They measure care. That is also what writing scoring applied to emails tries to make objective.

Neston is in early access.

The assistant plugs into Outlook, learns your style from your emails and prepares a reply that you read and approve before it is sent. Early access is free, on a waiting list.

Join the waiting list →

Windows 10/11 · Outlook · Optional Mistral EU

Further reading

FAQ: AI text detectors and professional emails

Can a client really tell that I used AI to write my email?
They cannot know: at best they can suspect it, or run your text through a detector that will return a probability, not proof. The independent measurements published by Epoch AI in 2026 show that two of the three tools tested wrongly accused none of the 495 human texts submitted, the third being wrong in 3.84 per cent of cases, but that 10 to 18 per cent of passages written by an AI asked to imitate an author's style escape them. In practice, a client does not analyse their emails. What alerts them is a gap with the way you usually write to them.
Is it a problem to use AI to write a professional email, if I read it over and approve it?
No, provided the text you send is genuinely yours: your wording, your level of detail, your commitments. What your counterparts judge is not the tool you opened, it is the message they receive and the person who signs it. An email that has been read over, corrected and owned commits its sender exactly as much as an email typed from start to finish. The problem only appears in the opposite case: a generic text sent without review, which looks like neither you nor your case.
What triggers suspicion in an email, even without any detection tool?
Three signals keep coming back: a register that is too uniform and very different from your previous messages to the same person; a length out of all proportion to the question asked; and the absence of details specific to the case, the ones no machine can know in your place. The email that draws attention is the one that could have been sent to anyone. Conversely, a short, precise message anchored in a real exchange triggers no suspicion at all.
A note on reading these figures. The rates quoted by vendors come from their own protocols and are not comparable with one another. The Epoch AI measurements relate to texts of around 500 words published outside a professional context, not to emails. No figure in this article makes it possible to predict what a detector would return on a short message.
YB
Yvan Bosser
Founder of Neston · Ex-founder of Comptasanté (IK Partners exit 2023)
Yvan designs Neston, the AI email assistant integrated with Outlook, based on his own experience as an executive. Contact: yvan@neston.fr · LinkedIn.

🔬 Sources

Published 31 August 2026 · Reading time: 6 minutes · approx. 1,400 words