Every figure here carries the date it was read and the source it came from. How scoring works →

Scoring AI Detectors on the Accuracy Somebody Else Measured, Not the One They Advertise

4 tools measured · Vouch Score data collected 8 August 2026 · highest composite first

Every detector publishes an accuracy figure about itself. This category scores them on figures published by people with nothing to sell: academic papers that state their corpus, their sample size and their method, and can be opened and checked. Two numbers matter and vendors tend to quote only the first: how often a detector catches AI text, and how often it accuses a human. The second is the one that ends careers and academic appeals, so it is published beside the first here rather than behind it.

Top score: Copyleaks. It holds the highest Vouch Score composite on this page, 69.6/100, measured 8 August 2026. The label is computed from the cards, not chosen by us, and it moves the day another tool measures higher. Read what each card measured before you take it as advice for your own work.

1Copyleakstop score · highest composite69.6/100 · average

AI-text and AI-image detector bundled with a plagiarism checker, a grammar checker, text moderation and Codeleaks source-code detection, sold to individuals, schools and enterprises on one shared credit meter.

Copyleaks is an AI-text and plagiarism detector from Copyleaks Technologies LTD, sold on one shared credit meter. Its own page states over 99% accuracy and an industry-low .03% false positive rate. Its Terms of Use, last revised 9 September 2025, say the result may incorrectly flag content and does not constitute a determination that anything was AI-generated. Our composite reads 69.6/100, 8 August 2026.

Visit CopyleaksRead the full review →Personal $16.99/mo, $13.99 billed annually ($167.88) · Pro $99.99/mo, $74.99 billed annually ($899.88), 25 seats · Education and Enterprise: no published price, Talk to Sales only
2Originality.ai52.6/100 · low

AI-text detector, plagiarism checker, fact-checker and readability scorer sold to agencies, publishers and educators; scans are billed in credits, one credit per 100 words.

Originality.ai is a credit-metered AI-text and plagiarism detector with no free tier; our card reads 52.6 of 100 on a full card of four. A preprint scores it above the other commercial detectors in its own sample on clean Arabic articles, 92%, falling to 12% after light polishing. Its 85% RAID claim checks out, and the same paper records a 0.62% false-positive floor.

Visit Originality.aiRead the full review →Pro $14.95/mo, $12.95 billed annually (2,000 credits) · Enterprise $179/mo, $136.58 annually (15,000 credits) · NO free tier and no trial tile (originality.ai/pricing, 9 August 2026)
3GPTZero45.3/100 · low

GPTZero scores 45.3 out of 100, computed 23 August 2026. Asked how often it calls human writing AI, four sources answer differently: the vendor reports 1%, and peer-reviewed studies report 10%, 16%, and 52 false positives across 114 student essays written before ChatGPT existed. Its own Terms of Use disclaim all warranties about accuracy.

4Winston AI43.8/100 · low

Winston AI scores 43.8 out of 100, computed 2 September 2026 on a full card of four dimensions. In a 14-tool academic evaluation it produced no false accusations across 18 human-written and translated documents, and correctly classified one of the nine documents that had been rewritten with a paraphraser. Those two results are the same property measured twice.

Outbound links may be affiliate links and can earn us a commission: they never touch a score, and the order on this page is the composite order, computed at build time from the cards themselves.

The score cards in full

What each card measured and what it did not: every number dated, sourced and reproducible.

#ToolScore /100GradePricing
1
AI-text and AI-image detector bundled with a plagiarism checker, a grammar checker, text moderation and Codeleaks source-code detection, sold to individuals, schools and enterprises on one shared credit meter.
Read the full review →
69.6
averagePersonal $16.99/mo, $13.99 billed annually ($167.88) · Pro $99.99/mo, $74.99 billed annually ($899.88), 25 seats · Education and Enterprise: no published price, Talk to Sales only
2
AI-text detector, plagiarism checker, fact-checker and readability scorer sold to agencies, publishers and educators; scans are billed in credits, one credit per 100 words.
Read the full review →
52.6
lowPro $14.95/mo, $12.95 billed annually (2,000 credits) · Enterprise $179/mo, $136.58 annually (15,000 credits) · NO free tier and no trial tile (originality.ai/pricing, 9 August 2026)
3
Read the full review →
45.3
lown/a
4
Read the full review →
43.8
lown/a

Read the dimensions, not the composite

Before you buy, settle what happens when the tool is wrong about a person. Read the vendor's terms for who carries the liability when a detector's output is used in an academic, employment or disciplinary decision: some contracts put it on the buyer, not the company. And check what any published accuracy figure was measured on: a score from clean, unedited text says little about text somebody lightly rewrote, which is what a detector actually meets.

Where these numbers come from

Capability is sourced from public output-quality arenas (blind pairwise votes) or from a published accuracy study where no arena covers the category, usability from review-crowd aggregates weighted by sample size, value from verified pricing, and commercial terms from clause positions read off the vendor’s own legal documents on a stated date. Where none of those exists for a tool, the row carries a named criterion instead: a different measurement, taken from saved sources under its own rubric, and the card prints that criterion’s name and the question it answers in place of the axis heading, so the row is never read as the axis it could not fill. Full detail: methodology.

This page is ordered by composite, largest first, and the order is computed from the scores at build time rather than written here. READ THE CAPABILITY COLUMN WITH ONE CAVEAT, because it is not a single instrument end to end. Copyleaks, Originality.ai and GPTZero were all scored from one table in Orenstrakh et al. 2023, on the same 114 pre-ChatGPT human submissions and the same forty ChatGPT ones, so those three are comparable to each other row for row. Winston AI is not in that study. Its cell comes from Weber-Wulff et al. 2023, whose machine-paraphrased class was produced with the same paraphrasing tool at its default settings, so the question asked is the same one and the team asking it is different. The pairing is checkable rather than assumed: GPTZero is the tool both teams put through this test, reading 20.0% in one and 33.3% in the other. That is the neighbourhood, not the identity, and this note exists so nobody reads the column as tighter than it is. THE HISTORY OF THIS COLUMN IS WORTH KEEPING, because it is why the caveat above is written down rather than trusted to memory. Until 9 August 2026 Copyleaks' cell converted an English paraphrasing arm while Originality.ai's converted an Arabic-language polishing study: different corpora, different languages, different kinds of degradation, published side by side on a page that ordered by score. The records said the pair was not a like-for-like reading and the layout invited readers to take it as one. Moving both onto Orenstrakh moved Originality.ai's composite as a consequence rather than as its purpose. The Arabic study did not leave the site; it is reported in full on that tool's page, 92% on clean articles against 12% after light polishing, and it remains the sharper finding about edited text. What changed was which of two measurements a single cell is allowed to stand for. The commercial-terms dimension was rebased on 12 August 2026 onto four clause positions read from each vendor's own legal documents. In this category the output-rights position does not apply, because a detector returns a verdict rather than a generated work, and no vendor here publishes a free-tier commercial permission either way. So the cell rests on the indemnity direction and the refund posture, and the spread between the cards on those two clauses is wide.