Two things happened on 2 August 2026 that changed what AI content disclosure means.
AI Content Detection Statistics 2026: 90 Sourced Numbers

The EU AI Act’s transparency obligations became applicable, requiring generative AI providers to mark their outputs in a machine-readable format. The same day, California’s AI Transparency Act became operative after a one-year delay, requiring covered providers to ship a free public AI detection tool. Disclosure stopped being a best practice and became law in the EU and California on the same day.
Here’s the awkward part. The detectors those laws lean on still don’t work well enough to settle an argument. The largest neutral benchmark published with its own dataset found commercial detectors “easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models.” The most-cited bias study found that seven detectors flagged 61.3% of essays by non-native English writers as AI-generated. And the company that built ChatGPT killed its own detector in 2023 because it caught 26% of AI text while falsely accusing 9% of human writers.
California’s half of that requires covered providers to ship a working public detection tool, which is a capability the independent testing says nobody has. The rest of both laws sidesteps the problem by making provenance metadata a property of the file instead. That split is the story of this page.
Below are 90 AI content detection statistics for 2026, grouped by the question each one actually answers. Every figure was pulled from the organization that produced it in September 2026: peer-reviewed papers, arXiv preprints, regulatory text, and company transparency centers. Where a source is a vendor measuring itself, I say so on the line.
Numbers that didn’t survive that check are gone. The most widely repeated stat in this category, that AI writing crossed 50% of new articles in November 2024 and hit 52% by May 2025, has been retracted in effect by the firm that published it. The Sources and Methodology section at the bottom lists every exclusion and why.
Key Takeaways
- AI’s share of new articles stalled, it didn’t take over. Graphite’s revised measurement puts primarily AI-generated articles at 49.94% in Q1 2026, roughly where they’ve sat for five straight quarters. The famous “52% and climbing” figure came from a study Graphite has since superseded.
- Detector accuracy claims and detector accuracy are different things. Vendors cluster at 99%+. The RAID benchmark, over 6 million generations across 11 models, found those same detectors easily fooled by paraphrasing and unseen models.
- The false-positive problem lands on non-native English writers. Seven detectors averaged a 61.3% false-positive rate on TOEFL essays while classifying US student essays correctly.
- Using AI doesn’t cost you rankings, but pure AI rarely wins them. Ahrefs measured a 0.011 correlation between AI share and ranking position. Graphite found only 7% of pages ranking first are AI-generated.
- AI Overviews cite AI-assisted pages at a higher rate than Google ranks them. 91.4% of AI Overview citations contain at least some AI-generated content, against 86.5% of top-ranking pages on the same detector and threshold.
- Provenance is scaling faster than detection. Google has watermarked over 100 billion images and videos with SynthID. No text detector has anything like that reach.
- The crawl is one-directional. 80% of AI crawling is for training, 18% for search, 2% for user actions. Anthropic crawled 38,000 pages for every visitor it referred back.
Top AI Content Detection Statistics for 2026
If you carry four numbers out of this page into your next meeting, carry these.

The Numbers Worth Quoting
1. 49.94% of new online articles published in Q1 2026 were primarily AI-generated, against 50.06% human-written (Graphite, May 2026, n=55,400 Common Crawl URLs).
2. Seven widely used detectors produced an average false-positive rate of 61.3% on TOEFL essays written by non-native English speakers (Liang et al., Patterns, 2023, n=91).
3. Only 14% of articles ranking in Google Search are AI-generated; 86% are human-written (Graphite, October 2025).
4. 91.4% of pages cited in AI Overviews contain at least some AI-generated content (Ahrefs, July 2025, 1 million SERPs).
The Number That Ended OpenAI’s Own Detector
5. OpenAI’s AI text classifier correctly identified 26% of AI-written text while falsely flagging 9% of human-written text, which is why the company retired it on 20 July 2023 (OpenAI).
OpenAI shipped the model that started this argument and could not build a reliable detector for its own output. Every accuracy claim further down this page should be read against that.
How Accurate AI Content Detectors Actually Are
Every detector vendor publishes a number north of 99%. Every independent study that puts several detectors on the same dataset produces something far worse. Both sets of numbers are real, and the distance between them is entirely about who picked the test.
What The Vendors Publish
6. GPTZero claims 99% accuracy, a false-negative rate under 2%, and a false-positive rate under 1% (GPTZero, vendor-published).
7. GPTZero also reports 96.5% accuracy on mixed documents that combine human and AI writing (GPTZero, vendor-published).
8. In its own three-way comparison across 3,000 samples, GPTZero reported 99.3% overall accuracy and a 0.24% false-positive rate (GPTZero, vendor-published).
9. Turnitin states it has ensured “a high accuracy rate accompanied by a less than 1% false positive rate” (Turnitin, vendor-published).
10. Originality.ai reports 85% accuracy on the RAID base dataset against 80% for the closest competitor, and 96.7% on paraphrased content against a 59% average across all other detectors (Originality.ai, vendor-published).
11. Pangram Labs reports 99.85% accuracy at a 0.19% false-positive rate on a 1,976-document benchmark spanning 8 large language models and 10 text domains (Emi and Spero, arXiv 2402.14873, v3 July 2024, vendor-authored).
One clarification on stat 10, because it gets repeated wrongly everywhere. Originality ran that test itself against the publicly released RAID dataset. It is not a finding of the RAID paper, and Originality is not one of the detectors that paper evaluated. Anyone citing it as “the RAID study found Originality most accurate” is citing a vendor blog. The three detectors whose numbers appear most in this section have their own reviews here: Originality.ai, GPTZero and Copyleaks.
What Independent Studies Measure

12. Across 14 detection tools, the conclusion was that “detection tools for AI-generated text do fail, they are neither accurate nor reliable,” with all tools scoring below 80% accuracy and only five above 70% (Weber-Wulff et al., International Journal for Educational Integrity, 2023).
13. Those 14 tools scored 96% accuracy on human-written text but only 74% on unmodified AI-generated text (Weber-Wulff et al., 2023).
14. Accuracy fell to 42% when the AI text had been edited by a human, and to 26% when it had been machine-paraphrased (Weber-Wulff et al., 2023).
15. Roughly 20% of AI-generated texts would be misattributed to humans unmodified, rising to roughly 50% once a human edits them and higher still once a machine paraphrases them (Weber-Wulff et al., 2023).
16. Across those 14 tools, false positives ranged from 0% for Turnitin to 50% for GPT Zero, and false negatives from 8% for GPT Zero to 100% for Content at Scale (Weber-Wulff et al., 2023).
17. A separate peer-reviewed test of seven detectors across 797 valid tests measured a mean accuracy of 39.5% on unmanipulated AI-generated content (Perkins et al., International Journal of Educational Technology in Higher Education, September 2024).
18. The same study found detectors correctly classified only 67% of human-written control samples (Perkins et al., 2024).
19. Copyleaks led that field by detecting 64.8% of AI-generated texts, followed by Turnitin at 61%, with GPTZero lowest at roughly 26% (Perkins et al., 2024).
20. Simple evasion techniques cut detector accuracy by a further 17.4 points, taking the average down to 22.14% (Perkins et al., 2024).
21. Turnitin lost the most under those techniques, dropping 42.1 points and falling from second place to fifth of seven (Perkins et al., 2024).
22. In a peer-reviewed medical-text test of 50 pieces, 20 written by ChatGPT and 30 taken from published articles, GPTZero scored 0.80 accuracy, 0.65 sensitivity and 0.90 specificity (Habibzadeh, Journal of Korean Medical Science, 2023).
23. That test’s confidence interval on sensitivity runs from 0.41 to 0.85, so it could not rule out a detector missing more than half of all AI text (Habibzadeh, JKMS, 2023).
Line up stat 6 against stats 17 and 19. GPTZero publishes 99% accuracy. An independent lab measured the category average at 39.5% and put GPTZero at the bottom of it on roughly 26%. Neither party is lying. They ran different tests and only one of them published the dataset.
Where Detectors Break
24. RAID, the largest neutral benchmark of machine-generated text detectors, spans over 6 million generations across 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, evaluating 8 open-source plus 4 closed-source detectors (Dugan et al., ACL 2024). The paper body puts the exact count at 6,287,820.
25. The RAID authors note that “many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more)” and then find that “current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models” (Dugan et al., ACL 2024).
26. In Pangram’s cross-detector benchmark the false-positive rates were Pangram 0.19%, GPTZero 2.01%, Originality.ai 9.24% and DetectGPT 14.4% (arXiv 2402.14873, vendor-authored).
The consistent finding across stats 14, 20 and 25 is that detector accuracy is a property of the test, not of the detector. Paraphrase the text, edit it by hand, or generate it with a model the detector hasn’t seen, and the number collapses. A detector score is evidence about how the text reads, and nothing more.
False Positives And Who Gets Flagged
The research is consistent about who absorbs the cost of a wrong answer, and it is not the vendor.
The Non-Native English Problem
27. Seven widely used GPT detectors were evaluated on 91 TOEFL essays from non-native English writers and 88 US eighth-grade essays from native writers (Liang et al., Patterns, 2023).
28. The detectors “incorrectly labeled more than half of the TOEFL essays as AI-generated,” at an average false-positive rate of 61.3% (Liang et al., Patterns, 2023).
29. All seven detectors unanimously flagged 19.8% of the human-written TOEFL essays as AI-authored (Liang et al., Patterns, 2023).
30. At least one of the seven detectors flagged 97.8% of the TOEFL essays (Liang et al., Patterns, 2023).
31. The same detectors classified the US student essays accurately, which is what makes this a bias finding rather than a general accuracy finding (Liang et al., Patterns, 2023).
32. On the same 91 TOEFL essays held out of training, Pangram reported a 0% false-positive rate against 7.7% for GPTZero (arXiv 2402.14873, vendor-authored).
The Liang study is from 2023 and every detector in it has shipped new models since. The mechanism it identified has not changed: these systems key on low text perplexity, low perplexity tracks a narrower vocabulary, and a narrower vocabulary tracks writing English as a second language. Stat 32 is the honest counterweight, and it comes from a vendor with an interest in the result.
What Vendors Do About False Positives
33. Turnitin shows no score and no highlights for AI detection results between 1% and 19%, displaying an asterisk instead, specifically to avoid surfacing false positives in that band (Turnitin).
34. Turnitin requires a minimum of 300 words before it will return an AI writing score at all (Turnitin).
35. OpenAI’s retired classifier was calibrated to a 9% false-positive rate on human text and still only caught 26% of AI text, a trade-off the company judged unacceptable enough to withdraw the product (OpenAI, 20 July 2023).
Suppressing scores under 20% is the most honest thing any vendor in this category does, and it’s an admission. If the bottom fifth of your scale is too noisy to show a customer, the scale isn’t a measurement.
How Much Of The Web Is AI-Generated
This is the question everyone asks and the one with the least stable answer. Three respected measurements give three different numbers because each one draws the line in a different place.
The Prevalence Curve

36. 0.97% of new online articles were primarily AI-generated in Q1 2020 (Graphite, May 2026).
37. That share was still 4.6% in Q4 2022, the quarter ChatGPT launched (Graphite, May 2026).
38. It reached 35.92% by Q4 2023 and 47.04% by Q4 2024 (Graphite, May 2026).
39. It peaked at 50.9% in Q4 2025 and sat at 49.94% in Q1 2026 (Graphite, May 2026).
40. Graphite’s own summary is that “the proportion of primarily AI-generated articles has remained relatively stable, near 50%, over the last five quarters” (Graphite, May 2026).
41. That study classified 55,400 Common Crawl URLs published between January 2020 and March 2026 using three detectors, Pangram, Copyleaks and GPTZero (Graphite, May 2026).
42. The three detectors’ false-positive rates in that study were 1.844% for Pangram, 1.836% for Copyleaks and 1.355% for GPTZero (Graphite, May 2026).
The flattening in stats 38 to 40 is the part nobody quotes. AI writing did not eat the web. It got to roughly half of new articles by the end of 2024 and stopped there.
Why The Three Big Numbers Disagree

43. 74.2% of 900,000 newly indexed web pages contained AI-generated content; 25.8% were classified as pure human and 2.5% as pure AI (Ahrefs, published May 2025 on an April 2025 crawl).
44. 86.5% of top-ranking pages contain some AI-generated content, while 4.6% are entirely AI and 13.5% are entirely human (Ahrefs, July 2025, n=600,000 URLs across 100,000 keywords).
45. 91.4% of pages cited in AI Overviews contain at least some AI content: 3.6% pure AI and 87.8% mixed, leaving 8.6% pure human (Ahrefs, July 2025, 1 million SERPs and 1.9 million cited URLs).
46. Graphite counts an article as AI-generated only when the majority of it is flagged. Ahrefs counts a page as containing AI content when a single sentence is flagged (Graphite, May 2026; Ahrefs, May 2025).
So “half the web is AI” and “three quarters of the web is AI” are both true statements about different questions. The first asks how many articles are mostly machine-written. The second asks how many pages have been touched by a machine at all. Pick the one that matches what you’re actually arguing about, and say which threshold you used.
AI Content Farms
47. NewsGuard has identified 3,749 AI content farm news and information sites (NewsGuard, last updated 23 June 2026).
48. Those sites span 16 languages: Arabic, Chinese, Czech, Dutch, English, French, German, Indonesian, Italian, Korean, Portuguese, Russian, Spanish, Tagalog, Thai and Turkish (NewsGuard, June 2026).
This is a count of sites identified by analysts, not a percentage of anything, and it’s a floor rather than a total. It’s the number to use when someone asks about AI slop specifically, rather than AI assistance generally.
What Ranks And What Gets Cited
Prevalence in publication and prevalence in retrieval are separate numbers, and the gap between them is the most commercially useful thing on this page.
AI Content In Google Search
49. 86% of articles ranking in Google Search are human-written and 14% are AI-generated (Graphite, October 2025).
50. In Graphite’s 2024 measurement the split was 12% AI and 88% human, so the AI share of ranking articles moved 2 points in a year while the AI share of published articles was near 50% (Graphite, October 2025).
51. Only 7% of the articles ranking first are AI-generated (Graphite, October 2025).
52. The correlation between how much AI-generated content a page contains and how highly it ranks is 0.011, which Ahrefs describes as no clear relationship (Ahrefs, July 2025).
Stats 51 and 52 look contradictory and aren’t. AI assistance carries no ranking penalty. Fully machine-written pages just rarely turn out to be the best answer, which is a quality outcome rather than a detection outcome.
AI Content In AI Overviews And Chatbots
53. 82% of articles cited by ChatGPT with web search enabled are human-written, and 18% are AI-generated (Graphite, October 2025).
54. Perplexity shows the same 82% human and 18% AI split in its citations (Graphite, October 2025).
55. Against those figures, 91.4% of AI Overview citations contain at least some AI content, because Ahrefs and Graphite are measuring different things: any flagged sentence versus a majority-AI article (Ahrefs, July 2025).
If you want your pages cited by an assistant, stats 53 and 54 are the ones to plan around. Mostly-machine articles are cited at roughly a third of the rate they’re published at.
What Google Actually Says
56. Google’s position, unchanged and still live since February 2023, is that “appropriate use of AI or automation is not against our guidelines” (Google Search Central).
57. Google adds that “using AI doesn’t give content any special gains. It’s just content. If it is useful, helpful, original, and satisfies aspects of E-E-A-T, it might do well in search” (Google Search Central).
58. Google’s March 2024 core update set out to “reduce low-quality, unoriginal content in search results by 40%” (Google).
59. On 26 April 2024 Google reported the rollout had completed and users would “see 45% less low-quality, unoriginal content in search results versus the 40% improvement we expected” (Google).
60. Google’s current spam policy defines scaled content abuse as generating many pages to manipulate rankings “no matter how it’s created,” and its first listed example is using generative AI to produce many pages without adding value (Google Search Central, last updated 28 August 2026).
61. Google’s evergreen guidance page on AI-generated content carries a last-updated date of 10 December 2025 (Google Search Central).
Note the wording change in stat 60. Google used to frame this as “whether automation or humans are involved.” The current text says “no matter how it’s created.” Same policy, and it removes the last excuse to read this as an AI rule. It’s a volume-and-value rule.
AI Crawler Traffic And What It Costs You
Crawler traffic is not AI-generated content, and it measures something different from every other number on this page. It belongs here because it is the part of the AI content economy that shows up in your own server logs and referral reports.
Who Crawls The Most

62. AI and search crawler traffic grew 18% from May 2024 to May 2025 (Cloudflare, July 2025).
63. By May 2025 the AI crawler share split GPTBot 30%, ClaudeBot 21%, Meta-ExternalAgent 19%, Amazonbot 11% and Bytespider 7.2% (Cloudflare, July 2025).
64. A year earlier the same ranking read Bytespider 42%, ClaudeBot 27%, Amazonbot 21%, GPTBot 5% and Applebot 4.1% (Cloudflare, July 2025).
65. Measured against all crawler traffic rather than AI crawlers alone, GPTBot went from 2.2% to 7.7%, a 305% rise in requests (Cloudflare, July 2025).
66. Googlebot’s share of all crawler traffic rose from 30% to 50% over the same year, growing 96% (Cloudflare, July 2025).
67. Around 30% of global web traffic now comes from bots (Cloudflare, July 2025).
Stats 63 and 64 are the same league table one year apart, and almost every position changed. Anything you decide about crawler policy on today’s numbers has a shelf life of about twelve months.
What The Crawling Is For

68. “Over the past 12 months, 80% of AI crawling was for training, compared with 18% for search and just 2% for user actions” (Cloudflare, August 2025).
Crawl-To-Referral Ratios
69. Anthropic crawled 38,000 pages for every visit it referred back in July 2025, down from 286,000 to 1 in January, an 87% improvement (Cloudflare, August 2025).
70. OpenAI’s ratio moved from 1,217 crawls per referred human in January 2025 to 1,091 in July (Cloudflare, August 2025).
71. Microsoft’s ratio was 38.5 in January and 40.7 in July, an order of magnitude better than either (Cloudflare, August 2025).
72. Perplexity moved the other way, from 54 crawls per human in January to 195 in July, a 256.7% increase (Cloudflare, August 2025).
Stat 71 is the one I’d put in front of anyone deciding whether to block AI crawlers. Microsoft sends a visitor for every 41 pages it reads. OpenAI reads about 1,100 and Anthropic about 38,000. Those are different businesses wearing the same user-agent header, and a blanket robots.txt rule treats them identically.
Provenance, Labeling And The Law
Detection tries to work out where a file came from after the fact. Provenance stamps the answer in at creation. One of these is scaling and one isn’t.
Watermarking And Content Credentials
73. Google reports that SynthID has watermarked “over one hundred billion images and videos, along with sixty thousand years of audio assets” since launch (Google, I/O 2026 keynote, May 2026).
74. In the SynthID-Text deployment study, roughly 20 million watermarked and unwatermarked Gemini responses were compared; thumbs-up rates differed by 0.01% and thumbs-down rates by 0.02%, both statistically insignificant (Dathathri et al., Nature, 2024).
75. More than 6,000 C2PA members and affiliates have live applications of Content Credentials (C2PA, Content Credentials 2.3 announcement, February 2026).
76. The C2PA steering committee is Adobe, Amazon, BBC, Google, Meta, Microsoft, OpenAI, Publicis Groupe, Sony, TikTok and Truepic (C2PA).
77. Cloudflare, the first major CDN to implement Content Credentials, says “twenty percent of the internet runs through our network” (Content Authenticity Initiative).
Stat 73 against stat 75 is the whole argument. Google watermarks at a scale no text detector approaches, and it only verifies Google’s own output. Content Credentials cover more producers and survive fewer round trips, because re-encoding a file strips the metadata and most social platforms strip it by default.
Platform Disclosure Rules
78. YouTube exempts production assistance from its disclosure requirement, naming “a video outline, script, thumbnail, title, or infographic,” plus caption creation, sharpening, upscaling, repair and audio repair (YouTube Help).
79. Creators who consistently skip disclosure “may be subject to manual application of a label, or penalties from YouTube, including removal of content or suspension from the YouTube Partner Program” (YouTube Help).
80. In its one published measurement window, 1 to 29 October 2024, Meta recorded over 360 million labeled pieces of content with user label views on Facebook and over 330 million on Instagram, against over 380 billion label views on Facebook and over 1 trillion on Instagram (Meta Transparency Center).
81. Meta renamed its “Made with AI” label to “AI info” on 1 July 2024, after finding that minor edits made with retouching tools carried industry-standard AI indicators and got labeled anyway (Meta).
82. TikTok launched creator-applied AI labels in September 2023 and became “the first major social media platform to support the open C2PA standard” in May 2024 (Content Authenticity Initiative).
Stat 80 is nearly two years old and it’s still the only window Meta has published. A transparency report that hasn’t been refreshed since October 2024 tells you how seriously the labeling programme is being tracked.
What Changed On 2 August 2026
83. The EU AI Act’s Article 50 transparency obligations became applicable on 2 August 2026, requiring generative AI providers to mark outputs in a machine-readable format detectable as artificially generated (European Commission).
84. The final Code of Practice on marking and labelling AI-generated content was published on 10 June 2026 and confirmed by the Commission and the AI Board as an adequate voluntary route to demonstrating compliance (European Commission).
85. About 190 companies and organisations had signed that code by the end of July 2026 (European Commission).
86. California’s AI Transparency Act became operative on 2 August 2026, delayed by AB 853 from its original 1 January 2026 date, and applies to generative AI services with more than one million monthly visitors or users in California (SB 942).
87. Covered providers must offer a free, publicly accessible AI detection tool, embed latent machine-readable disclosure in outputs, and make manifest disclosure available to users, with civil penalties up to $5,000 per violation (SB 942).
88. AB 853 phases in obligations for large online platforms from 1 January 2027 and for capture-device manufacturers from 1 January 2028 (AB 853).
89. The US Copyright Office holds that images “generated by the Midjourney technology are not the product of human authorship,” a position it applied by cancelling and reissuing the Zarya of the Dawn registration (US Copyright Office, 21 February 2023).
90. The DC Circuit affirmed in Thaler v. Perlmutter (No. 23-5233, D.C. Cir., 18 March 2025) that the Copyright Act “requires all eligible work to be authored in the first instance by a human being.” Thaler named his AI system as sole author with no human author claimed, so the holding is narrower than a general rule about AI-assisted work.
Here’s what I’d take from stats 83 to 87 as a marketer rather than a lawyer. Both rules land on the model providers, not on you. Neither requires you to label an AI-assisted blog post. What they do is make provenance metadata a default property of generated files, which over time makes the detector argument moot.
Frequently Asked Questions
How Accurate Are AI Content Detectors In 2026?
Far less accurate than the marketing says. Two peer-reviewed studies put the category between 39.5% and 74% accuracy on unmodified AI text, and both found accuracy falls further once the text is edited or paraphrased. Weber-Wulff et al. tested 14 tools and concluded they “are neither accurate nor reliable,” with every tool scoring under 80%. Perkins et al. tested seven detectors and measured a 39.5% mean, dropping to 22.14% after simple evasion techniques. Vendor figures of 99% or better are measured on the vendor’s own held-out benchmark and don’t transfer.
Does Google Penalize AI-Generated Content?
No. Google’s position since February 2023 is that “appropriate use of AI or automation is not against our guidelines,” and it has never retracted that. What Google does police is scaled content abuse, which its current spam policy defines as generating many pages to manipulate rankings “no matter how it’s created.” The March 2024 update built on that and reduced low-quality, unoriginal content by 45% against a 40% goal. Nothing in the policy turns on whether a machine wrote the words. The measured evidence agrees: Ahrefs found a 0.011 correlation between AI content share and ranking position.
What Percentage Of Web Content Is AI-Generated?
It depends entirely on where you draw the line, and the two most-cited numbers use opposite thresholds. Graphite, counting an article as AI only when most of it is machine-written, puts new articles at 49.94% for Q1 2026 and describes the share as stable near 50% for five quarters. Ahrefs, counting a page as AI-touched if a single sentence is flagged, found 74.2% of 900,000 new pages. Both are correct. Say which threshold you’re using or the number means nothing.
Are AI Detectors Biased Against Non-Native English Writers?
Yes. The Stanford study that established this evaluated seven detectors on 91 TOEFL essays and 88 essays by US eighth-graders. The detectors classified the US essays accurately but produced a 61.3% average false-positive rate on the TOEFL essays, with all seven unanimously misflagging 19.8% of them and at least one flagging 97.8%. That work is from 2023 and detectors have shipped new models since, but the mechanism has not changed: these systems key on low text perplexity, which tracks vocabulary range, which tracks whether English is your first language.
Why Did OpenAI Shut Down Its AI Text Classifier?
Accuracy. OpenAI retired the classifier on 20 July 2023 “due to its low rate of accuracy.” At launch it correctly identified 26% of AI-written text while incorrectly flagging 9% of human-written text as AI. OpenAI said at the time it was researching more effective provenance techniques for text instead, which is the direction the whole field has since taken.
What Does The EU AI Act Require For AI Content?
Article 50’s transparency obligations became applicable on 2 August 2026. Providers of generative AI must mark outputs in a machine-readable format detectable as artificially generated, and deployers must label deepfakes and AI-generated text on matters of public interest. The final Code of Practice on marking and labelling was published on 10 June 2026, and about 190 organisations had signed it by the end of July. The obligation lands on the model providers, not on you as a publisher, and there’s an exemption for content under human editorial responsibility.
Do I Have To Label AI-Assisted Blog Posts?
Not under the EU AI Act or California’s AI Transparency Act, both of which put the marking obligation on the providers of the generative systems. Platform rules are narrower than people assume too: YouTube explicitly exempts production assistance, naming outlines, scripts, thumbnails, titles, infographics and captions, and only requires disclosure for realistic synthetic content. Whether you disclose beyond that is an editorial and trust decision rather than a legal one.
Can AI Detectors Tell The Difference Between AI-Assisted And AI-Generated?
Poorly. This is the single weakest point in the technology and the one that matters most for content teams, because assisted writing is the normal case. Weber-Wulff et al. measured accuracy dropping from 74% on raw AI output to 42% once a human edited it. GPTZero claims 96.5% on mixed documents, but that figure is vendor-measured on a vendor-chosen dataset. Any workflow where a human drafts, edits or fact-checks alongside a model lands in exactly the band where independent testing shows detectors are least reliable.
Sources And Methodology
This page uses primary sources only: peer-reviewed papers, preprints with published methodology, regulatory text, and company-owned transparency centers and policy pages. Secondary press coverage was used to locate sources and is never the source of record. Every figure was checked against its origin document in September 2026, and the retrieval date is stated wherever a source updates continuously.
Vendor-published figures are labeled inline as vendor-published or vendor-authored. They are included where the methodology is public enough to be argued with, and excluded where the only artifact is a marketing page. Copyleaks, Winston AI, Sapling and Writer publish accuracy claims without any accompanying methodology document, dataset or technical report, so their numbers do not appear here.
Several figures carried by earlier versions of this research were removed rather than updated:
The most widely circulated prevalence statistic in this category, that AI-generated articles crossed 50% in November 2024 and reached 52% by May 2025, comes from a Graphite study that Graphite has since superseded. Its own replacement, published in May 2026 on a larger sample and three detectors rather than one, revises the entire time series downward and shows the crossover happening later and flattening rather than climbing. The superseded figures are not used on this page.
A claim that the Jisc National Centre for AI evaluated 16 detectors in June 2025 and found accuracy between 33% and 65% with false-positive rates of 10% to 14% was removed. The June 2025 Jisc post is a literature review citing several third-party studies, not a Jisc-run evaluation, and no “16 detectors” framing or 10% to 14% false-positive band appears in it. The underlying peer-reviewed studies are cited directly instead.
Two Cloudflare figures, a 147% rise in GPTBot requests and an 843% rise for Meta-ExternalAgent between July 2024 and July 2025, were removed because neither appears on the Cloudflare pages they were attributed to. Cloudflare’s own comparable figure, a 305% rise in GPTBot requests measured against all crawler traffic, is used instead.
A Turnitin missed-AI rate of roughly 15% and a claim that its minimum word count rose from 150 to 300 words were removed. Neither appears on any Turnitin primary source; the 15% figure traces only to a third party paraphrasing a company spokesperson.
Figures older than 24 months are date-stamped explicitly. The Liang et al. bias study and the OpenAI classifier numbers both predate 2024 and are retained, marked as historical, because they remain the canonical evidence for findings no newer work has overturned.
The D.C. Circuit opinion in Thaler v. Perlmutter is cited by court, docket number and date rather than by link, because the court’s own opinion endpoint for that docket no longer resolves.
Related statistics pages on this site: content marketing statistics, SEO statistics and website statistics.

Chintan Zalani
Hey, I’m Chintan, a creator and the founder of Elite Content Marketer. I make a living writing from cafes, traveling to mountains, and hopping across cities. Join me on this site to learn how you can make a living as a sustainable creator.
View Profile →Related Articles

AI Hallucination Rate 2026: 60+ Sourced Statistics
The real AI hallucination rate in 2026, across 64 sourced statistics: OpenAI’s o3 scores 5.7% on FActScore with browsing on and 24.2% with it off.

94 Generative Engine Optimization Statistics (2026)
94 generative engine optimization statistics for 2026, each checked against its primary source: 88% of ChatGPT citations come from ordinary web search.

71 YouTube Statistics for Creators and Marketers 2026
71 YouTube statistics for 2026, every number traced to its source: YouTube passed $60 billion in revenue and took 13.8% of all U.S. TV viewing time.