Are AI Detectors Accurate? From Substack and LinkedIn to Campus False Positives, Understanding the Trust Problem in AI Writing Detection

目錄

This article reflects information available as of August 2026. AI detection models continue to evolve, and the studies cited below should not be treated as direct measurements of the current accuracy of every product available in 2026.

The Verge recently described today’s AI writing environment as a new “age of distrust”: readers are becoming accustomed to wondering whether the article in front of them was written by AI, while human creators are beginning to worry that their own sentences may look too polished, too structured, or too formal and therefore be mistaken for AI-generated writing. This technology is no longer confined to schools. In July, Substack integrated Pangram’s AI text detection into Reader and its iOS app. LinkedIn also introduced a “Seems like AI slop” reporting option while using its own classification systems to reduce the reach of high-volume, automated content that lacks personal perspective. AI detection is gradually moving beyond tools used by teachers grading assignments and into the everyday interfaces people use to read articles, browse social posts, and judge whether creators are trustworthy.

The difficult question is not whether these tools have any ability to detect AI-generated text. It is how much authority a detection result should be allowed to carry. Pangram, GPTZero, and Turnitin all report relatively low false-positive rates, and newer tools have clearly improved since the first wave of products released in 2023. At the same time, independent research, earlier tests, and real campus cases have already shown that false positives are not zero and that performance can vary significantly depending on the text, language background, length, and detector. If a percentage is only a reference for a reader, a 1% error rate may seem acceptable. If the same percentage can affect a grade, suspension, publishing contract, or a creator’s reputation, the stakes are completely different.

AI Detection Has Moved Beyond Schools and Into Substack and LinkedIn

Substack Put Pangram Directly Into the Reading Interface, Allowing Readers to Scan Posts and Notes

Substack’s Scan for AI text feature currently uses Pangram to analyze text and applies to Posts and Notes published after July 21, 2026. Users can launch the detector directly from the menu on a post or Note in the Substack Reader web interface and iOS app, and they can also scan Comments and Replies. The Android version is still in development. Pangram returns an estimate of what percentage of the text is considered human-written or AI-assisted. Creators can also add a How I make this disclosure explaining how they produce their work, and they can disable AI detection for individual posts. When detection is disabled, readers are shown a notice that the content cannot be scanned.

This design is very different from older systems where AI detection remained hidden inside teacher or editor dashboards. Now an ordinary reader can click once and receive an AI probability, and that number can easily shape their impression before they have even evaluated the writing itself. Substack has not said that AI-assisted writing automatically means low quality or deception. Its messaging around the feature has focused on transparency rather than banning AI. The problem is that once an AI detection button becomes part of the reading interface, creators have to deal with another issue beyond writing the article itself: if the tool is wrong, how will readers interpret the result?

LinkedIn’s “Seems like AI slop” Is Not Simply a Test of Whether AI Was Used

LinkedIn’s 2026 feature needs to be distinguished carefully. The company has not said that any use of AI in writing violates its rules. In fact, LinkedIn explicitly says AI can be used to help revise text. What it is trying to reduce is large-scale automated publishing and polished-looking content that lacks genuine perspective, which it describes as AI slop. LinkedIn’s own classification system analyzes whether a post lacks perspective, context, and professional experience. In early testing, the company said the system could identify this type of generic content at a 94% rate. Once content is classified as low-quality AI material, the main effect is reduced recommendation beyond the poster’s existing social network.

The platform has also added a “Seems like AI slop” option that allows users to report posts that appear to be low-quality AI-generated material. For that reason, simplifying the button as “report an AI post” is not quite accurate. LinkedIn is primarily targeting mass-produced content with little practical value, not every piece of writing that has received AI assistance. Even so, once a social platform gives readers a direct button saying “this looks like AI,” the trust problem described by The Verge becomes easier to see: judging whether a piece of writing is good begins to blur together with guessing whether the author used AI.

Are AI Detectors Actually Accurate? The Answer Depends on the Tool, the Dataset, and the Use Case

Turnitin, Pangram, and GPTZero All Report Low False-Positive Rates

If you look only at the numbers currently published by vendors, AI detectors do not appear to be simply guessing. Turnitin has long stated that its AI writing detector has a false-positive rate below 1% for complete human-written texts that meet its analysis requirements. GPTZero currently says its AI/Human classification system also keeps false positives below 1%. Pangram reports an even lower figure, currently claiming an overall false-positive rate of around 1 in 10,000, while the Pangram 4 technical report released in July 2026 reported a 0.0041% false-positive rate on its own test data.

These numbers should not automatically be dismissed as false, but they also cannot be directly compared as if they were produced under the same conditions. Each company uses different test datasets, thresholds, text lengths, and AI models. GPTZero itself emphasizes that no AI detector can guarantee 100% accuracy. Turnitin likewise warns that its AI score should not be used by itself as the basis for disciplinary action against a student. A more reasonable interpretation is that some commercial detectors in 2026 are clearly more mature than early products, but “very low false-positive rates on vendor test sets” and “enough evidence to determine that a specific human cheated” are still two different claims.

Independent Tests Show That Results Can Change When the Dataset Changes

Independent testing matters because it places detectors on data the developers did not select themselves. Bloomberg Businessweek previously tested GPTZero and Copyleaks on 500 college application essays submitted to Texas A&M University in 2022, before ChatGPT became publicly available. Because the essays were produced before the widespread use of generative AI, they could reasonably be treated as human-written. Even so, the two tools produced false-positive rates of roughly 1% to 2%. This does not completely contradict vendors’ claims of low error rates, but it does prove that even a 1% error rate can result in real human writing being mislabeled in high-stakes settings.

Other academic research has produced an even wider range of results. A 2023 study testing 14 tools concluded that AI text detectors at the time were generally not reliable enough to distinguish human writing from ChatGPT-generated text. Newer research from 2025 onward suggests performance has diverged considerably across detectors, with some updated commercial tools showing substantially better accuracy on certain benchmarks. In other words, both “all AI detectors are inaccurate” and “modern detectors are accurate enough to serve as evidence” are oversimplifications. The real questions are which version is being used, what kind of text is being tested, and what happens if the detector is wrong.

A 1% False-Positive Rate Sounds Small Until You Apply It at School Scale

Vanderbilt Used 75,000 Assignments to Estimate Around 750 Potential False Positives

When Vanderbilt University disabled Turnitin’s AI detector in August 2023, it used a very straightforward calculation to explain the concern. Turnitin was reporting a 1% false-positive rate at the time, while Vanderbilt had submitted roughly 75,000 student papers to Turnitin during 2022. If every document had been scanned for AI use, applying a 1% false-positive rate would imply approximately 750 human-written assignments could be incorrectly flagged. Vanderbilt ultimately disabled the feature because of concerns over false positives, insufficient technical transparency at the time, and the university’s conclusion that AI detection was not appropriate as a reliable academic integrity tool.

This does not mean Turnitin currently misclassifies exactly one out of every 100 papers, nor can a 2023 figure be used to predict the performance of its 2026 version. What Vanderbilt’s calculation illustrates is a base-rate problem: when a system scans large numbers of documents and most of them are not actually cheating, even a low false-positive rate can still create a substantial number of cases in which real people are forced to prove their innocence. If the detector simply tells a teacher, “this paper may deserve another look,” other evidence can still be considered. If the score itself becomes proof of misconduct, the cost of errors shifts entirely onto the person being flagged.

Are Non-Native English Writers Really More Likely to Be Misclassified by AI Detectors?

The 61.22% Figure from Stanford’s 2023 Study Is Real, but It Does Not Represent Every Detector in 2026

A 2023 study from a Stanford team remains one of the most frequently cited examples. Researchers tested seven commonly used GPT detectors at the time on 91 TOEFL essays written by non-native English speakers and compared them with English essays written by U.S.-born eighth-grade students. On average, the seven tools misclassified 61.22% of the TOEFL essays as AI-generated. Eighteen essays were incorrectly flagged by all seven detectors, and 89 of the 91 were flagged by at least one detector. By comparison, detection of the native-English writing samples was nearly perfect.

The researchers argued that some detection methods relied heavily on perplexity, or how predictable a piece of text appears to a language model. Non-native English writers often use more common vocabulary and may show less variation in sentence structure, making their writing statistically more predictable. Those characteristics overlap with some patterns seen in early AI-generated text. That is why “simpler English” does not mean AI was used, but it can still produce a higher AI score under detectors that rely heavily on those statistical signals.

However, the time gap needs to be made explicit. The Stanford study tested seven tools from 2023, not the newest versions available in 2026. GPTZero now says that after retraining on ESL data, its false-positive rate on TOEFL essays has fallen to around 1.1%. Pangram has also published its own ESL evaluation, reporting no false positives on the 91 TOEFL essays used in the Liang et al. study. These are the companies’ own current evaluations, showing that vendors have specifically worked on the bias exposed by early research, but independent validation is still necessary.

The Issue for Neurodivergent Writers Should Not Be Simplified as “They Are Definitely More Likely to Be Flagged”

The Verge also raises the possibility that AI detectors may be biased against neurodivergent writers, but the empirical evidence here is not as extensive as the research on non-native English writers. A more careful claim is that if a detector uses regular sentence structure, repeated patterns, formal phrasing, or textual predictability as AI signals, it may also affect people whose natural writing styles happen to be more consistent or structured. Some neurodivergent writers could fall into that category. This is a legitimate fairness concern, but it should not be presented as though there is already comprehensive prevalence data proving that neurodivergent students are routinely misclassified.

AI Detection False Positives Have Reached the Courts, but Two Cases Should Not Be Treated as the Same

Yale Case: GPTZero Was Part of the Initial Concern, but Yale Says the Final Sanction Was Not Based on the Detector Score

Thierry Rignol is a French student in Yale School of Management’s EMBA program. After a final exam in a 2024 course, his professor found that portions of his answers received high AI-generated likelihood scores in GPTZero, and the case was later referred to Yale’s Honor Committee. Rignol ultimately received a failing grade in the course and a one-year suspension, leading him to sue Yale in 2025. His lawsuit also argued that AI detection tools have known biases against non-native English speakers.

However, saying that “a professor used GPTZero to determine he used AI, so Yale directly failed and suspended him” leaves out later parts of the process. Court documents record that Yale’s Honor Committee said its final decision did not rely on the GPTZero scans submitted by the professor. Instead, it also considered Rignol’s failure to provide original documents as requested, explanations the committee found not credible, and significant similarities between his answers and ChatGPT responses. Rignol denies cheating and challenges the fairness of the process. One of the most significant public rulings so far is that a federal court denied his request for a preliminary injunction that would have immediately reinstated him in May 2025. The case therefore cannot be described as one in which he already won because GPTZero had produced a false positive.

Adelphi Case: The Student Won, and the Court Ordered the Academic Misconduct Finding Removed

The case involving Orion Newby at Adelphi University had a different outcome. Newby submitted an essay comparing Christianity and Islam in a World Civilizations 1 course and was accused of using AI. Inside Higher Ed, citing court records, reported that Turnitin’s AI detector labeled the paper as fully AI-written, while two other detection tools used by Newby classified it as human-written. Adelphi nevertheless upheld the academic misconduct finding, and Newby’s family later pursued legal action. In early 2026, a New York court ruled in Newby’s favor, finding that the university’s plagiarism determination lacked a valid basis and ordering Adelphi to remove the accusation from his record.

This case is important, but it still should not be simplified into “a court proved Turnitin is inaccurate.” The court was not evaluating only the detector. It also considered whether the university followed its own academic disciplinary procedures, whether Newby received an adequate opportunity to appeal, and whether the overall decision had a reasonable foundation. A more accurate conclusion is that once an AI detector result enters a high-stakes disciplinary process, it cannot be used in isolation from other evidence and basic procedural safeguards.

If an AI Detector Misclassifies Your Writing, Preserving Evidence of the Writing Process Is More Useful Than Trying to Make the Text “Look Less Like AI”

Version History Is More Persuasive Than Running the Text Through Another AI Detector

If an article, assignment, or manuscript might later face questions about AI use, the most practical approach is not to run every paragraph through multiple detectors. It is to preserve evidence showing how the text was created. Google Docs Version history, Microsoft Word version records, Notion Page history, Git commits, research notes, source materials, and drafts can all show how a piece of writing developed from a blank page into a finished work. The key difference between this kind of process evidence and a detector is that it is not another model making a probabilistic judgment about the final text. It records an actual production timeline.

If you normally write an entire piece somewhere else and paste it into the final document in one step, the version history will naturally show one large insertion. In situations involving schoolwork, freelance writing, publishing, or other contexts where authorship may need to be demonstrated, it can be useful to preserve original drafts, revision history, and research materials. There is no need to deliberately disrupt your normal workflow merely to create “human-looking evidence.”

If You Are Accused, First Find Out What Role the Detection Result Is Actually Playing

The second step is not immediately finding another detector that says “this was written by a human,” because disagreement between detectors is part of the problem in the first place. A more practical approach is to determine which tool was used, which version, what score it produced, whether the school or platform’s policy permits an AI score to be used by itself, and whether other evidence exists. Turnitin itself warns that AI writing scores can be inaccurate and should not be used alone to take adverse action against students. GPTZero likewise publicly states that no AI detector can achieve 100% accuracy.

In an educational setting, it is also useful to review the course’s original AI policy, academic integrity procedures, and appeal rules, then provide version history, drafts, citations, sources, and previous writing samples. What carries more weight is not simply saying “I did not use AI,” but presenting a set of records that can actually be checked.

Do Not Deliberately Make Your Writing Worse Just to Lower an AI Score

The Verge’s description of “humans trying not to sound like AI” captures one of the strangest side effects of the current detection culture. Some AI detectors treat highly regular sentence structure, formal tone, repeated patterns, or low perplexity as possible signals, so creators begin worrying that their writing is too polished. Some even introduce unnecessary mistakes or alter otherwise natural sentences in an attempt to lower their AI score.

There is no reliable evidence that this works consistently, and it can easily turn the detector into a new writing standard. AI models and detectors will continue to change, so a tactic that lowers the score today may not work in the next version. The awkward edits added to the article, however, remain permanently. For creators, preserving their own writing style and production records is more reasonable than rewriting for the benefit of a classifier.

The Real Problem with AI Detection Is Not Whether It Can Detect AI, but How Much a Score Is Allowed to Decide

AI detectors in 2026 should no longer be judged solely by the performance of the first wave of tools released in 2023. Published results from newer systems such as Pangram 4 and GPTZero show that the technology has continued to improve, with some benchmarks now reporting extremely low false-positive rates. At the same time, Stanford’s research on non-native English writers, Bloomberg’s testing of pre-ChatGPT human essays, and student cases that reached the courts are enough to show that “low false-positive rate” does not mean “never wrong.”

Substack and LinkedIn bring another layer to the issue by moving these judgments into everyday reading interfaces. In the past, a detector error might have been visible only to a teacher and a student. Now any reader may be able to press a button and immediately attach a “looks like AI” impression to someone’s work. Detection tools can be useful as signals, but once they begin affecting reputation, grades, or employment opportunities, they need to be combined with original documents, version history, human judgment, and a clear appeal process. That is more practical than searching for a detector that never makes mistakes, because at least for now, no vendor is willing to guarantee 100% accuracy.

FAQ

Are AI Detectors Accurate Now?

Some newer commercial tools report very low false-positive rates on their own benchmarks. Turnitin and GPTZero both say their false-positive rates are below 1%, while Pangram reports an even lower figure. However, results vary by detector, dataset, text length, and use case, so no single AI score should be treated as definitive proof that an author used AI.

Are Non-Native English Writers Really More Likely to Be Misclassified by AI Detectors?

In a 2023 Stanford study testing seven detectors available at the time, 61.22% of 91 TOEFL essays written by non-native English speakers were misclassified as AI-generated on average, while native-English samples were identified almost perfectly. However, those numbers describe 2023 tools and should not be directly applied to 2026 products. Companies including GPTZero and Pangram say they have since improved their systems specifically to reduce ESL bias.

How Can I Prove I Wrote Something Myself If an AI Detector Flags It?

Version history, drafts, research notes, citations, revision records, and previous writing samples are usually more useful than the score from another detector. If the issue involves school discipline, you should also confirm which tool was used, whether other evidence exists, and whether the institution’s academic integrity procedures allow decisions to be made solely from an AI detection score.

Have Schools Actually Stopped Using AI Detectors Because of False-Positive Concerns?

Yes. Vanderbilt University disabled Turnitin’s AI detector in 2023 and noted that with roughly 75,000 submitted documents in a year, even a 1% false-positive rate could theoretically result in around 750 incorrect flags. The Verge has also reported that institutions including Yale, Johns Hopkins, and Georgetown have disabled or restricted the use of such tools.

Is AI Detection Still Mainly a School Issue?

No. Substack now allows readers to use Pangram to scan Posts, Notes, Comments, and Replies, while LinkedIn has added a “Seems like AI slop” reporting option. AI text judgments are moving into mainstream social and publishing platforms, meaning creators may face not only internal review from schools or employers, but also immediate judgments from ordinary readers about how their work was produced.

SUPPORT FENGNIII

喜歡這篇文章嗎?

如果這篇內容對你有幫助,可以透過小額贊助支持本站持續整理更多日文、韓文、旅行與數位工具內容。

小額支持本站

付款將由藍新金流安全處理