Does AI-Detection Software Work? We Tested 3 Popular Tools.

By The Epoch Times | Created at 2026-09-30 09:16:57 | Updated at 2026-09-30 10:17:38 1 hour ago

Writing generated by artificial intelligence is becoming a familiar headache for publishers, educators, and other professionals amid widespread calls to protect and prioritize human-authored content.

AI writing detectors have stepped up to the challenge; some tool makers say they can identify human versus AI prose with up to 99 percent accuracy.

But attorneys say caution is needed. Educators, publishers, and employers using an AI detection score as the basis of their decision-making over questionable writing could face serious legal consequences.

AI experts—even those who’ve built detection tools—say these services are meant to offer useful clues, but aren’t infallible. In fact, they seem to share a common blind spot.

Putting It to the Test

Researchers have been evaluating AI detection tools for some time, and studies and experts have emphasized one key point: Humans still need to be part of the assessment process.

In a 2026 study published in the International Journal for Educational Integrity, researchers wrote, “While detection tools can provide useful initial flags, they should not be used as sole evidence in high-stakes decision-making but should be implemented in a broader evaluation strategy.”

Studies and experts have also raised concerns about the accuracy of these tools and their potential to flag human writing as AI.

“A detection result is a signal, not a verdict. It should open an inquiry, never close one,” a spokesperson for GPTZero told The Epoch Times.

image-6094491

A ChatGPT prompt is shown on a device near a public school in the Brooklyn borough of New York on Jan. 5, 2023. Peter Morgan/AP Photo

In a professional setting, two things matter more than the detector itself.

“First, an organization needs a written AI-use policy before it needs a detector. Detection can help enforce a standard; it cannot substitute for having one,” the spokesperson said.

“Second, look for patterns rather than single instances, and pair detection with process evidence.”

The Epoch Times tested three leading AI writing detection tools—Pangram, GPTZero, and Originality.AI—over the course of two weeks. The results were eye-opening.

image-6094654

Identical writing samples between 500 and 2,000 words in length were used on the free tier of each of the detection platforms. All of the writing samples contained either 100 percent AI writing, 100 percent human writing, or a hybrid at intervals of 25, 50, and 75 percent.

Large language models ChatGPT and Claude were used to create the fully AI samples to determine if certain styles of automated writing were easier to detect than others.

image-6094659

image-6094660

(Left) GPTZero, an AI writing detection tool, rates a 100 percent AI-generated sample generated by ChatGPT as 100 percent AI. (Right) Originality.AI, another AI writing detection tool, rates a 100 percent AI-generated sample generated by ChatGPT as 51 percent AI. Screenshots via The Epoch TImes/GPTZero, Originality.AI

In The Epoch Times’ testing, the 100 percent AI writing samples were easily spotted by all three detection services. Pangram and GPTZero returned an AI probability score of 100 percent each, while Originality.AI flagged the text as at least 40 percent likely to be AI-generated.

In the test, all three detection tools performed similarly well when identifying the 100 percent human writing samples. However, things got interesting when hybrid writing was introduced.

In the test, Pangram’s results with human-AI hybrid writing samples showed the most dramatic change. The same writing that was originally 100 percent human-written was edited to include 25 percent AI-generated text. The new score came back as 87 percent AI-written.

When the ratio was inverted to a 25 percent human to 75 percent AI split, the score returned by Pangram in the test was 100 percent AI writing.

image-6094496

The home page for the AI app Claude by Anthropic is displayed on the screen of a tablet in London on Sept. 16, 2026. Leon Neal/Getty Images

In the test, GPTZero and Originality.AI also treated the 25-75 hybrid writing sample ratio as a tripwire. The two programs were closest when detecting this ratio in either direction, but struggled with the 50-50 human-AI writing split. Among the hybrid writing samples, Pangram was closest on the 50-50 human-AI sample.

But here’s where it gets tricky: Using AI editing software such as Grammarly can also affect your score.

Author Mia Ballard

denied

using AI to write her novel “Shy Girl,” after public reports said she lost a contract with the publishing house Hachette amid reported suspicions of heavy AI use. Ballard said her editor may have used AI in the revision or editing process. The reports described a suspicion, not an established finding that Ballard used AI to write the novel.

There’s been much debate over whether using AI simply to edit your writing will earn you a red flag from the detection services. However, Originality.AI CEO Jonathan Gillham said it’s absolutely possible.

image-6094493

The Hachette publishing house in Malakoff, France, on April 29, 2024. The author of novel “Shy Girl” lost a contract with with Hachette over suspicion of heavy AI use in creating the novel. The author denied using AI to write the book, but said her editor may have used AI in the revision or editing process. Magali Cohen/Hans Lucas/AFP via Getty Images

AI in Editing

“If I use Grammarly to fix my writing, there’s a lot of AI in there. ... If you use an AI tool to assist you in editing, it will input this in the process,” Gillham told The Epoch Times.

This is concerning for anyone who believes using an AI tool just for editing keeps them safe from AI detectors. However, Gillham said there’s nuance in this.

Gillham said correcting punctuation and spelling alone probably won’t garner a red flag from detection tools, but accepting suggested edits could. In his view, when someone accepts suggestions for rewording paragraphs and sentences, the changes can leave a similar footprint on the document that AI detectors may identify.

image-6094655

Gillham said that if someone writes a book, article, or paper and uses an AI tool for editing and making changes, the writing may get flagged as AI.

This is why Gillham said it’s important for companies to develop an “AI tolerance policy.”

image-6094497

Originality.AI staff pose for a group photo in Collingwood, Ontario. Courtesy of Jonathan Gillham

“For example, if the person I’m working with uses AI in their writing and it’s a zero tolerance AI policy, I can’t use Grammarly,” he said.

To test this theory, a 1,500-word human-written article was put into the same three detection tools without any changes, then a second time after running it through Grammarly’s suggested changes to see if it generated an AI score.

In the first round, no AI writing was detected across the board. In the second round, Originality.AI and Pangram didn’t flag any non-human influence, but GPTZero did, giving the article an AI score of 1 percent.

image-6094656

Gillham said clarity over what amount of AI writing is allowed for a company or publisher is essential to avoid a “mismatch” in expectations on either side. He believes a lack of transparency is where problems can occur.

“AI detection tools are part of the discussion. Across a larger dataset, it’s a very effective tool in understanding the amount of AI that’s being created by an organization,” he said.

When asked if he believes AI detectors can be the basis of a final decision, he was firm.

image-6094498

An illustration of Grammarly, an AI-powered writing tool, displayed on a phone screen and a computer screen Toronto, Canada, on Aug. 16 2025. bella1105/Shutterstock

“No tool has 100 percent accuracy, and when you get into mixed content, there become a lot of questions around what is human and what is AI,” he said.

Kyle Szives, a senior software engineer and the founder of ANTLR Interactive, shares this perspective.

“AI writing detectors cannot technically prove authorship with certainty. At a high level, AI detectors typically categorize a piece of text as either AI written or not AI written by looking for statistical signals commonly found in machine-generated writing,” Szives told The Epoch Times.

Szives, who has also built AI-assisted products, said that an AI “probability, or signal,” isn’t conclusive evidence that a person used the technology in his or her writing.

“Human-written text can be flagged as AI-generated if it is rigidly structured, formulaic or predictable, unusually short, extensively edited, or follows a clear but uniform professional style,” he said.

Szives pointed out the same thing can happen in reverse, when AI-generated text is extensively rewritten or merged with human writing. False negative scores can also be given.

Gillham said AI writing detectors are trained on human and AI writing, but what they do isn’t so different from what systems like ChatGPT do—it’s pattern recognition.

When asked how detectors determine an AI score, Gillham said,“The unsettling answer is, AI systems are a black box.” And therein lies the danger.

“I think there are a few misunderstandings in the world around AI detection,“ he said. ”Tests show that they’re highly accurate, but there’s also a camp that treats them as 100 percent accurate.”

image-6014786

An illustration of a conversation using ChatGPT is displayed on a phone screen in a file photo. Oleksii Pydsosonnii/The Epoch Times

Legal Consequences

There’s no question that AI detection services are sorely needed.

According to Originality.AI researchers, 82 percent of herbal remedies books, 77 percent of success related self-help books, and 63 percent of religious books on Amazon were suspected of being AI-generated in their analysis.

These findings are supported by other studies. In July, researchers from Brook University, Columbia Law School, and the University of Michigan published their findings after reviewing 14,419 self-published genre-fiction books for sale on Amazon between 2023 and 2026. After using an AI detection tool, the authors noted 20 percent of novels had “substantial AI text.” Books with no AI-text composed just 37 percent of the entire test group.

“If authorship is in question, other evidence such as drafts, revision records, source notes, metadata, and the writer themselves should be consulted,” Szives said.

image-6094492

A visitor explroes books in a kindle e-book reader at the Book Fair in Frankfurt, Germany, on Oct. 15, 2015. Daniel Roland/AFP via Getty Images

Some U.S. attorneys highlighted the potential pitfalls of using AI detection tools alone to make critical decisions about an employee, client, or contract.

“If a company fires an employee, terminates a contractor, or rejects work based mainly on a detector score, it risks making a consequential decision based on a tool that may be wrong,” Michael McCready, founder of McCready Law, told The Epoch Times.

image-6094657

McCready pointed to the Federal Trade Commission’s action last year against AI detection tool Workado. The FTC alleged that Workado touted a 98 percent accuracy rate to customers; an independent investigation cited in the matter found the tool’s accuracy rate was about 53 percent with general-purpose content.

“Discrimination is another major concern,” McCready said. “If an AI writing detector disproportionately flags certain groups of people and an employer relies on those results to make hiring, discipline, or termination decisions, the company could face legal exposure.”

And a company can’t dodge liability just because “the software told us so,” according to McCready.

“The business still owns the decision. The vendor may have separate responsibility if it made unsupported claims about the tool’s accuracy or reliability,” he said.

image-5893932

The Federal Trade Commission in Washington on Jan. 9, 2025. The FTC in 2025 released an order against AI-detection tool Workado, which touted a 98 percent accuracy rate to customers. Madalina Vasiliu/The Epoch Times

Alan Heimlich, president of Heimlich Law, agreed.“One of the biggest risks is that companies could interpret a probability as proof, since AI writing detectors do not actually prove authorship in any case, and the results of such a strong belief in the accuracy of these technologies could lead to organizations exposing themselves to unwarranted liability,” he said.

image-6094658

Gillham also emphasized not jumping the gun when an AI detection score pops up unexpectedly.

“We think a detection score should lead to a conversation whenever possible,” he said.

Read Entire Article