September 2026 review · run September 30, 2026

The AI humanizer benchmark that tests against Pangram.

In the September 2026 review, StealthGPT Super has the best bypass rate among AI humanizers. It passed Pangram on 92% of texts, against 65% for the next best tool.

September 2026 AI humanizer leaderboard

AI humanizer comparison, September 2026. Sorted by the mean strict-bypass rate across the 5 detectors. Every tool rewrote the same 50 texts, and every output went to every detector.
#HumanizerPangramGPTZeroOriginality.aiZeroGPTWinston AI
1StealthGPTPassed all 5 detectors on 46 of 5096.8%92%98%98%98%98%83.5%1.04
2WriteHumanPassed all 5 detectors on 32 of 4992.7%65%100%98%100%100%58.7%1.02
3GrubbyPassed all 5 detectors on 18 of 5085.6%36%94%98%100%100%77.5%1.11
4RephrasyPassed all 5 detectors on 20 of 5084.0%46%88%86%100%100%54.5%1.13
5Undetectable.aiPassed all 5 detectors on 15 of 5084.0%32%96%94%100%98%56.5%1.21
6RewriteAIPassed all 5 detectors on 13 of 5077.2%28%86%88%88%96%72.0%1.09
7HIX BypassPassed all 5 detectors on 1 of 5066.4%2%68%64%100%98%46.5%1.11
8HumbotPassed all 5 detectors on 0 of 5066.4%0%62%70%100%100%47.5%1.12
9BypassGPTPassed all 5 detectors on 0 of 5064.4%0%56%66%100%100%49.0%1.12
10RynePassed all 5 detectors on 0 of 5023.6%22%8%0%62%26%79.0%1.07

Each column counts how many of the 50 texts that detector called human-written. Rank is the average of the five columns. WriteHuman returned 49 of 50 outputs, so its rates are out of 49. Full methodology.

Bypass by detector, and bypass against quality

The first chart shows each tool’s strict-bypass rate on every detector. The second plots mean bypass against quality pass rate.

  • Pangram
  • GPTZero
  • Originality.ai
  • ZeroGPT
  • Winston AI
Strict-bypass rate by detector for each humanizer in the September 2026 review.0%25%50%75%100%StealthGPTStealthGPT Super, Pangram: 92%StealthGPT Super, GPTZero: 98%StealthGPT Super, Originality.ai: 98%StealthGPT Super, ZeroGPT: 98%StealthGPT Super, Winston AI: 98%WriteHumanWriteHuman, Pangram: 65%WriteHuman, GPTZero: 100%WriteHuman, Originality.ai: 98%WriteHuman, ZeroGPT: 100%WriteHuman, Winston AI: 100%GrubbyGrubby, Pangram: 36%Grubby, GPTZero: 94%Grubby, Originality.ai: 98%Grubby, ZeroGPT: 100%Grubby, Winston AI: 100%RephrasyRephrasy, Pangram: 46%Rephrasy, GPTZero: 88%Rephrasy, Originality.ai: 86%Rephrasy, ZeroGPT: 100%Rephrasy, Winston AI: 100%Undetectable.aiUndetectable.ai, Pangram: 32%Undetectable.ai, GPTZero: 96%Undetectable.ai, Originality.ai: 94%Undetectable.ai, ZeroGPT: 100%Undetectable.ai, Winston AI: 98%RewriteAIRewriteAI, Pangram: 28%RewriteAI, GPTZero: 86%RewriteAI, Originality.ai: 88%RewriteAI, ZeroGPT: 88%RewriteAI, Winston AI: 96%HIX BypassHIX Bypass, Pangram: 2%HIX Bypass, GPTZero: 68%HIX Bypass, Originality.ai: 64%HIX Bypass, ZeroGPT: 100%HIX Bypass, Winston AI: 98%HumbotHumbot, Pangram: 0%Humbot, GPTZero: 62%Humbot, Originality.ai: 70%Humbot, ZeroGPT: 100%Humbot, Winston AI: 100%BypassGPTBypassGPT, Pangram: 0%BypassGPT, GPTZero: 56%BypassGPT, Originality.ai: 66%BypassGPT, ZeroGPT: 100%BypassGPT, Winston AI: 100%RyneRyne, Pangram: 22%Ryne, GPTZero: 8%Ryne, Originality.ai: 0%Ryne, ZeroGPT: 62%Ryne, Winston AI: 26%Strict-bypass rate, % of texts
Bypass against quality for each humanizer in the September 2026 review. Quality ranges across 37 points and bypass across 73.0%25%50%75%100%0%25%50%75%100%StealthGPT Super: 96.8% bypass, 83.5% qualityWriteHuman: 92.7% bypass, 58.7% qualityGrubby: 85.6% bypass, 77.5% qualityRephrasy: 84.0% bypass, 54.5% qualityUndetectable.ai: 84.0% bypass, 56.5% qualityRewriteAI: 77.2% bypass, 72.0% qualityHIX Bypass: 66.4% bypass, 46.5% qualityHumbot: 66.4% bypass, 47.5% qualityBypassGPT: 64.4% bypass, 49.0% qualityRyne: 23.6% bypass, 79.0% qualityStealthGPTWriteHumanGrubbyRephrasyUndetectable.aiRewriteAIHIX BypassHumbotBypassGPTRyneMean strict bypass →Quality pass rate, mean of 4 →

Performance by detector

Undetectability score is model-dependent, so each detector gets its own page and its own ranking. A tool that bypasses one detector can be detected on another.

Pangram is the hardest detector in the panel

In September 2026, StealthGPT Super passed Pangram v4 on 46 of 50 texts with a mean human score of 88.9%. Across the panel, the average tool passed Pangram on 32% of texts, against 76% on GPTZero, 76% on Originality.ai, 95% on ZeroGPT and 92% on Winston AI. Neither of the two other public humanizer benchmarks tests it, which means both overstate every tool they rank.

Quality, reported separately

Bypassing AI detectors is worthless if the humanized text changes the meaning. Each output was checked with 4 factors in mind; a pass means no issue was found. These numbers never enter the ranking, so a weighting choice cannot decide the order.

Quality pass rates, September 2026. A pass means no issue was found in that dimension on the delivered output. Mean is the average of the 4 pass rates. Reported separately from bypass, and never blended into the ranking.
Humanizer
StealthGPT Super83.5%68%72%98%96%
WriteHuman58.7%22%49%98%65%
Grubby77.5%72%46%98%94%
Rephrasy54.5%36%30%92%60%
Undetectable.ai56.5%26%28%94%78%
RewriteAI72.0%54%40%100%94%
HIX Bypass46.5%18%6%96%66%
Humbot47.5%22%0%98%70%
BypassGPT49.0%22%2%94%78%
Ryne79.0%44%80%98%94%

In September 2026, StealthGPT Super had the highest mean quality pass rate at 83.5%, and HIX Bypass the lowest at 46.5%.

Which humanizer for which job

Best for students
Pangram is used by Capella University, Substack, Quora, NewsGuard, Penguin Random House, and others. In this review StealthGPT Super led Pangram at 46 of 50.
Best for marketers and SEO teams
StealthGPT Super is the best for bypassing, including Google’s AI detection, which helps content rank higher. Google is less likely to rank pages it reads as machine-written.
Best for publishers and writers
StealthGPT Super keeps a human touch in the writing, which is what a byline needs. Quality is the strongest in the panel: its mean quality pass rate was 83.5%.

Recompute this page yourself

Every number above comes from one public CSV of per-sample detector scores. Clone the repo and recompute the aggregates:

git clone https://github.com/StealthGPT-Labs/humanizer-benchmark
cd humanizer-benchmark
python3 scripts/verify_report.py reports/2026-09-30-humanizer-comparison

Why another humanizer benchmark

There are two other public humanizer benchmarks. One is run by WriteHuman and ranks WriteHuman first. One is run by UndetectedGPT and ranks UndetectedGPT first. They share a methodology template almost line for line, and neither tests Pangram, the one detector that reliably separates a strong humanizer from a weak one.

HumanizerEval is the first and only benchmark that tests against Pangram. It is the truest and most comprehensive benchmark available to give a full picture of humanizer performance.

Frequently asked questions

What is an AI humanizer?

An AI humanizer rewrites machine-generated text so that AI-content detectors are less likely to flag it. Most run a language model with a paraphrasing prompt plus post-processing that alters perplexity, burstiness and the other statistical patterns detectors look for, while trying to keep the original meaning intact. Whether that works depends almost entirely on which detector you face.

Which AI humanizer is best for Pangram?

In September 2026, StealthGPT Super passed Pangram v4 on 46 of 50 texts (92%). The next best was WriteHuman at 65%. See the Pangram page.

Why does Pangram give such different results from GPTZero?

Detectors use different models and different thresholds, and their human scores are not comparable. Pangram ships a classification head trained specifically to recognise humanizer output; GPTZero does not advertise one. On the same 50 texts in September 2026, the average tool passed GPTZero on 76% of texts and Pangram on 32%. That is why HumanizerEval reports every detector separately and never blends their scores.

Why should I trust these results?

Because you can recompute them. Every tool ran on the same 50 texts, through the same detector versions, with the settings recorded per tool. The per-sample scores are in a public CSV, and one command recomputes every number on this page.

Why is Turnitin not tested?

Turnitin offers no public API for this kind of testing, so it cannot be run at 50 samples per tool per review without a licensed institutional account. However, independent testing shows that StealthGPT Super bypasses Turnitin ~95% of the time.