September 2026 review · run September 30, 2026
The AI humanizer benchmark that tests against Pangram.
In the September 2026 review, StealthGPT Super has the best bypass rate among AI humanizers. It passed Pangram on 92% of texts, against 65% for the next best tool.
September 2026 AI humanizer leaderboard
| # | Humanizer | Pangram | GPTZero | Originality.ai | ZeroGPT | Winston AI | |||
|---|---|---|---|---|---|---|---|---|---|
| 1 | StealthGPTPassed all 5 detectors on 46 of 50 | 96.8% | 92% | 98% | 98% | 98% | 98% | 83.5% | 1.04 |
| 2 | WriteHumanPassed all 5 detectors on 32 of 49 | 92.7% | 65% | 100% | 98% | 100% | 100% | 58.7% | 1.02 |
| 3 | GrubbyPassed all 5 detectors on 18 of 50 | 85.6% | 36% | 94% | 98% | 100% | 100% | 77.5% | 1.11 |
| 4 | RephrasyPassed all 5 detectors on 20 of 50 | 84.0% | 46% | 88% | 86% | 100% | 100% | 54.5% | 1.13 |
| 5 | Undetectable.aiPassed all 5 detectors on 15 of 50 | 84.0% | 32% | 96% | 94% | 100% | 98% | 56.5% | 1.21 |
| 6 | RewriteAIPassed all 5 detectors on 13 of 50 | 77.2% | 28% | 86% | 88% | 88% | 96% | 72.0% | 1.09 |
| 7 | HIX BypassPassed all 5 detectors on 1 of 50 | 66.4% | 2% | 68% | 64% | 100% | 98% | 46.5% | 1.11 |
| 8 | HumbotPassed all 5 detectors on 0 of 50 | 66.4% | 0% | 62% | 70% | 100% | 100% | 47.5% | 1.12 |
| 9 | BypassGPTPassed all 5 detectors on 0 of 50 | 64.4% | 0% | 56% | 66% | 100% | 100% | 49.0% | 1.12 |
| 10 | RynePassed all 5 detectors on 0 of 50 | 23.6% | 22% | 8% | 0% | 62% | 26% | 79.0% | 1.07 |
Each column counts how many of the 50 texts that detector called human-written. Rank is the average of the five columns. WriteHuman returned 49 of 50 outputs, so its rates are out of 49. Full methodology.
Bypass by detector, and bypass against quality
The first chart shows each tool’s strict-bypass rate on every detector. The second plots mean bypass against quality pass rate.
- Pangram
- GPTZero
- Originality.ai
- ZeroGPT
- Winston AI
Performance by detector
Undetectability score is model-dependent, so each detector gets its own page and its own ranking. A tool that bypasses one detector can be detected on another.
Best AI humanizer for Pangram
92%
StealthGPT Super
Best AI humanizer for GPTZero
98%
StealthGPT Super (tied)
Best AI humanizer for Originality.ai
98%
StealthGPT Super (tied)
Best AI humanizer for ZeroGPT
98%
StealthGPT Super (tied)
Best AI humanizer for Winston AI
98%
StealthGPT Super (tied)
Pangram is the hardest detector in the panel
In September 2026, StealthGPT Super passed Pangram v4 on 46 of 50 texts with a mean human score of 88.9%. Across the panel, the average tool passed Pangram on 32% of texts, against 76% on GPTZero, 76% on Originality.ai, 95% on ZeroGPT and 92% on Winston AI. Neither of the two other public humanizer benchmarks tests it, which means both overstate every tool they rank.
Quality, reported separately
Bypassing AI detectors is worthless if the humanized text changes the meaning. Each output was checked with 4 factors in mind; a pass means no issue was found. These numbers never enter the ranking, so a weighting choice cannot decide the order.
| Humanizer | |||||
|---|---|---|---|---|---|
| StealthGPT Super | 83.5% | 68% | 72% | 98% | 96% |
| WriteHuman | 58.7% | 22% | 49% | 98% | 65% |
| Grubby | 77.5% | 72% | 46% | 98% | 94% |
| Rephrasy | 54.5% | 36% | 30% | 92% | 60% |
| Undetectable.ai | 56.5% | 26% | 28% | 94% | 78% |
| RewriteAI | 72.0% | 54% | 40% | 100% | 94% |
| HIX Bypass | 46.5% | 18% | 6% | 96% | 66% |
| Humbot | 47.5% | 22% | 0% | 98% | 70% |
| BypassGPT | 49.0% | 22% | 2% | 94% | 78% |
| Ryne | 79.0% | 44% | 80% | 98% | 94% |
In September 2026, StealthGPT Super had the highest mean quality pass rate at 83.5%, and HIX Bypass the lowest at 46.5%.
Which humanizer for which job
- Best for students
- Pangram is used by Capella University, Substack, Quora, NewsGuard, Penguin Random House, and others. In this review StealthGPT Super led Pangram at 46 of 50.
- Best for marketers and SEO teams
- StealthGPT Super is the best for bypassing, including Google’s AI detection, which helps content rank higher. Google is less likely to rank pages it reads as machine-written.
- Best for publishers and writers
- StealthGPT Super keeps a human touch in the writing, which is what a byline needs. Quality is the strongest in the panel: its mean quality pass rate was 83.5%.
Recompute this page yourself
Every number above comes from one public CSV of per-sample detector scores. Clone the repo and recompute the aggregates:
git clone https://github.com/StealthGPT-Labs/humanizer-benchmark
cd humanizer-benchmark
python3 scripts/verify_report.py reports/2026-09-30-humanizer-comparisonscores.csv
50 rows of anonymized per-sample detector scores, output word counts and quality scores for all 10 tools. SHA-256 on the review archive.
leaderboard.json
The full ranking, per-detector results, quality pass rates and pairwise counts as machine-readable JSON.
Methodology v1.0
Detector versions, the strict-bypass definition, the sort rule, the settings rule, and the gaps we have not closed yet.
Why another humanizer benchmark
There are two other public humanizer benchmarks. One is run by WriteHuman and ranks WriteHuman first. One is run by UndetectedGPT and ranks UndetectedGPT first. They share a methodology template almost line for line, and neither tests Pangram, the one detector that reliably separates a strong humanizer from a weak one.
HumanizerEval is the first and only benchmark that tests against Pangram. It is the truest and most comprehensive benchmark available to give a full picture of humanizer performance.
Frequently asked questions
What is an AI humanizer?
An AI humanizer rewrites machine-generated text so that AI-content detectors are less likely to flag it. Most run a language model with a paraphrasing prompt plus post-processing that alters perplexity, burstiness and the other statistical patterns detectors look for, while trying to keep the original meaning intact. Whether that works depends almost entirely on which detector you face.
Which AI humanizer is best for Pangram?
In September 2026, StealthGPT Super passed Pangram v4 on 46 of 50 texts (92%). The next best was WriteHuman at 65%. See the Pangram page.
Why does Pangram give such different results from GPTZero?
Detectors use different models and different thresholds, and their human scores are not comparable. Pangram ships a classification head trained specifically to recognise humanizer output; GPTZero does not advertise one. On the same 50 texts in September 2026, the average tool passed GPTZero on 76% of texts and Pangram on 32%. That is why HumanizerEval reports every detector separately and never blends their scores.
Why should I trust these results?
Because you can recompute them. Every tool ran on the same 50 texts, through the same detector versions, with the settings recorded per tool. The per-sample scores are in a public CSV, and one command recomputes every number on this page.
Why is Turnitin not tested?
Turnitin offers no public API for this kind of testing, so it cannot be run at 50 samples per tool per review without a licensed institutional account. However, independent testing shows that StealthGPT Super bypasses Turnitin ~95% of the time.