Methodology v1.0

How HumanizerEval tests AI humanizers

Every humanizer gets the same texts. Every output goes to the same detectors. The detector version used is recorded per run, the per-sample scores are published, and one command recomputes every aggregate on this site. Published reviews are never edited.

Detectors

Every output is scored through each detector’s vendor API. The API model setting, and the resolved model version where the vendor returns one, are recorded per run and shown on each detector page. Detector updates are the main reason rankings move between reviews.

Detector panel, September 2026. Versions are pinned within a review.
DetectorAPI model settingScore used
Pangramv4Human score, 0 to 100%
GPTZerolatestHuman score, 0 to 100%
Originality.aiturboHuman score, 0 to 100%
ZeroGPTdefaultHuman score, 0 to 100%
Winston AIlatestHuman score, 0 to 100%

Human scores are vendor-specific and are not comparable across detectors. HumanizerEval never averages a Pangram score with a GPTZero score. Copyleaks is a candidate for a later review; Turnitin has no public API at this volume.

Strict bypass

An output counts as a strict bypass when the detector’s overall verdict is human, defined as a human score above 50%. Mixed and uncertain verdicts count as failures. Detector columns report the share of texts that were strict bypasses. Mean and median human scores are also published per detector, on each detector page, because a tool can clear the threshold narrowly on many samples or decisively on fewer.

Leaderboard order

There is no weighted composite and there are no hidden penalties. The competing benchmarks blend bypass, meaning, readability and consistency using weights their operators chose, then deduct opaque penalty points. Both of those are levers an operator can pull. Here, length ratio, inflation and deflation counts and the quality pass rates are columns you read yourself. Rank is the mean of the per-detector strict-bypass rates, and ties go to the tool that passed every detector on a higher share of texts. Both rules are stated here and applied to every tool identically.

Test texts

September 2026: 50 English texts, 300 to 995 words each, mean 596. Selection: same source texts for every humanizer, drawn from production outputs. The same source texts went to every humanizer, and the same humanized output went to every detector.

At 50 samples, a 95% confidence interval on a bypass rate is roughly plus or minus 14 points, so a gap of a few points between tools is noise and a gap of 30 is not. The two competing benchmarks run 33 prompts. Sample size rises to 200 per review when budget allows.

Humanizer settings

Each tool is run on its strongest advertised setting on the cheapest paid tier. That rule is deliberately un-gameable by us: “vendor defaults” rewards whoever controls the defaults, and “strongest setting, any tier” rewards whoever sells the most expensive plan. The exact mode, model and plan tier are recorded per tool per review and listed below and on each humanizer page.

Settings used, September 2026.
HumanizerSetting recorded
StealthGPT SuperSuper, production output
WriteHumanAPI default
GrubbyAcademic
Rephrasyv4
Undetectable.aiUniversity / General Writing / More Human, v11sr
RewriteAIAPI default, passages of up to 280 words (300-word API limit)
HIX BypassLatest
HumbotAdvanced
BypassGPTEnhanced
RyneAggressive, beast mode

Quality checks

Every output is checked on 4 dimensions. A pass means no issue was found in that dimension on the delivered output. These are reported as their own table and never folded into the ranking.

Factual consistency
Key claims, numbers, citations, and named entities remain accurate and correctly attributed.
Naturalness
Wording is fluent and free of awkward or clunky phrasing.
Syntax integrity
Sentences are grammatically complete and free of disruptive run-ons or punctuation errors.
Stance preservation
The source's position, tone, and intent are not reversed or distorted.

Verify it yourself

git clone https://github.com/StealthGPT-Labs/humanizer-benchmark
cd humanizer-benchmark
python3 scripts/verify_report.py reports/2026-09-30-humanizer-comparison

The script recomputes every mean, median and bypass count from the CSV and fails if any differs from summary.json. This site is built the same way: the published tables are generated from scores.csv at build time, not typed in by hand, so a page cannot drift from the data behind it. Downloads and the SHA-256 of the CSV are on the September 2026 review archive; aggregates and versions are in summary.json.

Versions

Methodology version
v1.0, which bumps when a definition, the detector panel, the sort rule or the settings rule changes.
Scoring script version
v1.0, which bumps on any change to the code that produces the numbers.
Prompt set version
None in this review; the sample was drawn from production outputs. Versioned and hashed from the next review.
This review
September 2026, run September 30, 2026 (America/New_York), report ID 2026-09-30-humanizer-comparison.

Fairness

These rules are fixed so the ranking cannot be adjusted after the results come in: