September 2026 review · same 50 texts

StealthGPT Super vs Undetectable.ai

In the September 2026 HumanizerEval review, StealthGPT Super and Undetectable.ai rewrote the same 50 English texts. StealthGPT Super cleared all 5 detectors on 46 of them; Undetectable.ai on 15. The difference is mostly Pangram: on 30 samples StealthGPT passed it where Undetectable.ai failed, and the reverse never happened. On GPTZero, Originality.ai, ZeroGPT and Winston AI the two were close to even.

Per-sample agreement

Per-sample agreement, September 2026. Counts are out of the 50 shared samples, computed at the 50% human-score threshold. Margin counts the samples only one tool passed, higher count first.
DetectorPassed more oftenMarginBoth passedNeither passed
Pangram 4.0StealthGPT30 to 0164
GPTZero 2026-09-13-baseStealthGPT2 to 1470
Originality.ai turboStealthGPT2 to 0471
ZeroGPT defaultUndetectable.ai1 to 0490
Winston AI 5.0Tied1 to 1480
All 5 detectorsStealthGPT31 to 0154

Every count is computed from scores.csv at build time, not transcribed. “Neither passed” counts the texts that defeated both tools; on Pangram that was 4 of 50 here.

Quality and output discipline

Quality and output discipline, September 2026.
MeasureStealthGPT SuperUndetectable.ai
68%26%
72%28%
98%94%
96%78%
1.041.21
07
00
Super, production outputUniversity / General Writing / More Human, v11sr
Websitestealthgpt.ai undetectable.ai

On per-sample quality scores, StealthGPT scored higher on 37 samples, Undetectable.ai on 2, and the two tied on 11. StealthGPT’s output was shorter on 43 of 50 samples; mean length ratios were 1.04 and 1.21, or 4% longer than the input and 21% longer than the input respectively.

Where StealthGPT outperforms Undetectable.ai

StealthGPT

  • Pangram: passed 46 of 50 samples, against 16 for Undetectable.ai.
  • GPTZero: passed 49 of 50 samples, against 48 for Undetectable.ai.
  • Originality.ai: passed 49 of 50 samples, against 47 for Undetectable.ai.
  • Winston AI: higher mean human score than Undetectable.ai, 97.7% against 97.5%, with both passing 49 of 50 samples.
  • Factual consistency: kept facts accurate in 68% of outputs, against 26% for Undetectable.ai.
  • Naturalness: read fluently in 72% of outputs, against 28% for Undetectable.ai.
  • Syntax integrity: stayed grammatically clean in 98% of outputs, against 94% for Undetectable.ai.
  • Stance preservation: preserved the source’s stance in 96% of outputs, against 78% for Undetectable.ai.
  • Length: shorter than Undetectable.ai on 43 of 50 samples, which matters under a word limit.

Undetectable.ai

  • Pangram: failed 34 of 50 samples, 30 of which StealthGPT passed.
  • GPTZero: failed 2 of 50 samples, 2 of which StealthGPT passed.
  • Originality.ai: failed 3 of 50 samples, 2 of which StealthGPT passed.
  • Winston AI: lower mean human score than StealthGPT, 97.5% against 97.7%, with both passing 49 of 50 samples.
  • Factual consistency: introduced factual errors in 74% of outputs.
  • Naturalness: had awkward or clunky phrasing in 72% of outputs.
  • Syntax integrity: had grammar or punctuation errors in 6% of outputs.
  • Stance preservation: distorted the source’s stance in 22% of outputs.
  • Length: longer than StealthGPT on 43 of 50 samples, so more to cut under a word limit.

StealthGPT vs Undetectable.ai FAQ

Is StealthGPT better than Undetectable.ai?

Yes. In the September 2026 HumanizerEval review, StealthGPT Super cleared all 5 detectors on 46 of the 50 shared samples and Undetectable.ai on 15. On GPTZero, Originality.ai, ZeroGPT and Winston AI the two were within 6 points of each other. Quality ran 37 samples to 2, with 11 ties.

Which humanizer is best for Pangram, StealthGPT or Undetectable.ai?

StealthGPT Super. In September 2026, StealthGPT Super passed Pangram on 46 of the 50 shared samples and Undetectable.ai on 16. On 30 samples StealthGPT passed where Undetectable.ai failed; the reverse never happened. Both failed on 4.

Which humanizer changes the text less?

StealthGPT Super. It scored higher on per-sample quality on 37 of 50 samples, against 2. StealthGPT Super’s outputs averaged 1.04× the input length and passed factual consistency on 68%; Undetectable.ai’s averaged 1.21× and passed on 26%. On stance preservation the split was 96% to 78%. StealthGPT produced the shorter output on 43 of 50 samples.

More comparisons