September 2026 review · same 50 texts

Grubby vs Undetectable.ai

In the September 2026 HumanizerEval review, Grubby and Undetectable.ai rewrote the same 50 English texts. Grubby cleared all 5 detectors on 18 of them; Undetectable.ai on 15. The difference is mostly Pangram: on 9 samples Grubby passed it where Undetectable.ai failed, and the reverse happened on 7. On GPTZero, Originality.ai, ZeroGPT and Winston AI the two were close to even.

Per-sample agreement

Per-sample agreement, September 2026. Counts are out of the 50 shared samples, computed at the 50% human-score threshold. Margin counts the samples only one tool passed, higher count first.
DetectorPassed more oftenMarginBoth passedNeither passed
Pangram 4.0Grubby9 to 7925
GPTZero 2026-09-13-baseUndetectable.ai1 to 0472
Originality.ai turboGrubby3 to 1460
ZeroGPT defaultTied0 to 0500
Winston AI 5.0Grubby1 to 0490
All 5 detectorsGrubby9 to 6926

Every count is computed from scores.csv at build time, not transcribed. “Neither passed” counts the texts that defeated both tools; on Pangram that was 25 of 50 here.

Quality and output discipline

Quality and output discipline, September 2026.
MeasureGrubbyUndetectable.ai
72%26%
46%28%
98%94%
94%78%
1.111.21
07
00
AcademicUniversity / General Writing / More Human, v11sr
Websitegrubby.ai undetectable.ai

On per-sample quality scores, Grubby scored higher on 33 samples, Undetectable.ai on 4, and the two tied on 13. Grubby’s output was shorter on 34 of 50 samples; mean length ratios were 1.11 and 1.21, or 11% longer than the input and 21% longer than the input respectively.

Where each one outperforms the other

Grubby

  • Pangram: passed 18 of 50 against 16, including 9 samples where Undetectable.ai failed, against 7 the other way.
  • Originality.ai: passed 49 of 50 against 47, including 3 samples where Undetectable.ai failed, against 1 the other way.
  • ZeroGPT: tied at 50 of 50 but on a higher mean human score, 96.0% to 95.0%. It passed 0 samples where Undetectable.ai failed.
  • Winston AI: passed 50 of 50 against 49, including 1 sample where Undetectable.ai failed, against 0 the other way.
  • Factual consistency: 72% against 26%.
  • Naturalness: 46% against 28%.
  • Syntax integrity: 98% against 94%.
  • Stance preservation: 94% against 78%.
  • Length: shorter output on 34 of 50 samples, which matters under a word limit.

Undetectable.ai

  • Pangram: passed 7 samples where Grubby failed, even though it was detected more often overall, passing 16 to 18.
  • GPTZero: passed 48 of 50 against 47, including 1 sample where Grubby failed, against 0 the other way.
  • Originality.ai: passed 1 sample where Grubby failed, even though it was detected more often overall, passing 47 to 49.

Grubby vs Undetectable.ai FAQ

Is Grubby better than Undetectable.ai?

In the September 2026 HumanizerEval review, Grubby cleared all 5 detectors on 18 of the 50 shared samples and Undetectable.ai on 15. On GPTZero, Originality.ai, ZeroGPT and Winston AI the two were within 6 points of each other. Quality ran 33 samples to 4, with 13 ties.

Which humanizer is best for Pangram, Grubby or Undetectable.ai?

Grubby. In September 2026, Grubby passed Pangram on 18 of the 50 shared samples and Undetectable.ai on 16. On 9 samples Grubby passed where Undetectable.ai failed; the reverse happened on 7. Both failed on 25.

Which humanizer changes the text less?

Grubby. It scored higher on per-sample quality on 33 of 50 samples, against 4. Grubby’s outputs averaged 1.11× the input length and passed factual consistency on 72%; Undetectable.ai’s averaged 1.21× and passed on 26%. On stance preservation the split was 94% to 78%. Grubby produced the shorter output on 34 of 50 samples.

More comparisons