September 2026 review · same 50 texts

Grubby vs Humbot

In the September 2026 HumanizerEval review, Grubby and Humbot rewrote the same 50 English texts. Grubby cleared all 5 detectors on 18 of them; Humbot on 0. The difference is mostly Pangram: on 18 samples Grubby passed it where Humbot failed, and the reverse never happened. On ZeroGPT and Winston AI the two were close to even.

Per-sample agreement

Per-sample agreement, September 2026. Counts are out of the 50 shared samples, computed at the 50% human-score threshold. Margin counts the samples only one tool passed, higher count first.
DetectorPassed more oftenMarginBoth passedNeither passed
Pangram 4.0Grubby18 to 0032
GPTZero 2026-09-13-baseGrubby16 to 0313
Originality.ai turboGrubby15 to 1340
ZeroGPT defaultTied0 to 0500
Winston AI 5.0Tied0 to 0500
All 5 detectorsGrubby18 to 0032

Every count is computed from scores.csv at build time, not transcribed. “Neither passed” counts the texts that defeated both tools; on Pangram that was 32 of 50 here.

Quality and output discipline

Quality and output discipline, September 2026.
MeasureGrubbyHumbot
72%22%
46%0%
98%98%
94%70%
1.111.12
00
00
AcademicAdvanced
Websitegrubby.ai humbot.ai

On per-sample quality scores, Grubby scored higher on 39 samples, Humbot on 0, and the two tied on 11. Grubby’s output was shorter on 26 of 50 samples; mean length ratios were 1.11 and 1.12, or 11% longer than the input and 12% longer than the input respectively.

Where each one outperforms the other

Grubby

  • Pangram: passed 18 of 50 against 0, including 18 samples where Humbot failed, against 0 the other way.
  • GPTZero: passed 47 of 50 against 31, including 16 samples where Humbot failed, against 0 the other way.
  • Originality.ai: passed 49 of 50 against 35, including 15 samples where Humbot failed, against 1 the other way.
  • Winston AI: tied at 50 of 50 but on a higher mean human score, 99.3% to 97.4%. It passed 0 samples where Humbot failed.
  • Factual consistency: 72% against 22%.
  • Naturalness: 46% against 0%.
  • Stance preservation: 94% against 70%.
  • Length: shorter output on 26 of 50 samples, which matters under a word limit.

Humbot

  • Originality.ai: passed 1 sample where Grubby failed, even though it was detected more often overall, passing 35 to 49.
  • ZeroGPT: tied at 50 of 50 but on a higher mean human score, 96.7% to 96.0%. It passed 0 samples where Grubby failed.

Grubby vs Humbot FAQ

Is Grubby better than Humbot?

In the September 2026 HumanizerEval review, Grubby cleared all 5 detectors on 18 of the 50 shared samples and Humbot on 0. On ZeroGPT and Winston AI the two were within 6 points of each other. Quality ran 39 samples to 0, with 11 ties.

Which humanizer is best for Pangram, Grubby or Humbot?

Grubby. In September 2026, Grubby passed Pangram on 18 of the 50 shared samples and Humbot on 0. On 18 samples Grubby passed where Humbot failed; the reverse never happened. Both failed on 32.

Which humanizer changes the text less?

Grubby. It scored higher on per-sample quality on 39 of 50 samples, against 0. Grubby’s outputs averaged 1.11× the input length and passed factual consistency on 72%; Humbot’s averaged 1.12× and passed on 22%. On stance preservation the split was 94% to 70%. Grubby produced the shorter output on 26 of 50 samples.

More comparisons