We built the detector we wanted to be judged by
Most AI detectors hand back a percentage and leave you to argue about it. GBT Zero shows the working: which sentences moved the score, and which measurement made them move.
Why a single number was never enough
A detector that says "87% AI" gives you nothing to act on. You cannot rewrite a percentage. You cannot show a percentage to a writer and have a useful conversation about it. And when the number is wrong, which it sometimes will be, there is no way to see why.
So GBT Zero reports per sentence. Red marks lines with strong statistical markers of generation, amber marks mixed passages, green marks natural human cadence. Underneath each result sits the measurement that produced it: how predictable the wording was, how much the sentence lengths varied, how repetitive the structure became.
That changes what the tool is for. It stops being a verdict machine and starts being a diagnostic, which is the only honest thing a statistical classifier can be.
It also makes the tool useful on writing that is entirely your own. Flat burstiness is a real prose problem whether or not a model wrote the draft.
How a detection baseline gets built
A baseline is the reference distribution a passage gets compared against. Getting it wrong is how detectors end up flagging honest writers.
-
01
Collect matched pairs
For each model family and each language, we gather human writing and machine writing on the same topics, in the same registers, at the same lengths. Comparing a technical manual against casual blog prose teaches the model the wrong thing.
-
02
Measure, do not guess
Every passage is scored on the four signals: word predictability, sentence-length variance, structural repetition and semantic movement. Those distributions become the reference, per model, per language.
-
03
Calibrate the threshold against false positives
The cost of wrongly flagging a real writer is much higher than the cost of missing a generated passage. Thresholds are set with that asymmetry in mind rather than tuned for a headline accuracy figure.
-
04
Recheck when models change
Every major model release shifts the distribution. A baseline calibrated against last year of GPT output quietly stops working. Recalibration is maintenance, not a feature.
Claims we will not make
The detection market runs on confident numbers. Here is what we think those numbers are worth, and where we have decided to stay quiet instead.
- A headline accuracy percentage
- Accuracy depends entirely on the evaluation set. Any vendor can pick a set that produces a flattering number. Until we publish the set alongside the figure, the figure would be marketing, not evidence.
- Proof of who wrote something
- A detector measures statistical properties of text. It has no access to authorship. Treating output as proof is the single most common way this technology causes harm.
- Reliable results on short passages
- Under roughly 150 words there is not enough variance to measure. We would rather tell you the sample is too short than return a confident number built on nothing.
- Equal reliability for every writer
- Detection over-flags non-native English writing across the entire industry, ours included. Simpler vocabulary and steadier sentence length genuinely resemble low-perplexity output. We flag this rather than hide it.
Your text is not our training data
Scans run in memory and the text is discarded when the response is sent. Nothing is written to a database, nothing is kept for review, and nothing is used to train a model. That is a design decision, not a plan setting, and it applies to free scans exactly as it applies to everything else.
Read the privacy policy- Text retained after a scan
- None
- Account required
- No
- Used for model training
- Never
Try it on something you wrote
The fastest way to judge a detector is to run it on text whose origin you already know.