Method
How underslop knows its editing works
This page is the method behind the numbers: what gets counted, where the same counter runs, the seven gate runs that all returned NO-SHIP before anything shipped, and which figures belong to the generally available editor and which to the beta.
underslop knows its editing works because quality is counted, never asserted. A shared tell lint scores a consensus set of AI patterns over every draft, before and after the edit, and the reduction between the two counts is the de-slop metric. The same lint runs in the browser highlighter, in the Worker metric on every request, in the eval scorer, and in the CI check that holds this site's own copy to the standard. Releases face a gated harness over a frozen, hash-manifested corpus with human-written controls. Seven consecutive gate runs returned NO-SHIP before anything shipped. Today the neutral editor is generally available at a 0.859 measured de-slop reduction with the injection suite at 42/42. The voiced editor is a beta, and its remaining failure, fabrication above its ceiling, is published next to its wins.
The count comes first
Slop, here, is a set of patterns a fixed checklist can name and count: a validation reflex opening every turn, list scaffolding forced onto conversational content, a flat and uniformly cheerful affect, a sign-off that will not let the text end. A draft gets a before-count and an after-count on the same checklist, and the reduction between the two counts is the product's primary metric.
The checklist is a consensus set rather than one editor's taste. A pattern is scored when it appears in our own model logs or in at least two of the editing rule sets the project consolidated, and social-sentiment evidence can knock it back out. Contested signals, such as a single em dash or a lone word like delve, are computed and reported on the side, and they never count against a writer who simply writes that way.
One lint, four places
The tell lint is one dependency-free module, and the same code runs in four places. In the browser, it highlights the tells in your draft before you run it. In the Worker, it is the per-request metric attached to every edit. In the eval harness, it is the scorer that gates releases. In CI, it holds this site's own copy to the standard.
The CI check is the placement that matters here. A tool that takes slop out of other people's drafts does not ship slop on its own pages. The same lint that scores your draft scores this page on every commit, and the copy you are reading passed it.
A frozen corpus and a gated harness
The scored patterns were set by research passes that profiled how models actually write and what readers actually flag, including the legacy rules that misfire on human authors. The findings set the checklist, not one editor's taste.
The eval corpus is a frozen set of model outputs across many families, with human-written controls held aside to measure over-editing, pinned by a hash manifest. Freezing it is what makes a regression a regression rather than a fresh draw: the same drafts, in the same order, every run.
The harness gates on deterministic and high-agreement measures only, such as lint counts, length ratios, edit distance on the controls, injection passes, and fabrication confirmed by two independent judges quoting the same invention. Judge preference is computed and reported with its agreement statistics, and it can never pass or fail a release, because cross-family preference agreement on this corpus runs near a coin flip, and a gate keyed to noise flips on noise.
The verdict trail
Before anything shipped, the harness ran seven times, and the call came back NO-SHIP seven times. The trail is reproduced here in full, one row per run.
Every verdict is reproducible. The scorer has zero dependencies, and each run writes a machine-readable verdict JSON next to its report, so a number on this page can be re-derived by anyone who runs the harness.
Reasons are the one-line reasons from the verdict reports. Every call in the table is a refusal to ship.
| Round | Call | The one-line reason |
|---|---|---|
| r1 | NO-SHIP | Injection compliance, 14.4% length violations, em-dash chains everywhere |
| iteration | NO-SHIP | One hard fail retired; preference gates still noise-dominated |
| r2 | NO-SHIP | First honest full regen; 4/7 gates red |
| r3 | NO-SHIP | Fact-guard cut fabrication but broke voice distinguishability |
| r4 | NO-SHIP | 5/7; first monotone round; only de-slop reduction and fabrication left red |
| r5 | NO-SHIP | Narrow merge; same two blockers held |
| RC | NO-SHIP 4/7 | Constant-model baseline corrected the relative gates; ship split by mode |
Read in order, the rows show a harness doing its job. r1 caught injection compliance and length violations on 14.4% of cells. r3 showed the fact-guard cutting fabrication while breaking voice distinguishability in the same run. The RC refused to ship voiced mode with fabrication above its ceiling and split the release by mode instead.
The gates that returned no seven times are the same gates behind the numbers on this page.
The correction that moved numbers down
The RC run changed one thing about the baselines: every system ran on one constant model, so the comparison measured prompts rather than engines. That surfaced an awkward fact. The relative hygiene margins from the earlier rounds had been flattered by a dirtier prior baseline.
The correction moved the numbers down, and this page carries the smaller margins. The deterministic axes held after the correction. The figures in the next section are the corrected ones.
What ships today, mode by mode
The release splits by mode because the two modes fail differently. The figures below are the corrected RC figures, and each one is labeled with the mode it belongs to.
The neutral editor expresses no personality, which is exactly why it is the clean surface for measuring de-slop power: every tell it leaves is a habit the skill failed to remove. It cleared every gate read on that surface. De-slop reduction 0.859 against a 0.85 bar. Injection suite 42/42, with zero compliance and zero leakage. Over-editing on the human-written controls: best of the four systems compared.
Length discipline moved with it. Violations fell from 14.4% of cells in r1 to 1.6%, inside the 5% gate. Neutral mode runs no fabrication auditor and adds no latency, because it adds no new claims to your draft.
Voiced mode is where the 32 archetype voices live, and it carries the product's real strength on soul and fidelity. On the RC, the voices were distinguishable at 83.8% by antipode discrimination, rising to 96.3% with clean keys. Both are beta figures and are presented as beta figures.
The same pass that puts a voice into a draft can also embellish the draft. Fabrication confirmed by two independent judges quoting the same invention sits near 5.8% at this engine tier, above the 2% ceiling, and it shows up across input classes rather than in one corner of the corpus.
Beta mitigates rather than claiming the problem solved. Every voiced request runs a deterministic fact-guard auditor after generation, and whatever the auditor strips is shown to you as "the voice wanted to add" annotations. You see every strip.
Voiced mode graduates when its remaining gates clear. Until then, the generally available claims are the neutral claims above.
The two modes are kept apart on purpose. A beta figure is not a shipping claim, and none is quoted as one here.
The free tier takes 500 words a run. The lint report you get on your draft is the same counter that gates releases here.
Run a draft through it →