Home / Library / Articles / Models & data
Models & data
Slop, measured: why AI prose sounds like AI
You know slop when you read it. The testament to. The palpable tension. The air that's thick with everything. What you might not know: someone actually measured it... built word lists from model output, compared them against human writing, and put numbers on the smell. Here's what the data says, and what a storyteller can do with it.
You know it when you read it
Every long-form AI player develops the allergy eventually. The character whose eyes SPARKLE with mischief, twice a scene. The tension you could cut with a knife. The barely above a whisper. Individually harmless... cumulatively, the unmistakable smell of text no human chose.
The community named it slop, and for years it stayed a vibe. Then the people behind EQ-Bench did something better than complaining: they counted.
How you measure a smell
The Slop Score works on a simple idea: some words and patterns occur far more often in AI text than in human text, and you can catch them statistically. The lists were built by analyzing outputs from ten different language models against human-authored writing, using an open toolkit (slop-forensics)... so this isn't one person's pet peeves, it's over-representation you can reproduce.
The score itself is a weighted composite:
| Component | Weight | What it catches |
|---|---|---|
| Slop words | 60% | Individual words LLMs overuse (the sparkles, the testaments) |
| Not-x-but-y patterns | 25% | The contrast tic: "it wasn't fear, but something deeper" |
| Slop trigrams | 15% | Three-word phrases that fingerprint AI prose |
The not-x-but-y line deserves its 25%. Once you see the pattern, you cannot unsee it... it's the single most reliable tell in AI fiction, and models lean on it hardest exactly when they're trying to sound profound.
The leaderboard ranks models on this, lower meaning more human-like patterns, and there's an analysis tool that scores any text you paste. Two honest caveats from the authors themselves: it's not an AI detector, and it's tuned for creative writing and essays. Which is to say... tuned for exactly what we do.
Why roleplay gets it worst
Slop compounds. In a one-shot essay, a testament to slides by. In a 200-turn story, the model's favorite phrases come back around like a carousel, and repetition is the fastest immersion-killer there is... worse than any plot hole, because it breaks the VOICE, and the voice is the whole game.
That's also why prose quality is a legitimate reason to pick one model over another. The differences are real and measurable: models vary widely on slop and repetition metrics, and community finetunes often exist precisely because someone got sick of a base model's tics.
What you can actually do
Pick with your eyes open. Before committing an evening to a model, check its slop and repetition numbers on the Creative Writing leaderboard and read its sample outputs. The samples tell you more than the score.
Tune the samplers. Slop is partly a sampling problem: the overused word is overused because it's always the highest-probability token. This is exactly what min-p and temperature shape, and it's why mi·do·na's Imagination dial exists... pushed higher, the model reaches past its first-choice words. It won't cure slop. It audibly thins it.
Steer the card. A character card with strong, specific sample dialogue gives the model a voice to imitate that ISN'T its default register. Generic cards get generic prose.
And rewind. When a turn comes back sloppy, redo it. In mi·do·na that's one tap... the fastest slop filter ever shipped is the retry button with your taste attached.
The honest close
No model is slop-free... the lists were built from ALL of them. And a score can't measure charm, humor, or whether a character sounds like themselves at turn 190. Treat the numbers as a floor check, not a verdict, and your own eyes as the real benchmark. Running locally makes that test cheap: load a model, run YOUR scene, read ten turns. You'll know.
Your taste is the benchmark.
Load a model on your machine, run a scene that matters to you, and let your own eyes judge the prose. mi·do·na makes the test cheap.