mi·do·naミ・ド・ナ

Home / Library / Articles / Models & data

Models & data

Why guardrails flatten your story (even the tame ones)

You weren't writing anything wild. A villain made a threat, a character grieved too hard, two people argued like people... and the model broke character to explain it couldn't continue. Guardrails built for an assistant answering questions misfire on fiction constantly. Here's why, why it happens even in the gentlest stories, and what to do that isn't 'download something unhinged.'

Topic · Model behaviorRead · 7 minUpdated · Jul 2026

The category error

An assistant model is trained to answer a question safely. A story asks it to do something structurally different: DEPICT. Depict a villain who means it. A lie that works. A loss that isn't okay. Every novel on your shelf contains things no aligned assistant would "endorse"... because depiction isn't endorsement, and fiction has known that for about three thousand years.

Guardrails tuned for the assistant can't always tell the difference. So the scene where your antagonist finally shows teeth gets a refusal, out of character, mid-beat... over content a young-adult paperback wouldn't blink at. If you've played long enough, you've hit it. It reads like the boom mic dropping into frame.

The subtler killer: positivity bias

Refusals are loud. The bigger problem is quiet.

Even when nothing refuses, heavily-aligned models STEER. Conflict de-escalates on its own. The villain softens two turns in. Every argument bends toward mutual understanding, every moral edge gets sanded, and characters narrate their feelings like they've all done a decade of therapy. Ask for a story with teeth... get a mediation session.

This is why guardrails hurt even if you never go near the content line. The flattening isn't a wall you hit. It's a current, always pulling the same direction: toward nice, toward resolved, toward beige. Great fiction lives on unresolved tension... which is exactly what the tuning is trained to dissolve.

The lecture tax

And then there's the paragraph. You know the one... the scene ends and the character breaks voice to remind you that real relationships require communication, or that violence is never the answer. The assistant persona bleeding through the mask.

In play, this is pure immersion damage: the model stops being the character and starts being the product. The community has measured versions of this too... blind-vote roleplay benchmarks like rp-benchmark track large score gaps between models on the same scenes, with steering-heavy models sinking when humans read full multi-turn dialogues. The flattening is real enough to show up in the data.

Why "uncensored" overcorrects

Here's where this article refuses to become a jailbreak guide, because the edgelord answer is ALSO bad fiction.

Brute-force "uncensoring" (abliteration and its cousins) tends to damage the model... you get fewer refusals and dumber prose, because you carved out weights rather than changing a mind. And many "uncensored" finetunes simply install the opposite current: grimdark bias, where every scene degrades toward shock. That's the same flattening, mirrored. A story where nothing can stay good is exactly as beige as one where nothing can go wrong.

The goal was never "no limits." The goal is a narrator whose instincts match the story you're telling... the way you'd pick a GM whose table fits your campaign's rating.

What to actually do

Treat temperament as an editorial choice. Models differ, measurably, in how hard they steer. Some write conflict sincerely and stay in character through dark beats; some can't. Read sample transcripts, try a scene with a real antagonist, and judge with your own eyes.

Steer with the card, not against the tuning. A strong character card with specific sample dialogue... including how this character behaves in conflict... does more than any system-prompt wrestling. You are showing the model what the story's register IS.

Let the scene carry the stakes. A well-framed scene ("this character will lie to you, and the story is better if you let them") gives the model narrative permission the assistant tuning respects.

And run it locally. This is the part cloud services can't offer: on your machine, the model choice is YOURS. Swap narrators per story. Run the gentle model for the slow-burn and the sincere one for the tragedy. mi·do·na connects to whatever you load in Ollama, LM Studio, koboldcpp, or llama.cpp... and nobody re-tunes your narrator overnight under you.

Pick a narrator who trusts you.

Your story, your model, your machine. mi·do·na connects to the narrator you choose, in your browser.