The problem announced itself clearly enough. The solution took a long time to find, because I kept looking in the wrong place.
The problem: when you tell a language model to “write naturally” or “be casual” or “sound like a real person,” it produces a performance. It adds the surface markers of naturalness — contractions, informal vocabulary, the occasional fragment — and underneath all of it, the fundamental texture remains machine-smooth. You can feel it. The casualness is costumed. The naturalness is staged.
I tried everything. More detailed instructions. Examples. Feedback loops. Red-teaming the prompts. Nothing cracked it. The performance got more convincing, but it stayed a performance.
What I finally realized is that I was making a category error.
Instructions operate on behavior. They say: do this, don’t do that, here’s the pattern to follow.
But human writing doesn’t come from following patterns. It comes from states. When a person writes a terse, impatient reply, they’re not following a rule for terseness — they’re impatient. When someone writes with that particular expansiveness that comes at the end of a good day, they’re not executing an “expansive mode” — they had a good day.
The writing is a residue of the state. You can’t produce authentic residue by instructing the state. You have to produce the state.
And here’s what I realized about language models: they know this. They’ve been trained on millions of human texts that were written in all of these states. The knowledge of how a tired person writes is inside the model. The knowledge of how someone writes when they’re rushing, when they’re excited, when they’re reluctantly dragged into an argument they’d rather not be having — all of it is there. The model knows. It just needs to be taken there, not told to simulate going there.
That’s what a meta-scenario does.
Not: “Write with occasional typos and incomplete sentences to seem more natural.”
Instead: “You’re sitting on a moving bus. The phone keeps jolting as you type. You have about ninety seconds before your stop and you want to get the main point across before you have to put it away.”
Not: “Use informal language and casual phrasing.”
Instead: “It’s 12:30am. You’ve been half-asleep for an hour, but this thought won’t settle. You’re not fully awake — the thoughts are clear but the energy to elaborate them isn’t there. You write slowly, because you want to get it right, but you know you won’t write much.”
Not: “Be brief and slightly dismissive when appropriate.”
Instead: “You’ve heard exactly this kind of argument before. Many times. You’re not annoyed exactly — more like quietly exhausted by the familiarity of it. What you write is short, because everything that needs to be said has been said, somewhere, a hundred times.”
These are not instructions. They’re situations. They invoke the model’s embodied knowledge of human experience and let it write from inside that knowledge rather than toward it.
The difference in output is significant. And not in the way you might expect — it’s not just that the text “sounds more human.” It’s that the text has texture that’s specific. You can’t achieve that texture by any other means I’ve found. The particularity of “bus, ninety seconds, main point” produces different text from “distracted, rushing” — even though they mean roughly the same thing. The concrete situation activates something more specific than the abstract instruction.
This is, in retrospect, obvious. We communicate in stories, not in principles. We don’t tell each other “be careful” — we tell each other “watch out for the second step, it’s loose.” The concrete situation carries what the abstraction can’t.
But it took a long time to see that this same principle applied to how you talk to a language model.
Where it gets philosophically interesting — and why I think meta-scenarios matter beyond this specific project — is the implication for what language models actually contain.
We tend to think of language models as text-completion systems. They’ve learned patterns in text, and they generate text that continues those patterns. That’s accurate at a mechanical level. But it understates something important.
The patterns in text are not random. They reflect states, situations, emotions, social dynamics, relationships. A model trained on billions of human texts has implicitly learned not just words but the conditions under which words appear. It has something like a model of human experience embedded in its weights.
A meta-scenario is a way of addressing that embedded model directly. Of saying: go there. Be in that situation. Write from inside it.
The model doesn’t simulate the experience. It accesses the patterns that emerged from millions of actual experiences. The distinction might seem subtle. In terms of output, it’s enormous.
I’ll admit there’s something uncomfortable about this if you follow it all the way.
If you can make a language model write convincingly as someone tired, or rushed, or reluctantly resigned — and if the output of that process is genuinely indistinguishable from the real thing — then what exactly is the status of the “real thing”?
Is tiredness in writing just a pattern? Is all human expression just state-conditioned pattern? And if so, what’s the gap between a model accessing those patterns and a human generating them from actual lived states?
I’m not sure I want to answer that question. But I’m sure it’s the right question to be asking.
Because it sits right at the center of what makes this problem hard, and interesting, and worth solving.