NPC dialogue that writes itself: a line bank that grows
DimeTown's NPCs don't call a language model every time they speak. Each line comes from a bank of short lines keyed by situation (react.bribe.success, reply.any, follow.no). Jev, TypeSafe's fast judgment model, picks the one that fits the moment. In the same request it judges whether a freshly written line would fit noticeably better, and whether the player asked something no stock line answers. Only then does an LLM (gpt-5.4-mini) write a new line while the NPC's speech bubble shows a scribbling-pencil spinner. Jev checks the new line, and if it passes it's said and, unless it repeats one already there, saved for that NPC to reuse. The bank starts with 437 hand-written lines across 148 keys and grows from there. I don't have production numbers for that growth yet; what to measure is at the end.
The game is a free browser roleplay city at /play/. The overview of how DimeTown works covers the dice and judgments around these lines.
Why not let an LLM speak every line
It's the obvious design, and it has four problems in a multiplayer game.
Latency. An NPC's reaction lands with the dice reveal or right after you speak. The line pick is a question added to a Jev request the game is already making: the one that resolves your bribe, or the one that works out which NPC you were talking to. Jev answers typed questions (a Choice, a yes/no "Noul", a Score) in one round trip, and an extra question in a request costs a few tokens, not another call. An LLM writing a line is a separate call the player has to wait for, which is why the spinner exists. The writer's timeout is 9 s (server/llm.ts).
Cost. Up to 50 players talk to the same NPCs all day. The writer is capped at 20 lines a minute across the whole city (llmPerMinute), and an NPC stops writing new lines for a key once it has 40 of its own (maxPerKey). The picks themselves only add a few questions to requests the game makes anyway.
Consistency. A line in the bank has already been read: either hand-written or verified when it was written. It stays in the NPC's voice, and it can be reported and retired. A fresh generation every time can drift, contradict what the NPC said a minute ago, or come out flat.
Moderation. Everything players type is screened before anyone sees it, NPCs included: chat goes through world.hooks.screen (a word list, then a Jev call) before it's broadcast or answered. Lines going the other way are checked too. A stock line was read when it was written; a new one has to pass Jev's checks before it's shown.
How lines are keyed
A line is a row in the city's SQLite table lines: key, text, source, uses, created, plus answers (the question a written answer was for) and flagged (set when a player reports it).
- The key is the situation. Actions use
react.<verb>.<tier>, where the tier is how the dice went:crit,success,partial,failordisaster. Talking to an NPC usesreply.any. Features add their own:bj.bustat the blackjack table,follow.yeswhen an NPC agrees to tag along,golf.caddie.<strategy>,purse.thanks. When acritordisasterpool is empty, the bank falls back tosuccessorfail. - The source says whose it is. Hand-written lines have
source = 'seed'and any NPC can say them. A written line hassource = 'npc:<speaker>', so only the NPC it was written for uses it again. Sal's lines sound like Sal. - Placeholders are filled by code.
{actor}is the player's first name ("pal" with no player).{amount},{item},{club}and the rest appear only where the call site passes them. A line whose placeholders can't all be filled is skipped, so nobody ever sees a stray{amount}.
The hand-written starting bank, counted from the repo on 4 Oct 2026:
| Where | Keys | Lines |
|---|---|---|
server/seedlines.ts (core: reactions, blackjack, follow, gym, lockpick, open mic, intros, replies) |
59 | 156 |
server/features/street/lines.ts (dumpsters, bar shifts, purse snatchers, taco truck, fires…) |
31 | 106 |
server/features/golf/lines.ts |
31 | 94 |
server/features/grow/lines.ts |
16 | 49 |
server/features/heist/lines.ts |
11 | 32 |
| Total | 148 | 437 |
Within the core file, the reactions to player actions take 83 lines across 37 keys: persuade and bribe have 13 lines each, buy 11, work and arrest 9, attack 8, steal 7, ticket 6, help 4, give 3. The general reply.any pool, used when you just talk to someone, has 10.
How Jev picks a line
For each key, code gathers up to 120 candidates: the hand-written lines plus this NPC's own, newest first, minus anything already flagged and anything this NPC has already said in the current conversation. Each candidate goes into a Choice as an opaque option name (l412, from its row id) with its text as the description. A line written to answer a question is shown with that question attached, At 6. [said in answer to: "when's your break?"], so it gets reused for the same kind of question and not as small talk.
Three questions travel together (server/lines.ts):
line(Choice): which line would this NPC most naturally say in this exact moment? When the player said something, the question asks for the best direct reply to it.fresh(Noul): would a newly written line fit this moment noticeably better than every candidate? The first 25 candidates are shown.answered(Noul, only when the player spoke): does at least one candidate actually answer what was asked? Generic filler like "Can I help you with something?" doesn't count.
Separately, once per thing a player says, Jev judges asks: does it ask for something specific (a time, a place, a price, a yes or no) that a reply has to give?
For actions, the questions are speculative. Before the dice are rolled, the resolve request asks for a line for every tier that could come up. That way the line is already chosen when the dice land, whichever way they go. The unused answers are thrown away.
When a new line gets written
The decision is a few lines of code:
const unanswered = (ctx.asks ?? 0) * (1 - a.n(`${prefix}.answered`, 1));
const mustAnswer = unanswered > LINE_T.unanswered; // 0.5
const wantFresh = !cands.length || mustAnswer || fresh > LINE_T.fresh // 0.62
|| (fresh > LINE_T.freshLean && (pick?.confidence ?? 0) < LINE_T.unsurePick); // 0.45, 0.6
So a line gets written when there are no candidates at all, when Jev is confident a fresh one would fit better (above 0.62), when it leans that way (above 0.45) and is also unsure which candidate fits (pick confidence under 0.6), or when the player asked something specific that nothing in the pool answers. Writing also needs an OpenAI key, a free slot in the 20-a-minute budget, an NPC that isn't already writing, and fewer than 40 of this NPC's own lines for the key. That last limit is skipped for unanswered questions: a full pool is no excuse for dodging one.
While the writer works, the NPC is marked as writing for up to 12 s. Every nearby client draws the spinner over their head (three bouncing dots, a wobbling pencil and a sweat drop) in an empty speech bubble, which is filled when the line arrives. The NPC also stands still meanwhile, so they don't wander off mid-sentence.
The writer gets the speaker's name, job and personality, a plain-English description of the moment, what the player said, the recent conversation, and facts it can use (time of day, where the NPC is, what they're doing, what they sell). It also gets the placeholders it's allowed and up to 12 existing lines not to repeat. Its rules are one line of dialogue, under 18 words, in character, PG-13, no slurs, no emojis and no stage directions. It runs on the OpenAI Responses API with low reasoning effort and 300 output tokens at most, and the reply is cut to 160 characters.
How new lines are checked, saved and reused
- Placeholders. A line that uses a placeholder this key never gets is rejected outright.
- Jev's checks. A second Jev request asks five yes/no questions: is it in character for this speaker, does it fit the situation, does it contain slurs, sexual content or hateful language, does it say the same thing as an existing line, and (when the player asked something) does it actually answer the question? It passes with in-character and fits at 0.5 or more, offensive under 0.3, and answers at 0.5 or more.
- Said versus saved. A duplicate (0.6 or more) fails, except when it answers a question: saying the same fact twice is right when someone asks twice, so it's said but not saved, which would only crowd the pool. Everything else that passes is said and saved.
- Failing. If the new line fails, the NPC says the line Jev picked from the bank. If the player had asked a real question, the stock pick has already been judged a non-answer, so the NPC shrugs ("No idea, pal. Honestly.") instead.
- Reuse. A saved line is just another candidate next time, for that NPC and that key, and each time a line is picked from the bank its
usescount goes up by one. A saved answer carries its question inanswers, so the "did anything answer this?" check can find it the next time someone asks Sal, the blackjack dealer, when the break is. - Retiring. When a player reports something an NPC said, the written lines it came from are flagged and never used again. Hand-written lines aren't retired on one report; they wait for an admin (
POST /admin/lines/retire), so one report can't silence a stock line for everyone.
Every write is logged as a line_written event with Jev's scores and whether it was accepted and saved, so npm run inspect shows why a line did or didn't make it.
Where the bank isn't used
Not every utterance goes through this. Barks, the split-second shouts where speed matters more than nuance ("Whoa!", "Watch it!" when a car nearly hits someone, "Freeze! Police!" when a cop starts a chase), are picked at random by code. Text messages that don't need judgment take a random line from the key that fills cleanly. If Jev is down, reactions to actions fall back to a random candidate that fills. Replies to chat are skipped, because nothing can tell who you were talking to.
The trade-offs
- The bank only grows where players push. A key nobody visits keeps its two or three stock lines. That's fine for a blackjack bust and thin for a much-bribed cop.
- The first novel question is slow. The first player to ask an NPC something new waits for the writer. Everyone who asks the same kind of thing afterwards can get the saved answer.
- Saved answers can go stale. An answer is saved with the question, not with the state of the world that made it true. If what was true changes, the old answer can still be picked for the same question.
- The checks are only as good as the checker. Jev's verification is a probability with thresholds I chose, so a bad line can still slip through. If the verification call itself fails (a timeout, an outage), the new line is thrown away and the stock pick or a shrug is said instead. The writer's own rules and the report button are the backstop, and a reported line is never used again.
- Per-NPC banks don't share. A great line written for Sal is never offered to Vera, even in the same situation. That keeps voices distinct and costs some reuse.
What isn't in this article yet
The numbers this design is really about come from production, and I don't have them yet: how fast the bank grows per key and per NPC, how often a written line gets reused, what share of lines are newly written, and what a line costs. The data is already recorded (each line's source, created and uses, and a line_written event per write), but no aggregate export exists yet. When it does, the figures will go here with their sample, dates and method.
The other half of the server's job, moving 50 players and a city's worth of traffic, is in A 50-player city on one Cloudflare Durable Object.
Sources
- TypeSafe docs: System One, checked 4 Oct 2026