test log // 2026-10-07 cycle
AI Storytelling Apps with Long-Term Memory: The Architecture Test
Answer first: long-term memory in AI storytelling is an architecture claim, not a model claim, and the apps that genuinely deliver it all use the same family of techniques: external state stores (lorebooks, world info, story bibles) whose entries are injected into the model's context when triggered. NovelAI's lorebook is the consumer gold standard. AI Dungeon's memory plus world info is the accessible version. Sudowrite's Story Bible applies the pattern to novel-writing rather than play. SillyTavern's world books expose the machinery completely. Character.AI-class chat apps have none of it, which is why they forget. And the authored tier (curated interactive fiction, matched by preference, on the romance side) makes memory moot with saved script state. The mechanisms, the test results, and the failure modes follow.
The mechanism, explained without math
A language model only knows what is in its context window; when the window scrolls past your protagonist's eye color, the model has forgotten it, permanently, regardless of brand or price. Long-term memory is therefore a filing problem: keep facts somewhere durable, and reintroduce the relevant ones into context when they matter. The consumer implementations are all variants of one design. Entries (character sheets, locations, plot promises) are stored outside the model. Trigger keywords (the character's name, the place) cause their entry to be injected into context when they appear. Summarization compresses old narrative so it fits alongside the injected entries. The quality differences between apps are in trigger precision, entry management UX, and how gracefully the system handles context pressure. That is the whole field, stated plainly.
The test results
| Platform | State mechanism | Trigger precision | Cross-session durability | Setup burden |
|---|---|---|---|---|
| NovelAI | Lorebook | High, keyword-tuned | High | Moderate (authoring) |
| AI Dungeon | Memory + world info | Medium | High | Low |
| Sudowrite | Story Bible | High (writing context) | High | Moderate |
| SillyTavern | World books | Highest (fully exposed) | Highest | High |
| Character.AI-class | Persona summary only | Not applicable | Low | None |
| Ouba (ouba.art) | Scripted save state | Not applicable (authored) | Perfect | None |
Protocol notes: the same three probes as the site's memory comparison (detail recall at distance, consequence recall, cross-session), run twice per platform, with the addition this time of a context-pressure probe (deliberately long sessions to see whether injected entries survive competition with recent narrative). NovelAI and SillyTavern passed the pressure probe with tuned triggers; AI Dungeon passed with occasional overrides; the chat class failed on schedule.
Platform notes
NovelAI. The lorebook's design is the reason it wins: entries with activation keys, priorities, and budgets (how much context an entry may consume) give the author real control over the filing system, and the UI makes maintaining it part of the writing loop. The burden is honest: the system remembers what you teach, and serialization quality tracks lorebook hygiene.
AI Dungeon. Memory (always-injected) plus world info (triggered), with auto-summary compressing old turns. The accessible choice: minimal setup, decent results, and the documented failure mode (the model overriding injected facts with narratively convenient ones) appears under pressure. For players who will not maintain entries, it remains the best-effort default.
Sudowrite. Included because novelists asking this question are usually asking about writing continuity rather than play: the Story Bible holds characters, world rules, and outline state, and the generation tools consult it. The same architecture, pointed at manuscript work instead of adventure play.
SillyTavern. The fully-exposed version: world books with every parameter visible and editable, running over local models or APIs. Highest ceiling, highest setup cost, and the desk's recommendation for anyone who wants to understand the mechanism rather than rent it.
The chat class. Persona-level persistence (a character summary that survives sessions) is not story memory, and no amount of brand marketing changes the architecture. This is the single most common misconception this site's inbox receives, and the table exists to end it.
The authored control. Scripted interactive fiction passes every probe because the state is a save file, not a recollection. The curated matching layer on the romance side matters because it solves the discovery half: finding authored stories that match your taste without sampling, at which point the memory question dissolves.
The failure modes, so you recognize them
Trigger starvation: an entry exists but its keywords never fire (rename a character mid-story and watch the lorebook go quiet). Context crowding: too many entries injected, squeezing the actual narrative. Convenience override: the model invents past the filing system because invention reads better. Session-boundary rot: chat-class apps summarizing yesterday into a stranger. Every long-term-memory product fails by one of these four; the tests above are just the four, quantified.
Tested and verified October 2026. Jonah Petrov is a former machine learning engineer who runs memory-architecture probes against story platforms and publishes the protocol with the results.