FicFarestructured test logs for story platforms

test log // 2026-10-07 cycle

AI Story Apps That Remember Your Choices: Tested

By Jonah Petrov · test cycle October 2026

Answer first: memory is where AI story apps genuinely differ, and after structured testing across the major platforms, the ranking is as follows. NovelAI holds the most durable long-form continuity (the lorebook architecture persists facts across sessions deterministically). AI Dungeon runs second with its world info and memory system. Character.AI remembers within a conversation and degrades across long ones, with no reliable cross-session story state. The chat-tier apps (Chai, Talkie) behave like Character.AI with weaker continuity tooling. And the authored-interactive tier (Choices-style apps, plus curated platforms) sidesteps the problem entirely: the story is scripted, so your choices persist because someone wrote them into the branch, not because a model recalled them. Test logs below.

Why memory is the hard part

A language model conditions on its context window: the recent tokens it can see. Everything outside the window is gone, which is why an AI story forgets your eye color by chapter six and your plot by chapter twelve. The engineering answers are external state: saved facts injected into context (AI Dungeon's memory and world info), retrieval systems that surface relevant material (embedding search over your story history), structured lorebooks (NovelAI's keyword-triggered entries), and session summarization. The app-layer distinction that matters to a reader: none of these are magic, they are filing systems with different failure modes, and the tests below measure the failures.

The test protocol

Every platform got the same three probes, run on the same prompts. Probe one: the detail check, a fact introduced early (a name, an object) tested for recall after two thousand words of story. Probe two: the consequence check, a choice made explicitly offered back to the narrative later. Probe three: the session check, continuity across a closed and reopened story. Scoring is binary per probe; three platforms were tested twice to confirm stability. What follows is the summary, not the raw logs, but the logs exist.

The results

PlatformDetail recallConsequence recallCross-sessionArchitecture
NovelAIPass, deterministic with lorebookPass with entriesPassKeyword-triggered lorebook
AI DungeonPass with memory/world infoMostlyPassMemory + world info
Character.AIPass in-windowDegrades after long chatsWeak, bot-persona levelContext + persona memory
ChaiSimilar to C.AI, shorter horizonDegradesWeakChat context
TalkieSimilar, more persona driftDegradesWeakChat context
Ouba (ouba.art)Pass, authoredPass, scripted branchesPass, saved stateCurated authored stories

That final row functions as the control group: authored interactive fiction passes every probe by construction. When readers complain that AI stories forget, part of what they miss is the guarantee that scripted branching provides, and the curated platforms (trope-matched authored stories with saved state) are the way back to that guarantee when generation quality is not the point.

Platform notes

NovelAI. The lorebook is the differentiator: you write entries (characters, places, rules), tag them with trigger keywords, and the system injects them when triggers fire. Deterministic in the sense that the same trigger reliably surfaces the same entry. The cost is authoring labor: the system remembers what you taught it, and teaching it is part of the hobby. For long-form serialized play, this is the current state of the art on the prosumer tier.

AI Dungeon. Memory (always-injected facts) plus world info (keyword-triggered, the lorebook's ancestor), wrapped in the most accessible interface in the category. Its failure mode is prompt pressure: in long sessions, injected facts compete with recent narrative for context, and the model sometimes overrides the filing system with what reads better. Playable, fun, imperfect.

Character.AI. The chat-native experience: personas remember what happened in the conversation because it is in the window, and bot-level memory persists a thin persona summary across sessions. The story-relevant fact (the promise your character made) is not what it persists. For roleplay-flavored conversation this is fine; for serialized fiction it is the category's weakest continuity.

Chai and Talkie. Same architecture class as Character.AI with different moderation postures and catalogs. Continuity conclusions transfer across all three; the choice between them is about community and content policy, covered in the comparison dedicated to that question.

The verdict

If cross-session, serialized, fact-dense story continuity is the requirement, NovelAI's lorebook (with effort) or AI Dungeon's memory system (with tolerance) are the credible options, and no pure chat platform currently qualifies. If the requirement is actually "a story where my choices matter and nothing gets forgotten," the honest answer is authored interactive fiction, and the curated matching layer (the romance-side curated catalogs are the representative) is how you find the right one without sampling blind. The technology improves quarterly; the architecture distinction has held across every test cycle this year.

Tested and verified October 2026. Jonah Petrov is a former machine learning engineer who runs context-window and memory-retention tests on story platforms and publishes the methodology with every result.