How to Test Branching Stories: A Complete Template for State Matrices, Path Coverage, and Regression Testing
Test branching stories with a state dictionary, node test cases, pairwise coverage, and critical paths to establish a reproducible regression process.

Introduction
Testing branching stories cannot be completed by simply “playing through every route once.” The number of paths grows rapidly, and many defects arise from combinations of states. A more reliable approach is to verify state transitions first, then cover critical combinations through risk-prioritized path testing, and finally turn each defect into a regression test case.
1. Build a State Dictionary
| State | Range | Default | Written at | Read at | Visible Feedback |
|---|---|---|---|---|---|
| trust_A | -1/0/1 | 0 | N03/N06 | N08/END | Dialogue and assistance |
Remove variables that are never read, and supply the missing sources for conditions that read values with no defined origin.
2. Create Node Test Cases
Each test case records the starting node, preconditions, actions, expected next node, state changes, and screenshots or logs. Timed choices also require tests for timeouts, input at the time boundary, and repeated clicks; media nodes require tests for loading failures and recovery.
3. Use Pairwise Coverage Instead of Exhaustive Testing
Prioritize critical combinations of two variables, such as high or low relationship values and whether the player holds a key, or QTE success or failure and whether the player has obtained a clue. This cannot replace testing all high-risk combinations, but it can uncover common conditional errors at a lower cost.
4. Six Critical Paths
The default main route, the shortest route to each ending, the route with the fewest resources, the route where all QTEs fail, the save/chapter recovery route, and regression routes for past severe defects.
5. Defect Report Template
Version/platform:
Starting node and preconditions:
Steps to reproduce:
Actual/expected results:
Frequency:
Screenshots or logs:
Before release, at minimum, ensure that all severe defects are closed, every ending can be reproduced twice, save upgrades and chapter jumps have been verified, and the risks of combinations that remain uncovered are clearly identified.
The State Dictionary Must Come Before Full Route Testing
Testing with only a story graph tells you that the player moved from A to B, but not which variables were written. A state dictionary defines each variable's type, default value, valid range, and write and read nodes. Whenever a variable is renamed or its range changes, the test cases must be updated accordingly.
Relationship values are particularly prone to getting out of control. Do not simply write “trust increases”; specify the starting and ending values, whether there is a cap, which scenes read the value, and how the interface provides feedback. If an ending condition is trust > 3, both boundary values, 3 and 4, must be tested.
Node Test Cases Must Cover Abnormal Input
Beyond normal choices, test repeated clicks, input at the countdown boundary, network disconnections, switching to the background, media loading failures, rapid skipping, language switching, and save restoration. Common defects involve states being written repeatedly or not written at all after abnormal actions, rather than errors in the story itself.
Start each test case from a reproducible save. Merely writing “enter Chapter 3 and click on the left” is insufficient for reproduction, because Chapter 3 may inherit different relationships and items.
How to Tier Path Coverage
P0 covers the main route, all endings, save corruption, and content safety; P1 covers important relationships, items, QTEs, and chapter jumps; P2 covers local dialogue and low-risk visual differences. Run node tests and P0 smoke tests with every commit, then run P1/P2 tests on release candidates.
Do not report coverage solely as the number of nodes visited. Also report state transitions, ending conditions, recovery from abnormal situations, and untested combinations. Visiting a node does not mean all of its conditions are correct.
Regression Test Cases Come from Real Defects
Whenever a problem is fixed, add its original preconditions and actions to the regression suite. If the problem arose from default values during a chapter jump, every future version must verify that scenario; otherwise, similar errors will reappear after refactoring.
For example, a player obtains a key at N05, but after jumping to N08, the system uses the default has_key=false. The report should include a comparison of the full sequence and the chapter-jump sequence, the save version, and logs; after the fix, separately verify new saves, old saves, and chapter restarts.
What a Release Report Should Explain
The report lists the endings that passed, critical paths, blocking defects, remaining risks, and the version. Avoid unauditable statements such as “the entire flow has been tested.” The goal is not to claim that there are no bugs, but to let decision-makers know which causal relationships have been verified and which combinations remain unknown.
Assign Risk Levels to Paths
Classify the first playthrough of the main route, payment entry points, irreversible choices, and final endings as the highest risk, and test them with every build; run daily regression tests on common side routes and major convergence points; rotate rare combinations between versions. Priority is determined jointly by user impact, likelihood, repair cost, and past defects. Do not test only the routes that are easiest to automate.
Each test case should clearly specify the initial save, action steps, expected state, visible result, and cleanup procedure. On failure, preserve the build number, node, state snapshot, screenshots or video, and shortest reproduction path. If the only description is “sometimes jumps to the wrong ending,” developers will struggle to determine whether the problem lies in condition evaluation, save migration, or media loading.
Version Upgrades Must Be Tested with Old Saves
When adding states, renaming nodes, or adjusting default values, load old saves from a released version into the new build and check how missing fields are populated, whether unlocked content is retained, and whether reverting to an earlier save bypasses critical migrations. Passing tests with a fresh save does not mean an upgrade is safe. If compatibility cannot be maintained, explain the impact in advance and provide an understandable solution.
Acceptance meetings should address only evidence-backed results: which paths have been tested, which state boundaries have been covered, and why the remaining gaps are acceptable. Explicitly listing uncovered areas is more valuable than creating a sense of safety with a vague overall pass rate.


