How One Text Node Supports Four Endings: Mixed-Media Design That Reduces Video Costs in Interactive Film Games
Combine video, text, still images, chat, and voice to create four endings while preserving the consequences of player choices and controlling production costs.

Introduction
An all-video experience is not inherently more immersive. Filming evidence reading, internal judgments, phone chats, and changes across time as video increases costs, makes information harder to revisit, and slows revisions. The principle of mixed media is this: video handles performances and actions that must be seen, while text handles information that needs comparison, recall, and rapid iteration.
How one node can support four endings
Suppose the protagonist receives a chat log before the final confrontation. A text node displays three pieces of evidence already obtained and lets the player choose “Make everything public,” “Share only with the ally,” “Delete the log,” or “Pretend not to know.” This action is recorded in evidence_action. The final video can still reuse the same venue and opening performance, then deliver the four states through different phone screens, two lines of dialogue, and ending narration.
Four endings do not require filming four complete short films. The most expensive video carries the shared climax, while differences between states are presented through combinations of interchangeable short shots, text, sound, and epilogues.
What each medium does best
| Medium | Suitable for | Unsuitable for |
|---|---|---|
| Video | Facial expressions, actions, danger, relationship breakdowns | Long passages of searchable information |
| Text | Evidence, inner thoughts, summaries, rules | Climaxes that rely on physical performance |
| Still images | Establishing locations, files, memories | Continuous complex actions |
| Chat interface | Timelines, relationships, asynchronous conflict | Replacing face-to-face dialogue in every scene |
| Voice | Private information, off-screen space | Key evidence without subtitles |
How to transition between mixed-media nodes
Before entering text, use a character action to establish its physical source, such as picking up a phone or unfolding a document. When leaving it, have the next video begin with the same object and physical posture. With a narrative reason, a media switch will not feel like a loading screen.
A sample budget allocation
For a single 10-minute playthrough, around 70% of the total duration of unique assets could be reserved for the shared video backbone, 15% for key branching videos, and 15% for text, still images, chat, and epilogue variants. These proportions are only a planning example and should be adjusted according to performance value and tool costs; they should not be treated as industry pricing.
When text enhances rather than merely saves
When players need to stop and compare evidence, revisit names, or think about promises, text provides a greater sense of control than continuous footage. It can also support font sizing, read-aloud features, and multilingual adaptation. The goal of mixed media is not to “cobble together endings cheaply,” but to present each kind of content in the form best suited to understanding it.
How state combinations produce four endings
Combine evidence_action with ally_trust from the previous chapter, instead of letting the final button alone determine every outcome. Making everything public with high ally trust can lead to joint testimony; making everything public with low trust may lead the other person to deny being the source; sharing only with the ally leads to a private deal; deleting the log leads to an ending with missing information. The four outcomes come from traceable states, rather than four random epilogues.
The design sheet should list the prerequisites, triggering shots, required new media, and reusable parts for each ending. The shared venue video establishes the space, performance, and crisis; phone close-ups show how the evidence is handled; two interchangeable lines of dialogue reflect the relationship; and a final still image or audio segment explains the long-term consequences. This makes it possible to estimate actual assets, rather than simply count ending titles.
Media switches must have a narrative reason
When moving from video to text, first show the character picking up a phone, opening a file, or pausing to remember something; when returning to video, maintain continuity in the character’s position, the object in their hands, and time. If a long explanation suddenly appears in the interface, players will interpret it as a loading screen caused by an insufficient budget. The transition itself should tell users why this is the right moment to read or act.
Text nodes also need staging: information hierarchy, items appearing one by one, audio feedback, and state changes after a choice can all sustain the pacing. However, players must be able to pause and revisit key evidence; it must not be forced to disappear quickly for the sake of a “cinematic feel.”
Mixed-media acceptance checklist
Testers should be able to explain the source of each switch, understand all key information, complete choices at the largest font size and with a screen reader, and trace earlier states back from the ending. Then separately track work hours for video production, text editing, localization, and QA to confirm that complex interfaces and combination testing have not consumed the video budget savings.
When an emotion only works through an actor’s pauses and physical movements, do not replace it with text to save money; when information needs repeated comparison, do not force users to memorize it from video. Cost control should serve understanding and emotion, rather than the other way around.
Start with a prototype without art assets
Use placeholder videos, plain text, and temporary audio to run through the four state combinations and verify that the endings are reachable, the information is readable, and the transitions have a reason before investing in generation or filming. If the four outcomes are already hard to distinguish at the prototype stage, high-quality visuals will only conceal the structural problems, not fix them.
After the assets are complete, test subtitles, read-aloud features, pausing, returning, and saving again. The more media you use, the more important it is to ensure that the same key evidence cannot be lost because of a particular input method or accessibility setting.


