Characters, scenes and voices
Asset settings are the foundation of consistency. Approve a main appearance, define variants, locations, props and voices, then reference them correctly in each shot.
Main appearances and variants
- Open project Settings and inspect each extracted character.
- Establish the main appearance through generation, a library selection or an upload.
- Add costume or story-stage variations as variants of the same character, rather than unrelated characters.
- Check names and episode appearances so that shots can reference the right version.
- Compare generation history and select the current image. After changing the main appearance, inspect variants and existing footage for consistency.
Choose clear references: use visible facial features and recognizable clothing. Separate images help distinguish characters that would otherwise share one crowded picture. Explain each reference's purpose in the prompt and use assets you are authorized to use.
Locations and props
A location reference defines space, lighting and environment; a character reference defines identity. Create separate assets for important objects that recur across shots, such as letters, jewelry, weapons or handheld items.
Go beyond “luxurious room”: establish doors, windows, furniture and usable floor space. Switching between unrelated location images from shot to shot undermines continuity.
Configure voices
Depending on the interface, use the voice library, custom audio or intelligent voice generation. Intelligent voices use character information to produce a voice description. Edit that description and audition the result before assigning it.
- Open voice settings for a speaking character.
- Select a voice or provide clear audio containing a single speaker. Follow the upload dialog's format and duration requirements.
- Audition text in the target language. Check perceived age, delivery, speed, emotion and accent.
- Prefer completing voice setup before storyboard analysis.
- If voices are added later, use intelligent linking and inspect the actual voice references in the shots.
The September 2026 update supports intelligent voice linking. It does not replace checking the speaker and voice assignment in each shot. A character with no dialogue generally does not need an audio reference for that shot.
Before batch production
| Check | Ready state |
|---|---|
| Main appearance | Every principal character has an approved current image |
| Variants | Each belongs to the correct character and story moment |
| Locations | Main spaces are complete and support continuity |
| Props | Appearance, position and ownership are clear |
| Voices | Each speaker has the intended voice and an acceptable target-language audition |
| References | Each shot points to the correct character, image and voice assets |
Use asset libraries for cross-project reuse. For original-to-new character mapping and merging in remaster projects, see Replace and correct assets.
Automatic versus manual asset setup
Before entering setup from the script, choose whether to generate assets automatically. Automatic setup creates main character images and scene assets; listed outfits still need to be generated or configured manually. Manual setup suits projects with approved artwork: generate, select from a library or upload each asset before continuing to shots.
Replacing a main character, scene or prop image does not erase its derived assets. Those older derivatives may no longer match, so review consistency. Search within setup to locate assets, or use batch matching for larger projects.
Custom voice clips and intelligent voices
Custom character voice uploads can exceed 5 seconds, but the documented behavior takes the first 5 seconds as the voice sample. Put clean speech from one speaker at the beginning, without leading silence, music or overlapping speakers. This is separate from the 30-second total audio-reference limit for video generation.
Intelligent voice design uses Seed Audio 1.0. Review the generated voice description, edit it, audition a sample and select it before shot analysis. Narrated-drama mode uses a shared narrator voice configured in setup. If voices are added after analysis, run intelligent association and check each mention. Characters without dialogue default to no voice in that shot.
