Step 5: Audio
Configure voice (style, pacing, language, accent) and background music for your video.
What is the Audio step?
The Audio step collects your voice and music preferences and passes them as hints to the selected video model. Toggle voice on or off, choose a language, accent (for English), delivery style, and pacing. Then describe background music and balance the music volume against the voice. These are prompt inputs — whether the final clip contains spoken voice or music depends on the model you're using. This step is optional.
Enable Voice
Toggle spoken audio on or off. Off creates a silent or music-only video.
Voice Source
Three options: 'Character Speaks' uses a voice matching your selected character (this is the default), 'Narrator' uses a voiceover with no tie to a character, and 'Both' combines a character voice with a narrator voice.
Language
The spoken language: English, Spanish, French, German, Italian, Portuguese, Hindi, Japanese, Mandarin Chinese, Korean, Arabic, or Russian.
Accent
For English, choose an accent variant (e.g., American). The accent option is hidden for other languages.
Delivery Style
The tone and energy: Conversational, Energetic, Calm, Authoritative, or Whispery.
Pacing
Speaking speed: Slow, Normal, or Fast.
Music
A background-music description passed to the video model. Describe the music you want in the prompt box, or switch it off. The model may add a matching track.
Volume Mix
Balance between background music and speech. Default is 30% music — optimized for speech clarity.
Configuring Audio
Set the voice
Leave Enable Voice on (default) for spoken audio tied to your character, or toggle it off for a music-only video.
Choose language and accent
Pick the spoken language from 12 options. For English you can also pick an accent.
Pick delivery style and pacing
Choose the tone (Conversational, Energetic, Calm, Authoritative, Whispery) and speed (Slow, Normal, Fast).
Describe your music
In the Add Music & Mix section, describe the background music you want (e.g., 'upbeat acoustic with light percussion') to pass to the model, or turn music off.
Balance the mix
Use the volume sliders to mix music against speech. The 30% default works well; nudge louder for social content, quieter for educational content.
Voice and music are handled by the selected video model as part of generation — ScriptMotion passes your settings as prompt hints. There's no separate ScriptMotion audio pipeline, so whether audio appears depends on the model's capabilities.
Leaving the music prompt empty is fine — the model decides the audio, and your volume setting is applied if it produces a track.