Back to blog

How to Write AI Audio Prompts: A Director's Guide With Real Examples

Six techniques for writing audio scene prompts — emotion, sound-effect timing, character casting, paralanguage, music direction — each with a real generated clip you can hear.

Jul 23, 2026SeedAudioSeedAudio
How to Write AI Audio Prompts: A Director's Guide With Real Examples

Two prompts can describe the same scene and produce completely different audio. "A man says welcome" gets you a flat line read. "A warm male voice, low and unhurried, says: 'Welcome back' — over a soft piano, with a door clicking shut behind him" gets you a scene.

The difference isn't the model. It's the prompt. This guide breaks down how to write audio prompts like a director hands off a script — and unlike every other guide on this topic, you can listen to what each technique actually produces. Every example below is a real clip generated from the exact prompt shown.

An annotated script beside a vintage microphone, sound rising off the page

What makes an audio prompt different from a text prompt?

An audio prompt is stage direction, not a question. Text prompts ask a model to think; audio prompts tell a model how a scene sounds — who speaks, in what voice and mood, what music sits underneath, and which sounds land where. A model like Seed-Audio 1.0 generates all of those together in one pass, so everything you don't specify gets improvised. The craft of audio prompting is deciding what not to leave to chance: voice casting, emotional delivery, effect timing, and the music's role. Get those four right and a one-sentence description comes back as a finished, mixed scene.

The formula

Almost every good audio prompt reduces to one pattern:

[Music / mood] + [Character (voice, emotion)] says: "line" + [Sound effects, placed where they happen]

Example: Tense low strings. A man (deep, raspy, calm) says: "You came alone?" A door creaks open behind him.

Everything inside quotes is spoken. Everything outside quotes is direction — the model reads it as instructions for music, voices, effects, and pacing. The six techniques below are ways of getting more out of each part of that formula.

Three floating layers — voice, music, effects — merging into one wave

1. Direct the emotion, don't just name the line

The same sentence can be read a dozen ways. Put the emotional direction in the character parentheses and the model performs it instead of reading it.

Prompt used: A woman (adult, clear voice) says calmly: "The results are in." Then the same woman, now furious, shouts: "The results are IN!"

One voice, one line, two completely different scenes. Words like calm, furious, hesitant, teasing, exhausted are cheap to write and change everything.

2. Place sound effects where they happen, not in a list

Timing follows your text order. An effect written before a line plays before it; written after, it lands after. Don't stack effects at the end of the prompt like a shopping list — weave them into the moment.

Prompt used: A wooden door creaks open slowly. A man (middle-aged, wary) whispers: "Who's there?" Then two heavy footsteps approach on a wooden floor.

Creak → whisper → footsteps, exactly in written order. This is the single most common fix for prompts that "have the right sounds in the wrong places." For more on effect-first scenes, see the AI sound effect generator.

3. Cast characters like a director

Give each speaker a casting note — age, texture, energy — and they stay distinct and consistent for the whole clip. Vague casting ("a man", "a woman") produces interchangeable voices.

Prompt used: A boy (young, bright) asks excitedly: "Did you see that?!" An old man (deep, slow, amused) replies: "I did. And I still don't believe it." Quiet park ambience with distant birds.

Two voices, one generation, no stitching. This is the core trick behind the AI dialogue generator and the two-person dialogue generator — name the characters, and the model holds them apart.

Two voice waveforms conversing across a microphone

4. Ask for paralanguage — the sounds between words

Laughs, sighs, pauses, breath. These are what make a voice sound recorded instead of synthesized, and you can request them directly in the prose.

Prompt used: A young woman starts to speak, laughs mid-sentence, then continues: "Okay, okay — I'll tell you the story." She sighs softly at the end.

Write the human noises into the scene — she pauses, he exhales, they both laugh — and the read stops sounding like text-to-speech. This matters most for podcast-style conversation; the AI podcast generator leans on exactly this.

5. Direct the music like a soundtrack, not a background

Music can enter, duck under a voice, and swell back — if you tell it to. Treat the score as a character with its own entrances and exits.

Prompt used: Upbeat jazz piano starts alone for two seconds, then ducks under a male narrator (warm, smooth) who says: "Some mornings just sound better." The jazz swells back up to finish.

Verbs do the mixing for you: starts alone, ducks under, swells, fades out on. This is how ad reads and narration get that produced feel without a session in an editor.

6. Use distance and space

Voices exist in a room. Whispered close to the mic, called from across a hall, processed through a phone — spatial direction is part of the prompt vocabulary.

Prompt used: A man whispers very close to the microphone: "Can you hear me?" Then the same man calls out from far away, with room echo: "How about now?"

Phrases like close to the mic, from another room, over the phone, echoing in a hall change the perceived space without any post-processing.

Three mistakes that flatten a scene

Everything at once. A prompt that asks for four characters, three effects, and two music changes in one breath usually comes back muddy. One scene, one focus — generate long pieces in parts.

Naming a feeling without staging it. "A sad scene" is weaker than "she speaks slowly, voice catching, over a single held piano note." Stage the sadness; don't label it.

Leaving the music unassigned. If you mention music but never say what it does, it sits at one volume under everything. Give it entrances and exits (see technique 5).

A full worked example

Here's the pattern at production scale — the prompt behind our crime-thriller template: a phone buzz, birdsong ambience, a marimba-led suspense score, two cast voices (one processed as a phone line), and closing foley of a car and footsteps, all in written order. Read the full prompt and hear the result, or start from any of the twelve templates and swap in your own scene.

FAQ

How do I write a prompt for AI music? Describe instrumentation, mood, and movement rather than genre alone: "marimba lead over a synth pad, strings building, tense" beats "suspense music." If the music accompanies a voice, say how they interact — under, behind, swelling between lines.

How long should an audio prompt be? As long as the scene needs and no longer. A single narrated line needs one sentence; a two-character exchange with effects typically runs 60–120 words. Past ~600 words of prompt you're usually describing two scenes — split them.

Do these techniques work in other languages? Yes. The formula is language-independent; write the direction and the lines in the language you want spoken. Chinese scenes — including regional accents — are a particular strength of Seed-Audio 1.0.

Where can I try these prompts? In the studio — your first generation is free with no sign-up. Every example clip in this guide was generated there from the exact prompt shown.

Write one scene, hear it back

The fastest way to learn audio prompting is to generate one scene and listen. Open the studio, paste any prompt from this guide, and change one thing — the emotion, an effect's position, the music's entrance. The difference will teach you more than any article.