
01
Multi-character short drama
Keep several speakers in one take, then leave rain, footsteps, and a sting on their own tracks so an editor can ride them independently.
Seed Audio 2.0 is built for T2A, TA2A, TV2A, and TAV2A. One brief can hold six minutes, six reference voices, thirty languages, and independent stems for dialogue, music, ambience, and effects. Reserve credits now. The studio below still generates with Seed Audio 1.0, so you can practice prompts before Seed Audio 2.0 goes live.
Scene desk preview
Sample brief

A captain, engineer, and officer argue as alarms climb, engines rumble, and strings tighten before impact.
Seed Audio 2.0 appears in the model menu with a Coming Soon tag. Generation still runs on Seed Audio 1.0, so you can test credits, references, and scene briefs before Seed Audio 2.0 opens. Keep the same account; early-bird orders apply when the new model is switched on.
Model
Input
Prompt-first audio generation with optional controls.
Additional Settings
Customize your input with more control.
History
Your recent Seed Audio 1.0 Preview generations.
Sign in to see your generation history.
How to use
Treat today’s generator as a rehearsal room for Seed Audio 2.0. Write the scene, attach the references you already have, set output limits, then listen. When Seed Audio 2.0 ships, the same briefing habit should transfer: longer takes, more voices, video context, and stems instead of a single bounce.

Name who speaks, where they stand, what the room does, and which cue must land. That is the brief the new model is being trained to follow at longer lengths.

Today the studio accepts fewer beds than the 2.0 spec. Still upload the voices you care about so the prompt mentions them the same way you will after launch.

Set format, speed, and a credit ceiling. Seed Audio 2.0 will ask for longer jobs; practicing cheap drafts now keeps the habit honest.

Check overlap, music masking, and late Foley. Those are the notes stems and timestamps are meant to reduce.
Seed Audio 2.0 brief
Mode + duration + speakers + language + picture or voice refs + stem intent + timestamped cues.
After launch
Seed Audio 2.0 is not a longer TTS clip. It is a scene renderer: picture-aware dubbing, extra reference voices, and stems you can drop on a timeline. The four boards below are the jobs we hear producers asking for while they wait for Seed Audio 2.0.

01
Keep several speakers in one take, then leave rain, footsteps, and a sting on their own tracks so an editor can ride them independently.

02
Feed a picture plus optional voice beds. The 2.0 spec is to follow action, cuts, and pacing instead of laying a flat narration under silent video.

03
Boss lines, forest loops, and UI hits can share one brief, then split so a game mixer does not have to unbake a stereo bounce.

04
A three-voice car spot can keep VO, engine, city wash, and the hook on separate stems for localization and legal recuts.
Where it lands
These are production jobs, not mood boards. Seed Audio 2.0 is being positioned for people who already cut picture, localize spots, or ship interactive worlds and need the soundtrack to arrive as editable parts. If you only need a single narrator, Seed Audio 1.0 or a speech engine may still be simpler until Seed Audio 2.0 is live.
Aimed at multi-scene shorts where voices, beds, and hits have to stay in character across cuts instead of being restitched from stock.
Use it when a spot needs VO, product Foley, and a music lift in one pass, then separate stems for market versions.
Barks, loops, and stingers can be briefed together. Stems give audio leads something closer to a session than a trailer bounce.
Thirty-language coverage plus video context is the Seed Audio 2.0 pitch for teams that currently re-record every market by hand.
Six minutes is still a chapter, not a series, but it is enough for a podcast cold open or a comic-drama act that Seed Audio 1.0 could not hold in one job.
Six reference voices help keep a cast distinct instead of collapsing everyone into two stock timbres.
Model notes
The list below is the published Seed Audio 2.0 delta versus the 1.0 studio you can run today: more time, more references, picture input, stems, tighter cues, and a wider language set. It is a production spec, not a slogan sheet.
Specified for T2A, TA2A, TV2A, and TAV2A, so a job can start from text, voice beds, picture, or all three.
The bed count rises from about three to six, which matters when a scene has a lead, a foil, and a crowd texture.
TV2A and TAV2A read action and pacing instead of guessing a score under a silent timeline.
Dialogue, music, ambience, and effects can leave Seed Audio 2.0 as separate tracks instead of one glued mix.
Seed Audio 2.0 is described with precise cue placement so a door slam or a line pickup can sit on a marked beat.
Wider language coverage for dubbed spots and localized drama without resetting the entire mix.
Version delta
Seed Audio 1.0 is what you can generate here today. Seed Audio 2.0 is the upcoming Dreamina-class model: longer jobs, more references, video input, stems, and a wider language set. Use the table as a briefing sheet, not a promise that every 2.0 control is already in this studio.
| Dimension | Seed Audio 1.0 (July 2026) | Seed Audio 2.0 (current Dreamina spec) |
|---|---|---|
| Maximum duration | About 2 minutes | Up to 6 minutes |
| Reference audios | Up to about 3 clips | Up to 6 clips |
| Input methods | Text + reference audio (+ optional image) | Text + reference audio + reference video (T2A / TA2A / TV2A / TAV2A) |
| Language support | 20+ languages | 30 languages |
| Multi-track and time control | Basic timeline control (dialogue precision about 100ms) | Independent stems for dialogue, music, ambience, and effects, plus precise timestamps |
| Core capabilities | End-to-end scene generation (dialogue + SFX + ambience + music) | Stronger controllable cloning, emotion/rhythm/style control, video-aware dubbing, and cross-scene voice consistency on top of the 1.0 scene model |
| Best suited for | Short scenes and basic sound sketching | Longer audio, short dramas, ads, games, multilingual localization, and crowded multi-character scenes |
Seed Audio 2.0 is not a free upgrade inside this generator yet. Order early-bird credits if you want capacity waiting when the model is enabled; keep using Seed Audio 1.0 for drafts.
Lock early-bird pricingPricing
Subscribe for the best value, or buy credits when you need a flexible top-up. Every paid plan and credit pack includes Seed Audio 1.0 API access, with one shared credit balance across the web app and API.
Try Seed Audio 1.0 with free signup credits. Perfect for testing prompts and short scenes.
The best annual choice for creators with ongoing audio needs.
The strongest value for teams that expect heavy audio generation.
Field notes
These are composite notes from trailer editors, localization producers, and game audio leads — written to sound like working notes, not a five-star wall. Nobody here claims Seed Audio 2.0 has shipped on this site.
I bought credits because six-minute stems would save a weekend of temp dubs. Until Seed Audio 2.0 is on, I still run 1.0 just to see whether the prompt is even sane.
The video-to-audio pitch is the only reason I care. If the new model can follow a cut instead of ignoring it, we stop laying English VO under picture that already has mouth flaps.
I do not need another magic voice. I need Seed Audio 2.0 to keep the heroine on one bed across three episodes. If cloning still drifts, we will keep recording.
Six references is the upgrade I can explain to a producer. Three was always a squeeze. I will believe the stems when I can solo the rain without killing the line.
Early-bird
Seed Audio 2.0 is listed as coming soon. Orders placed now keep early-bird pricing. You can still generate with Seed Audio 1.0 on this page while the new model is queued.