Choose the tier by the stage
Explore on Mini, refine on Fast and keep Standard for the render you will deliver. The same brief carries across all three, so a tier change does not mean rewriting.
BestAIVideo
Seedance 2.0 is ByteDance Seed's multimodal video model, and here it comes in Standard, Fast and Mini tiers so cost can follow the stage of your project. Start from text, anchor a shot with frames, or hand it up to nine images, three video clips and three audio clips, and see the exact credit price before you generate.
ByteDance Seed describes Seedance 2.0 as a unified model for text, image, video and audio input with synchronized sound in the output. These are the settings the generator exposes.
Standard gives the finest detail and is the only tier with 1080p and 4K. Fast iterates quicker, and Mini is the lightweight choice for previews and batch exploration. Fast and Mini stop at 720p.
Pick any whole number of seconds in that range. The default is 5. Audio is generated with the picture, and a switch in the advanced options turns it off when you want a silent clip.
Text to video, frame generation with a required first frame and an optional last frame, or Omni Reference, which accepts up to 9 images, 3 video clips and 3 audio files in one request.
In text-to-video mode you can switch on web search so a prompt about a current topic can draw on fresh information. It is off by default and does not change the price.
ByteDance's guidance for Seedance 2.0 treats every upload as a numbered part of the brief. The habits below keep a crowded request readable.
Refer to Image 1, Video 2 and Audio 1, numbering each type from one in the order you uploaded it. Say what to take from each, such as the outfit from Image 1 or the camera move from Video 1.
Write each spoken line in double quotation marks, say who speaks, and point to an audio reference when you want a particular voice. Describe music by when it enters.
If a clip starts from a photo, set the ratio to adaptive or crop the photo to the target shape. A mismatch is the usual cause of stretching and jumps at the start.
ByteDance recommends staying under about 1,000 English words, because long prompts lose details. A time-coded list such as the first two seconds, then the next three, works well within that budget.
A woman in a red dress pegs white shirts to a line while the low sun flares across rooftops and a wooden basket sits beside her. It is a model of a quiet, single-subject brief where light and ambient sound do the work.
Two fighters, one in white robes and one in a straw cape and conical hat, circle and clash on a muddy floor under mist. Combat like this is easier to control when the prompt lists each exchange in order, with the camera position for each.
An overhead shot of fingers stroking faux fur beside acrylic plates and bubble wrap on a beige table, recorded with close, tactile sound. It shows how Seedance 2.0 can build a clip around texture and audio rather than around a story.
Explore on Mini, refine on Fast and keep Standard for the render you will deliver. The same brief carries across all three, so a tier change does not mean rewriting.
Frame generation and Omni Reference are separate modes. If a shot needs exact first and last frames, use frame generation, and describe the extra references in words.
ByteDance notes that 4K output uses an encoding some players and browsers cannot open directly, so check playback on your target device before you plan a 4K delivery.
Draft at 480p on Mini, where a 5 second clip costs 38 credits, and keep only the ideas worth a better render.
Standard is the tier for delivery, with 408 credits for a 5 second 1080p clip and 832 at 4K.
Feed up to nine images so a person, product or setting stays recognizable across several clips.
Upload a reference clip and describe the movement to borrow, at a lower credit rate than a no-video request.
Add up to three audio files alongside an image or video to steer voice, beat or mood.
Use a first frame and a last frame to decide where the clip starts and where it lands.
Seedance 2.0 is ByteDance Seed's multimodal video model, announced in February 2026. It takes text, images, video and audio as input and produces video with sound generated alongside the picture.
Standard has the most detail and offers 480p, 720p, 1080p and 4K. Fast is quicker and Mini is the lightest, and both stop at 720p. All three share the same modes and settings.
Price scales with length and resolution. For a 5 second clip at 720p it is 164 credits on Standard, 132 on Fast and 82 on Mini. At 480p it is 76, 62 and 38. The generator shows the exact figure before you submit.
Yes. When a request includes a reference video, a lower per-second rate applies to the seconds you generate. A 5 second 720p clip on Standard is 100 credits with a reference video against 164 without one.
Any whole number of seconds from 4 to 15, with 5 as the default. For longer scenes, Seedance 2.5 on this site goes beyond 15 seconds.
In Omni Reference mode, up to 9 images, 3 videos and 3 audio files. At least one file is needed, and images must be under 30 MB, videos under 50 MB and audio under 15 MB.
1:1, 4:3, 3:4, 16:9, 9:16, 21:9 and adaptive, with 16:9 preselected. For image-led clips, adaptive follows your picture.
Yes. Audio is on by default and a switch in the advanced options removes it. The setting does not change the price.
Seedance 2.0 makes clips of 4 to 15 seconds from up to 9 images, 3 videos and 3 audio files. Seedance 2.5 reaches 30 seconds and accepts many more references. On this site, 1080p and 4K are on Seedance 2.0 Standard only.
Put each line in double quotation marks, name the speaker and, if you uploaded a voice sample, say which audio file the voice should follow. Keep lines short enough for the clip length.
In text-to-video mode it lets the model look up current information while it interprets your prompt. It is optional, off by default and unavailable in the frame and reference modes.
Pick a tier, choose how to begin and see the credit price before you generate.