Seedance 2.5
ByteDance's current audiovisual model for coherent clips up to 30 seconds, with native sound, long-form storytelling, rich references, and targeted editing.
ByteDance's current audiovisual model for coherent clips up to 30 seconds, with native sound, long-form storytelling, rich references, and targeted editing.
Plain-English overview
Seedance 2.5 generates picture and sound together, so dialogue, movement, ambience, effects, and background music can follow the same timeline. It accepts text and can also use image, video, and audio references for more controlled production work.
This is a video model rather than a standalone music generator. Its audio is valuable when a finished clip needs synchronized sound, while Eleven Music, Lyria, or Seed-Music are better comparisons when the deliverable is an audio track by itself.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Example: 8-second, 720p clips with audio. Models that do not support these settings will not receive an estimate.
Seedance 2.5
ByteDance Seed
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Available through ByteDance creator products and partner-hosted text-, image-, and reference-to-video endpoints on fal.
ByteDance product availability and fal processing terms vary by account and region. ByteDance announced BytePlus ModelArk API access as coming soon at launch.
Check live availabilityData and training
fal stores request inputs and outputs by default. Its X-Fal-Store-IO: 0 request header prevents payload storage, but uploaded CDN files need separate lifecycle handling; other Seedance routes have their own terms.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
ByteDance's current audiovisual model for coherent clips up to 30 seconds, with native sound, long-form storytelling, rich references, and targeted editing. Seedance 2.5 generates picture and sound together, so dialogue, movement, ambience, effects, and background music can follow the same timeline. It accepts text and can also use image, video, and audio references for more controlled production work.
Seedance 2.5 was released on July 31, 2026 according to the cited provider materials.
Available through ByteDance creator products and partner-hosted text-, image-, and reference-to-video endpoints on fal. The access routes listed in this guide are ByteDance Seed and fal.
$0.0214 / 1K video tokens on fal. For common 16:9 output, fal estimates about $0.2205 per second at 480p or $0.4730 per second at 720p. The authoritative cost uses output dimensions, duration, frame rate, and the endpoint's token formula.
Available through ByteDance creator products and partner-hosted text-, image-, and reference-to-video endpoints on fal. ByteDance product availability and fal processing terms vary by account and region. ByteDance announced BytePlus ModelArk API access as coming soon at launch.
fal stores request inputs and outputs by default. Its X-Fal-Store-IO: 0 request header prevents payload storage, but uploaded CDN files need separate lifecycle handling; other Seedance routes have their own terms. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
OpenAI's legacy synced-audio video model for short text- or image-guided clips—still documented, but approaching a permanent API shutdown.
Open full comparisonGoogle's preview video model for text-, image-, and video-guided generation with native audio and output options up to 4K.
Open full comparisonGoogle's lower-cost Veo 3.1 variant for quickly creating short video clips with native audio and output options up to 4K.
Open full comparisonRunway's flagship developer video model for text- and image-guided clips, with professional formats and a predictable credit-per-second base rate.
Open full comparisonKuaishou's Pro video model delivered through fal for multi-shot text/image generation, optional native audio, references, and clips up to 15 seconds.
Open full comparisonA cost-conscious Kling model on fal for fluid five- or ten-second clips generated from text or a starting image.
Open full comparisonSpaceXAI's current video model for turning text, images, or references into short clips with optional generated speech and sound.
Open full comparisonAlibaba's unified video family for creating short clips with sound from text, a starting image, or a reference video.
Open full comparisonLightricks' quality-focused video model for making 720p or 1080p clips from text, images, or audio with native sound.
Open full comparisonLTX's speed-optimized video model for text, image, or audio input, native sound, automatic duration, and output up to 4K.
Open full comparisonLTX's mature full-workflow video model for generation plus Retake, Extend, Reframe, native audio, and first-to-last-frame control.
Open full comparison