Blog

Seedance 2.5 vs Seedance 2.0 - is it actually better?

Twice the take, four times the references, 2.6× the price, and still 720p. Where the new one wins, and where the old one still does.

Omri Ganor11 min read

ByteDance shipped Seedance 2.5 on the last day of July. The version number went up by half a point. The price per second, on most of the routes you can actually buy it through, went up by two and a half to three times.

So the question in the title isn't rhetorical. A new model is not automatically a better model, and a better model is not automatically the better choice for the shot in front of you. 2.5 is genuinely better at some things, no better at others, and it has quietly given up one thing 2.0 still has.

This post is what a month of hands-on tests, the official spec sheets and our own price list add up to. It's not a "just upgrade" post, and it's not a "the old one is fine" post. It's a map of where each one wins, so you stop paying 2.5 prices for 2.0 shots.

The short version

  • What changed: one take grows from 15 seconds to 30, the reference cast grows from 12 files to 50, and you can redraw a region of a finished clip instead of re-rolling the whole thing. ByteDance also claims 20% better prompt adherence - its number, nobody else's yet.
  • What it costs: about 2.6× per second on Goa at 720p. ByteDance's own cloud prices it nearer 1.5×; resellers run up to 3×.
  • The catch: the 4K from the June keynote hasn't reached the API. Today 2.5 renders at 720p (1080p on a few routes). 2.0 still renders 1080p and 4K.
  • The tests: 2.5 wins on acting, dialogue, holding a face across a long take, and cinematic camera moves. 2.0 wins or ties on product and mechanical motion, fast action, and anything that needs resolution. Neither has fixed two people touching.
  • The verdict: a better model, not a better default. Pay for 2.5 when the shot needs 30 seconds or a big cast of references. Otherwise 2.0, at whatever resolution you actually need.

What actually changed

Seedance 2.0 Seedance 2.5
Launched February 2026 31 July 2026 (announced 23 June)
One take 4-15 sec 4-30 sec, extendable; a 3-minute mode in beta
References 12 (9 images, 3 videos, 3 audio) 50 (30 images, 10 videos, 10 audio)
Resolution you can buy today 480p up to 4K 480p, 720p; 1080p on some routes
Editing Re-roll the clip Redraw a region or a timestamp, keep the rest
Audio Generated with the picture Same, tighter sync; lip-sync in 10+ languages
Technical report Model card published April 2026 None yet
Independent arena score Top five text-to-video, second image-to-video Not listed
On Goa, 720p $0.23 per second $0.61 per second

Three of those rows are real changes. The rest of this section is about them, and about one number that isn't a change at all.

Thirty seconds in one pass

This is the headline, and it's real. 2.0 stops at 15 seconds, so anything longer is two clips and a join - and a join is where the face changes, the light shifts and the music restarts. 2.5 does 30 in one go, with scene changes and tempo shifts inside the take, and can extend from there.

Two things to know. First, it isn't a 30-second-only model. Four seconds still works, and you should still ask for the length you need. Second, a longer take is a longer bet: a miss at 30 seconds costs you 30 seconds. One reviewer's short film took fifteen 30-second generations to reach 142 seconds of finished cut - a 3.2-to-1 shoot ratio, which is about par for this kind of work. At 2.5 prices, par adds up.

Fifty references

2.0 takes 12 files: nine images, three clips, three audio tracks. 2.5 takes 50: thirty, ten and ten. Blender and Maya assets, a clay-render "white model" pass and green-screen support come along with it.

Whether you need 50 is a different question. The reviewer above held a character across a two-minute film with three references - two portraits and a location. PixVerse's own guide warns that piling on conflicting assets makes the brief less precise, not more. And the rule from the prompting guide still stands: things start falling apart around eight. Check what your route exposes, too - through the API, most routes (ours included) still give 2.5 the same nine, three and three as 2.0.

Edit a region instead of re-rolling

The quiet one, and potentially the biggest saver of the three. On 2.0, one wrong prop means regenerating the whole clip and hoping the other 14 seconds come back the same. On 2.5 you circle the prop, describe the change, and the motion, lighting and audio around it stay as approved. It also takes edits at a timestamp - "at 0:12, she picks up the cup".

For now this lives in Dreamina's app rather than every API. If it's the thing you want, make sure it's on the route you're paying through.

The claimed 20%

ByteDance says 2.5 follows prompts 20% better than 2.0. It hasn't said how that was measured. There's no technical report to check - 2.0 got a model card on arXiv in April, 2.5 has nothing yet - and a month after launch it still isn't on the Artificial Analysis arena, where 2.0 sits in the top five for text-to-video and second for image-to-video.

That's not a red flag. It's just worth being clear that the only independent number in this comparison belongs to the old model.

What the tests say

Nobody has run a proper blind comparison yet. What exists is a month of hands-on tests from people who used both, and they agree more than you'd expect.

Where 2.5 is clearly better

  • Acting and dialogue. The most consistent finding. A two-hander in a bar held its performances across 30-second takes in a way 2.0's 15-second cuts never allowed, and a Spanish version of an ad came out with usable lip-sync in one pass.
  • Holding a face over a long take. In side-by-side tests, 2.5 keeps a character's appearance more stable across the shot. Faces, that is - one tester noted body proportions still wander between shots.
  • Cinematic camera. More depth, better lighting, and camera moves that carry momentum through the take rather than resetting every few seconds.
  • Deliberate hands. A finger-counting test on both models: 2.5 was more temporally consistent and matched the spoken number to the fingers more often. Both still fumbled two-hand coordination.

Where 2.0 wins, or it's a wash

  • Product and mechanical motion. In one side-by-side, 2.0 gave a steadier, more readable product transformation; 2.5 moved abruptly between positions and drifted the framing. If the shot is a thing turning on a plinth, the old model is the better model.
  • Resolution. 2.5 is 720p today. Close-ups on dialogue come out soft, and tight push-ins are unflattering. 2.0 renders the same shot at 1080p or 4K - and on Goa, 2.0 at 1080p ($0.59 a second) costs less than 2.5 at 720p.
  • Fast action. Morphing and decoherence in quick movement - characters go spiky and unstable - and it can't be upscaled away. 2.0 did this at launch too and was patched later, so this may age well. It hasn't yet.
  • Two people touching. ByteDance itself says complex physical motion and multi-subject interaction still need work. Both models misread gestures - one tester got unrequested kisses from both - and both occasionally garble speech.

Put together: the people who used both call it a meaningful step on performance and an incremental one on everything else. Not the leap 2.0 was is the line that keeps coming up. 2.0 was a ground-up rebuild that put text, pictures, clips and audio into one model. 2.5 refines a model that already worked.

The price, honestly

Prices vary a lot by route. Through ByteDance's own cloud, 2.5 costs roughly half again as much as 2.0. Through resellers, anywhere from 1.5× to 3×, depending on who you ask and when. On Goa, at 720p, it's $0.61 a second against $0.23 - call it 2.6× - and at 480p it's $0.29 against $0.10.

Shot Seedance 2.0 Seedance 2.5
10 sec, 720p $2.34 $6.15
15 sec, 1080p $8.78 not available
30 sec, 720p two 15-sec clips, $7.02, plus the join $18.45
30 sec, 480p draft - $8.60

So a 30-second take on 2.5 costs about two and a half times what two joined 15-second takes on 2.0 cost. Whether that's worth it depends entirely on what the join would have cost you - in re-rolls to match the two halves, and in the cut you couldn't hide.

Which brings back the arithmetic from the last post: what you pay is price per attempt times attempts to a keeper. On dialogue and long takes, 2.5 needs fewer attempts, so the 2.6× may well come back. On a six-second product shot, it needs the same number or more, and you've paid 2.6× for nothing.

Two billing details worth knowing. When you hand 2.5 a reference video, the per-second rate for the output drops (to $0.37 at 720p on Goa) but every second of the reference clip is billed as well - so trim references to the part that matters. And if you let it choose the length, you're quoted for the longest it might run. Give it a number.

Your 2.0 prompts don't carry over

This surprised people. Prompts tuned for 2.0 can come back visibly broken on 2.5 - glitchy, unstable, not what was asked. Three things changed:

  • It wants scenes that breathe. 2.5 resists rapid-cut, dense-beat prompting. Give a 30-second clip four or five beats, not twelve.
  • It wants more structure, and more of it. Prompts run longer and need explicit order. The checklist from the prompting guide still applies - who, what, where, look, camera, sound - it just needs every beat written out for a 30-second take.
  • It casts voices from the picture. If you don't say, 2.5 infers what a character should sound like from the reference, and one tester got an unrequested British accent. "American" fixed it. "American English" didn't, because the word "English" pulled it straight back.

Two smaller ones. The rights screen is touchy - "a 35mm film look" got a clip refused until it was reworded in plain language. And generation time swings, from about a minute to nearly four for identical requests.

The practical rule: don't draft on 2.0 and finish on 2.5. It's the double-pay trap from the last post - a different model preserves nothing, and here it doesn't even preserve the prompt. If you're finishing on 2.5, draft on 2.5 at 480p. Same model, one tier down, under half the price.

When to pay for 2.5

  • The shot is longer than 15 seconds and has to be one take. A conversation. A walk through three rooms. Anything where a cut would be a cheat.
  • A face has to survive that long. Long-take character consistency is the thing it's measurably best at.
  • It's a performance. Dialogue, reactions, timing. The gap over 2.0 is widest here.
  • You'll fix rather than re-roll. If one prop is wrong in a 28-second take you otherwise love, region editing pays for itself the first time.
  • You genuinely need the cast. Several characters, a location, a motion clip and a voice, all in one take.

When 2.0 is still the right call

  • The clip is under 15 seconds. Most short-form is. Most of your shots are.
  • You need 1080p or 4K. 2.5 can't sell you that today.
  • It's a product, a mechanism, or a thing turning. 2.0 tested steadier.
  • It's fast action. Until the morphing is patched, 2.0 is the safer buy.
  • You're iterating on a short shot at all. 2.0 at 480p is $0.10 a second - the cheapest way to see an idea move in this family.

The verdict

Is Seedance 2.5 actually better than 2.0? As a model, yes - and it's better in the places that were hardest: long takes, performance, a face that stays itself for half a minute. As a default, no. It costs 2.6× as much, renders at 720p, needs its prompts re-learned, and on the short mechanical shots most short-form is built from it ties the model it replaced.

The mistake is treating the version number as a ranking. Treat it as a second tool. 2.5 is the long-take, big-cast, performance model. 2.0 is the short, sharp, high-resolution one. Most videos want both, in different places.

Upgrade the shot, not the account.

How Goa handles this for you

You shouldn't need to hold any of this in your head, and with Goa you don't.

Both models are in the studio, and Goa casts per shot rather than per project. The 20-second conversation goes to 2.5. The six-second product hero goes to 2.0 at 1080p, for less than 2.5 would charge at 720p. The draft of either goes to the same model a tier down, so what you approve is what you get. Nothing pays 2.5 prices for a 2.0 shot.

It absorbs the prompt problem too. You describe the scene the way you'd tell a person; Goa writes it in the shape each version wants - beats that breathe for 2.5, the tighter checklist for 2.0, references addressed by name for both. When 2.5's habits change again in a patch, ours change with them and yours don't.

And every choice shows its price in plain dollars before you generate - including the "this 30-second take is $18" moment, while it's still a decision and not a surprise. If you'd rather have your own AI agent drive the studio, it can.

You decide what the shot is. Goa decides which Seedance it needs.


This space moves fast. Everything here is accurate as of August 2026.

Common questions

What is the difference between Seedance 2.5 and Seedance 2.0?
Three things. Seedance 2.5 generates up to 30 seconds in one take where 2.0 stops at 15; it accepts up to 50 reference files (30 images, 10 videos, 10 audio) where 2.0 takes 12; and it can redraw a region or a timestamp of a finished clip instead of regenerating the whole thing. ByteDance also claims 20% better prompt adherence. What 2.5 gives up is resolution - it ships at 720p today, while 2.0 renders 1080p and 4K - and it costs roughly 2.6 times as much per second.
Is Seedance 2.5 actually better than Seedance 2.0?
In specific places, yes. Hands-on tests agree it is better at acting and dialogue, at keeping a face consistent across a long take, and at cinematic camera moves. It is no better, and sometimes worse, at product and mechanical motion, at fast action (where it can morph), and at two people interacting. Reviewers who used both call it a meaningful step on performance and an incremental one elsewhere - not the leap 2.0 was. As a model it is better; as a default for every shot it is not.
How much more does Seedance 2.5 cost than Seedance 2.0?
It depends on the route. Through ByteDance's own cloud, roughly half again as much per second. Through resellers, anywhere from 1.5× to 3×. On Goa at 720p it is $0.61 a second against $0.23, about 2.6×, so a 10-second clip is $6.15 instead of $2.34 and a 30-second take is about $18. Because 2.5 needs fewer retries on dialogue and long takes, the premium can pay for itself there; on short product shots it usually does not.
Does Seedance 2.5 output 4K?
Not on any route you can buy today. ByteDance announced native 4K at the June keynote, but the API rate cards list 480p and 720p, with 1080p on a few platforms, and hands-on reviewers report a 720p cap. Seedance 2.0 still renders 1080p and 4K. If a shot needs resolution more than it needs length, 2.0 is the model for it - and 2.0 at 1080p costs less per second than 2.5 at 720p.
Do Seedance 2.0 prompts work on Seedance 2.5?
Not reliably. Prompts tuned for 2.0 can come back glitchy on 2.5. The new model resists dense, rapid-cut prompting and wants longer, more structured prompts with scenes that breathe - four or five beats in 30 seconds rather than twelve. It also infers a character's voice from the reference picture unless you name the accent. Rewrite rather than paste, and if you plan to finish on 2.5, draft on 2.5 at 480p rather than on 2.0, so the draft actually predicts the result.

Contact: hello@getgoa.io

Let Goa write the prompt.

Describe the shot the way you'd describe it to a person. Goa picks the model and writes it in the shape that model wants.

← All postsHow credits workConnect your agentRead as markdown