Why iterating with a cheaper model costs you more
Cost per attempt is not cost per result - and the second number has the wider range.
There is a version of this everyone has lived. You pick the cheap model, because it costs a fraction of the good one and you're only experimenting. Forty generations later you still don't have the shot, you've spent an hour watching progress bars, and you've spent more than if you'd used the good model twice.
This post is the arithmetic behind that hour.
It is not an argument for always buying the expensive tier. Sometimes the cheap model genuinely is the right call, and there's a section below on exactly when. It's an argument for measuring the right thing - because the number on the pricing page is not the number you pay.
The short version
- What you pay is not the price of a clip. It's the price of a clip times the number of tries it takes to get one you'd keep.
- Across documented productions that second number is about 3 tries per usable shot, and it blows out to 8 or 10 when characters and locations start drifting.
- So the test is a ratio: a model that's 3× cheaper has to need fewer than 3× the attempts. Often it does. Sometimes it very much doesn't.
- The cheap route loses in four specific places: paying twice, retrying something the model can't do, re-rolling instead of revising, and your own time.
- The single highest-leverage habit: draft down a tier, not down a model.
Cost per keeper, not cost per clip
One line of arithmetic sits under all of this:
what you pay = price per attempt × attempts to a keeper
Everyone shops on the first term. It's printed on the pricing page, it's easy to compare, and it feels like the thing you control. The second term is invisible until the end of the month.
Here's the problem: they're the same size of lever.
Take the price term first. From the most efficient tier to the most expensive, what a clip costs spans more than 20×. That's a wide lever, and it's the one everybody pulls.
Now the attempts term. On one documented 3-minute animated production, 164 clips were generated and 41 made the cut - a 25% keep rate, about three generations for every shot that stayed in. And the keep rate is not fixed. It tracks your standard, not your budget: roughly 60-85% for social-tier work, 40-60% for narrative, and as little as 4-5% for broadcast and commercial.
That's 1.2 attempts per keeper at the top of the range and 20 at the bottom. A span of roughly 17×.
Two levers, near enough the same size, and only one of them is printed on the page you're shopping on.
The ten-second test
Forget dollars for a second and count in multiples of the cheap clip. Say the standard tier costs three times what the efficient one does:
| Efficient model | Standard model | |
|---|---|---|
| Price per attempt | 1× | 3× |
| Tries before you keep one | 7 | 2 |
| Cost per keeper | 7× | 6× |
The cheap model is a third of the price and still costs more, because it needed three and a half times the tries. Put your own numbers in - the shape is the point, and the rule that falls out of it is worth memorising:
A model that's 3× cheaper stops being cheaper the moment it needs more than 3× the tries.
Flip one cell and the answer flips with it. Six tries is the break-even here; if the efficient tier lands in four, it's 4× against 6× and it wins comfortably. Which is why "cheap models are false economy" is bad advice, and so is "always draft cheap." Neither is a rule. The ratio is the rule, and you can run it in your head.
The rest of this post is about the four things that push the attempts term up without you noticing.
The four leaks
1. You pay for the draft and the finish
This is the expensive one, and it's the reason the title of this post is worded the way it is.
You iterate on the cheap model. You find a version you like. You re-render it on the good model - and the good model makes different choices. Different framing, different pace, different hands. So you iterate again, at full price.
Now you've paid for both rounds. In the multiples above, that's seven cheap tries plus two good ones - 13×, against 6× for going straight to the model you were always going to ship. Drafting didn't save you anything. It cost you double.
The fix is a distinction worth internalising: a draft is only a draft if it's the same model.
Dropping the resolution, or using the fast variant of the model you intend to finish on, preserves the things you're actually deciding between - the composition, the blocking, the timing, where the eye goes. What you give up is fine texture and colour fidelity, and you weren't choosing between those anyway.
Switching to a different, weaker model preserves nothing. Different physics, different prompt habits, different failure modes. Its output doesn't predict the good model's output, so you're not previewing. You're guessing, one cheap guess at a time, and then paying full price to find out.
Two footnotes, because both cut against the easy version of this advice:
- Promote the draft, don't re-type it. Where a tool offers a documented path from the draft you picked to a full-quality render of that draft, take it. Re-typing the same prompt one tier up doesn't promote anything - it generates a new video that happens to share a description.
- If you only want one version, don't draft at all. A draft plus a finish costs more than a finish. The ladder starts paying from the second variant onward, which is most real work, but not all of it.
2. You keep retrying something the model can't do
Some misses are didn't. Others are can't.
Didn't is fixable with words. It put the cup in the wrong hand, it ignored the time of day, it wandered off your reference. Say it more precisely and the next attempt lands.
Can't is a wall. Legible words on screen, a hand doing something specific with an object, contact physics, one face holding across three shots, lip sync past a few seconds. Cheaper models hit these walls earlier and harder. No amount of rephrasing crosses a wall - but because every attempt comes back slightly different, each one feels like it might be the one. That's how people spend twenty generations proving a model can't spell.
The tell is simple:
If three attempts fail in the same way, it's a wall. If they fail in different ways, it's your prompt.
Three misses that all mangle the same word on the same sign is not bad luck. Go around it: a different model, a different shot, or do that bit on the timeline where text is just text.
3. You re-roll instead of revising
Pressing generate on an unchanged prompt is not iteration. It's a second ticket in the same lottery.
And it's worth knowing why it's so hard to stop. A generate button with a random seed is a variable-ratio reward schedule - rewards arriving after an unpredictable number of attempts. That's the schedule slot machines run on, and it produces the most persistent, hardest-to-extinguish behaviour psychologists have ever measured. Most outputs are fine, some are wrong, and every so often one is wonderful. Your gut learns the only lesson available: pull again.
That isn't a character flaw. It's how the mechanism works on everyone. So the defence has to be a rule rather than willpower:
Every attempt changes exactly one thing, and you can say what it was out loud.
If you can't name the change, you're paying to sample the same distribution twice.
The honest exception: a genuine re-roll is fair when the prompt is right and the model is close. Budget two. Then change something.
4. The clock
A clip takes somewhere between a minute and a half and four minutes to come back, depending on the model, the resolution and how busy the queue is - and resolution moves that number more than clip length does.
So seven attempts isn't just a bigger number on the bill. It's also twenty-odd minutes of watching a progress bar, during which your sense of what you were going for quietly drifts.
The clearest measurement of this comes from the software side of AI rather than video, but it transfers exactly. In one published comparison of the same workload run two ways, the cheaper route cut the vendor bill by 54% and raised review and correction time by 133% - landing at $68 per finished outcome against $35. The invoice went down. The cost roughly doubled.
Time never appears on an invoice. That's precisely why it's the term that gets ignored.
When the cheap model is the right call
Plenty of the time. This is not a "buy premium" argument, and pretending otherwise would be selling you something.
When you don't know what you want yet. Early on, every output teaches you something, so a wrong one isn't a failure - it's information that cost you almost nothing. That's the phase efficient tiers exist for, and it's the phase most people skip.
When the cheap model is genuinely good at your specific thing. Price is a poor proxy for prompt adherence. Some inexpensive models follow instructions better than models several times their price, and the ranking changes every few months. So pick on the difficulty in your shot - dialogue, on-screen text, fast motion, holding a face - and then buy the cheapest tier of the model that handles it. Capability first, price second. That order is the whole trick.
When it isn't the hero shot. Background plates, b-roll, anything out of focus, anything on screen for under a second. Nobody has ever noticed the resolution of a two-frame cutaway.
When your bar is deliberately low. This is the underrated one. Keep rates run 60-85% on social-tier work and 4-5% on broadcast. The higher your standard, the more the attempts term dominates, and the earlier the better model wins. Cheap models are cheapest for people who are easy to please - which is a perfectly good place to be, as long as you know which one you are today.
Nine habits that cut the bill
- Decide what a keeper is before you generate. One sentence: what does this shot have to do? Most retries aren't failures, they're drift - you never named the target, so nothing ever hit it.
- Change one thing per attempt, and name it. Out loud or in a note. Unnameable changes are re-rolls wearing a costume.
- Three strikes, change the plan. Three misses on one prompt means the prompt is wrong, the shot is wrong, or the model can't. All three are fixed by changing something structural, never by a fourth roll.
- Spend upstream, where it's cheap. An image costs roughly a tenth of what a clip costs. Lock the frame first - the face, the outfit, the room, the light - and animate from that. Productions that lock their characters before shooting hold around three attempts per shot; the ones that don't drift to eight or ten. The cheapest thing on the menu prevents the most expensive thing on the menu.
- Draft down a tier, not down a model. Lower resolution, or the fast variant of the model you'll ship. Not a different model.
- Promote the draft, don't re-type it. If there's a path from the version you picked to a full-quality render of that same version, take it.
- Ask whether it's "didn't" or "can't." The same failure three times is a wall. Walls are walked around, not argued with.
- Don't generate what you can compose. Timeline work costs nothing. A still that slowly pushes in, a caption that types on, a cut that lands on the beat - these fill time for the price of an image. A 30-second video does not need 30 seconds of generated video, and the good ones usually don't have it.
- Count keepers, not generations. For one week, log two numbers per shot: how many attempts, and whether you kept it. Your own hit rate per model is the only figure that tells you what anything actually costs you. It's usually a surprise, and it's usually not the model you assumed.
How Goa handles this for you
All of the above is a discipline, and disciplines are the first thing to go at 1am when the shot is nearly right.
So Goa carries most of it for you.
It casts the model per shot instead of running everything through one default, so a talking head, a background plate and a hero move each go to the thing that's good at them - and the cheap parts stay cheap, because nothing sends b-roll to a premium tier. It drafts on the fast version of the model it's going to finish on, so what you approve is what you get. It fills time on the timeline, where time is free, rather than generating video you didn't need.
And it prices in plain dollars. One credit is one dollar, so you can actually know how much you're spending and not be surprised you suddenly spent your whole allowance. "Is that worth another try" becomes a question you can answer at the moment you're asking it, rather than one you work out from a balance after it's gone. If you'd rather hand the whole loop to your own AI agent, it can drive the studio too.
None of this is about making generation cheap. It's about making it deliberate. You bring the idea and the taste; Goa does the heavy lifting and shows you the price of each choice before you make it.
Never a lever you pull and pray at.
Because the cheapest generation is still the one you didn't need.
This space moves fast. Everything here is accurate as of August 2026.
