---
title: "Different models, different prompts - save money by generating less"
dek: "How to prompt AI video models: Seedance, Kling, MiniMax H3 and Veo."
description: "Seedance, Kling, MiniMax H3 and Veo each read prompts differently. A short per-model guide to the prompt each one wants - and to paying for fewer retries."
published: "2026-08-16"
updated: "2026-08-16"
author: "Omri Ganor"
tags: ["AI video", "prompting", "Seedance", "Kling", "MiniMax H3", "Veo"]
ogHeadline: "Four models, four prompt formats."
faq:
  - q: "Why does the same prompt give different results on different AI video models?"
    a: "Because each model was trained to listen for a different prompt shape. Seedance expects an ordered checklist, Kling expects a script with named speakers, MiniMax H3 expects labelled reference files, and Veo expects a single shot with the camera named first. The idea travels between models; the format does not, so a prompt written for one model quietly loses half its instructions in another."
  - q: "Which AI video model is best for dialogue?"
    a: "Kling. It is the model built for people talking, and it rewards a script page over a paragraph: give every speaker a fixed name tag, put the action before the line, add a tone to each line, and keep lines short so the lip sync holds. Veo also does speech well inside a single 8-second shot, using plain quote marks rather than tags."
  - q: "How long can an AI-generated video clip be?"
    a: "It depends on the model. Seedance runs 4 to 15 seconds, and up to 30 seconds on 2.5. Kling runs 3 to 15 seconds. MiniMax H3 runs 5 to 15 seconds. Veo 3.1 is fixed at 4, 6 or 8 seconds. A long clip has to be written as steps, with one thing happening in each - a single long sentence looks good for ten seconds and then falls apart."
  - q: "How do I stop wasting money on AI video retries?"
    a: "Find your look on the cheap, low-resolution tier first, then run the winning prompt once at full quality. Write in the format the model expects so it does not ignore half your instructions, and change one thing per retry instead of piling on adjectives. Drafting cheap and finishing once cuts most bills roughly in half."
  - q: "Do AI video models generate sound as well as picture?"
    a: "Yes - Seedance, Kling, MiniMax H3 and Veo all make audio in the same pass as the picture. If you say nothing about it you usually get generic library music, so name the sounds you want, or write \"no music\" if you want it quiet."
---

Here is the fastest way to waste money on AI video: write one prompt, paste it into every model, and hit generate ten times until something looks right.

The problem is that these models don't speak the same language. They were built by different teams, and each one listens for different things. Seedance wants a list. Kling wants a script. MiniMax H3 wants you to label your files. Veo wants one clean shot.

Write the prompt the way the model likes, and you get what you wanted on the first or second try. Write it the wrong way, and the model quietly ignores half of what you said - and you pay for every retry.

This guide is short on purpose. Read the part for the model you use.

## The short version {#short-version}

- **Seedance** reads a prompt as an ordered checklist - who, what they do, where, the look, the camera, the sound - and uses brackets to keep music, sound effects, speech and on-screen text apart.
- **Kling** is built for dialogue. Write it as a script page: a fixed name tag and a tone on every line, and the action before the line.
- **MiniMax H3** takes photos, clips and voice recordings together, but only pays attention if you say what each file is for.
- **Veo** makes one shot of 4, 6 or 8 seconds, and wants the camera named first.
- Whichever model you use, draft on the cheap tier and run only your winner at full quality. That one habit cuts most bills in half.

## The four models at a glance {#at-a-glance}

| Model | Think of your prompt as | Clip length |
|---|---|---|
| **Seedance** (2.0, 2.5, Fast, Mini) | A checklist | 4-15 sec (up to 30 sec on 2.5) |
| **Kling** (2.6, 3.0, O3) | A movie script | 3-15 sec |
| **MiniMax H3** | A brief with labelled files | 5-15 sec |
| **Veo 3.1** | One single shot | 4, 6 or 8 sec |

All four make sound at the same time as the picture. That matters more than you think - more on that below.

## Seedance - write a checklist {#seedance}

Seedance reads a prompt as an ordered checklist, and it likes the information in a fixed order. Put things in this order and it works better:

**Who → what they do → where → the look → the camera → the sound**

You don't need all six. The first two are the only ones that really matter.

> A ceramicist in a linen apron lifts a bowl from the wheel and turns it slowly. A cluttered studio at golden hour, clay dust in the air. Warm film look. Slow push in that settles on her hands. (sparse piano) &lt;wheel slowing, clay scraping&gt;

Those round and pointy brackets are real. Seedance uses four of them to keep sound and text apart:

- `( )` = music
- `< >` = sound effects
- `{ }` = someone speaking out loud
- `【 】` = words on the screen

Speaking and on-screen text are two different things. If you put a line in `{ }`, it gets spoken but it does **not** appear on screen. Want both? Write it twice, once in each.

If you upload photos or clips, Seedance numbers them. You then say what each one is for:

> @Image 1 is her face and hair only. Do not use her clothes or the background.
> @Video 1 is the camera movement only. One slow move, no cuts.

**Do:**

- Say what each photo is for **and** what to ignore in it
- Keep it shorter than 1000 characters, ideally no more than 100-200 words
- For long clips, break the video into steps, with one thing happening in each step
- Test on Mini or Fast (the cheap versions), then run your winner on the full model

**Don't do:**

- Don't upload eight photos and hope it figures them out
- Don't pack in every film word you know - "handheld documentary" and "locked-off symmetrical" fight each other and you get mush
- Don't expect an edit to change the shape of the video. If you need vertical, make it vertical from the start

## Kling - write a script {#kling}

Kling is the actor of the group. It's built for people talking. So write it like a script page, not a paragraph.

> A dim kitchen late at night. Only the fridge hum.
> A plate is set down too hard.
>
> **[Character A: Tired Wife, shaky frustrated voice]:** "You never listen to me."
>
> Immediately, he turns around, eyes wide.
>
> **[Character B: Defensive Husband, shouting]:** "Because you never stop blaming!"

Four small habits do most of the work here:

1. **Give each person a name tag** and use the exact same tag every time. Never write "he" or "she."
2. **Put the action before the line.** Describe the plate slam, then the words. Otherwise Kling doesn't know who did it.
3. **Add a tone to every line.** "shaky," "shouting," "whispering." "He says" is not enough for it to pick a voice.
4. **Use joining words between lines.** "Immediately," or "After a pause," - without them the two voices blur into one.

Kling 3.0 can also do several shots in one go - up to six. Just number them and give each one a job.

**Do:**

- Describe your main character at the very top, and describe them the same way every time
- Say exactly how the camera moves: tracking, panning, holding still, following
- When you start from a photo, treat the photo as the anchor and only describe what changes from it

**Don't do:**

- Don't write long speeches. Short lines stay in sync with the lips; long ones drift out
- Don't use "he" and "she" once two people are in the scene
- Don't write one big paragraph when you want several shots. Number them

## MiniMax H3 - label your files {#minimax-h3}

MiniMax H3's whole trick is that you can hand it photos, clips and voice recordings all at once. But it only works if you say what each file is for.

> Use Image 1 for the mood and the location. Image 2 is the model. Image 3 is the bag.
> Match Video 1 for the cutting speed. Match the voice in Audio 1.

Four photos with four jobs beats four photos and a nice description, every time.

H3 also lets you write a timed list, which stops the video from turning into a slideshow:

> [0-2 sec] Overhead shot. She sits and looks up at the camera.
> [2-4 sec] Push in on her right arm. A menu slides in from the right.
> [10-15 sec] She stands and the whole city loads around her.

**Do:**

- Give every photo, clip and voice file one clear job
- Write out the timing if the clip is longer than one beat
- Say what you **don't** want. H3 listens to this well: "no soft fades," "no black frames," "no spelling mistakes in the text"
- Lock a character by listing their details: "same black hair, silver clip, blue coat, white shoes"

**Don't do:**

- Don't give two photos the same job. If two pictures could both be "the jacket," it's a coin flip
- Don't just say "keep her consistent." List the details instead
- Don't name a transition - describe it. "Whip pan" does less than "blur, smear, then snap back into focus"

## Veo - one clean shot, camera first {#veo}

Veo makes short clips: 4, 6 or 8 seconds. So don't try to fit a story in there. Fit one shot.

And unlike the others, Veo wants the **camera first**:

**Camera → who → what they're doing → where → the mood**

> Medium shot, a tired office worker rubbing his temples, in front of a bulky 1980s computer in a messy office late at night. Lit by harsh ceiling lights and the green glow of the screen. Retro look, slightly grainy.

For sound, Veo uses plain labels instead of brackets:

- Speech goes in quote marks: *A woman says, "We have to leave now."*
- `SFX: thunder cracks in the distance.`
- `Ambient noise: the quiet hum of a spaceship.`

**Do:**

- Start with the shot type: wide shot, close-up, low angle, tracking shot
- Describe what you want instead of what you don't. "An empty desert with no buildings or roads" works much better than "no buildings"
- Make your first and last picture in an image tool, then let Veo fill in the movement between them

**Don't do:**

- Don't ask for a whole story in 8 seconds
- Don't write a long paragraph. Veo's prompt box is small, and short prompts do better anyway
- Don't fix a weak result by adding more adjectives. Change the camera or the lighting instead

## Common mistakes {#common-mistakes}

These come up over and over, on every model.

**1. Piling on pretty words.** "Beautiful cinematic 8K masterpiece, epic." This says nothing, so the model gives you the most boring version of your idea. Describe movement instead: not "a stunning dancer," but "a dancer drops into a low spin, the skirt flares, then she snaps upright."

**2. Reusing the same prompt everywhere.** Kling's `[Character A]:` tags do nothing in Veo. Seedance's `{ }` brackets do nothing in H3. The idea travels. The format does not.

**3. Forgetting that the sound comes free.** All four models make audio in the same pass. If you say nothing about it, you usually get music that sounds like a car advert. Name the sounds you want. And if you want it quiet, write "no music."

**4. Uploading files with no instructions.** A photo carries everything in it - the face, the clothes, the background, the lighting. Say which part you want.

**5. Asking for too much at once.** One change per moment. If you ask for a new pose, a new outfit and a new camera angle in the same beat, all three come out soft.

**6. Treating the maximum as the goal.** Some models accept 50 reference files. Things start falling apart at around eight. Use what you need.

**7. Testing at full quality.** Find your look on the cheap, low-resolution setting first. Then run the good version once. This one change alone cuts most people's bill in half.

**8. Writing a long clip like a short one.** A 30-second clip needs steps. One long sentence looks great for ten seconds and then falls apart.

## How Goa handles prompting for you {#how-goa-helps}

You've just read four different sets of rules. Nobody wants to memorise that - you want to make a video.

That's the part [Goa](/) takes off your plate.

You describe the shot the way you'd describe it to a person: who's in it, what happens, how it should feel. Goa turns that into the right shape for whichever model is making it - brackets for Seedance, character tags for Kling, labelled files for H3, camera-first for Veo. Same idea, four translations, and you never have to learn the difference.

It also handles the boring money part. Goa drafts on the cheap, fast versions of these models so you can look at options, and only spends real money on the version you actually pick. That's the whole point of the title: you save money by generating less, not by picking a cheaper model. What a render costs, and what your plan buys, is [spelled out in plain dollars](/credits.html) - and if you'd rather have your own AI agent drive the studio, [it can](/agents.html).

And when a new model shows up next month - because one always does - your way of describing a shot doesn't change. Ours does.

---

*This space moves fast. Everything here is accurate as of August 2026.*
