Seedance 2.5 is here. Every model is included on every plan.
Dolly
Sign inSign up

AI video generation: turning a sentence — or a still image — into video

6 min read

How text-to-video and image-to-video models work, what Seedance, Kling and Veo are good at, and how to control motion, duration and cost.

AI video generation turns a text prompt — or a still image — into a short video clip. It is the most impressive and the most demanding kind of AI media: what took a camera, a set and an editor can now start from a sentence. It is also the most expensive to run, so it pays to understand how it works before you spend credits.

Text-to-video vs. image-to-video

There are two ways in. Text-to-video generates the whole clip from a description: "a drone shot flying over a misty pine forest at dawn." You get motion and composition for free, but less control over exactly how the first frame looks. Image-to-video starts from a still you already have — a photo, or an image you just generated — and animates it. This is the more reliable path for brand work, because you lock the look first and only ask the model to add motion. On the video studio you can do either, and animating one of your own generated images is a single click.

The main models, and what they're for

  • Seedance — a strong, well-priced default for general text-to-video and image-to-video, with support for reference images and a range of durations.
  • Kling — known for fluid, cinematic motion and character movement; a good pick when the feel of the motion matters most.
  • Veo — Google's model, strong at prompt adherence and, in some modes, synchronized audio.

As with images, switching models is often the fastest way to fix a disappointing result. Using a platform that carries several of them — ideally from different vendors — also means one provider's outage doesn't stop your work.

The controls that matter

  • Duration — most clips are 5–10 seconds. Cost scales with length, so generate short, then extend what works.
  • Resolution — 720p is the sweet spot for drafts; go higher only once the motion is right.
  • First / last frame — some models let you pin the opening (and closing) frame to an image, which is the tightest control you get over the result.
  • Reference images — feed in a character or product so it stays consistent across shots.

Why video costs what it does

Video is priced per second of output, and a single second is far more compute than a whole image — which is why a few seconds of video can cost more than fifty images. Two habits keep the bill sane: prototype the look in the cheap image studio first, and only animate a still you already like; and generate short clips before committing to longer ones. And confirm you are never charged for failed generations — ours refund automatically. Rates are on the pricing page.

A reliable first result

If your first text-to-video attempt looks generic, don't reword endlessly — instead generate a strong still image of your subject, then animate that image. You will get a far more controllable result. From there, add a voiceover and a music bed to finish the piece. New accounts include 200 free credits to experiment with.