Seedance 2.5 is here. Every model is included on every plan.
Dolly
Sign inSign up

How to make a product video with AI, starting from one photo

7 min read

From a single product photo to a short video ad: which models animate stills well, how references keep your product accurate, and what it all costs.

Most product videos start life as a photo. You already have good stills, what you need is motion: the bottle turning, steam rising off the mug, fabric moving in wind. Modern video models are genuinely good at exactly this, and the image-to-video route is cheaper and far more controllable than asking a model to invent your product from a text description.

The path below goes from one photo to a clip you can run as an ad, with realistic prices at every step.

Start from a still

Text-to-video will happily generate "a sleek water bottle on a mountain ledge." It will not generate your water bottle. The label, the cap shape, the exact blue: models drift on all of it. Handing the model a real photo as the first frame pins the product down, and the model spends its effort on motion instead of invention.

In the video studio, dock your photo as the first frame and describe only what should happen: "slow 180 degree turn on a marble counter, soft studio light." Keep the motion simple. One camera move or one subject action per clip reads as intentional. Three reads as chaos.

Which model to use

For product work the differences that matter are motion quality, clip length and price. A reasonable ladder:

  • Draft cheap. A short clip on a small model (Seedance Mini runs 86 credits for five seconds) tells you quickly whether your motion idea works at all.
  • Finish on a flagship. Seedance 2.5, the current default, is 250 credits for a five second 720p clip and handles product motion and lighting well. Kling and Veo are strong alternatives with different motion character, and trying the same frame on two models is often the fastest way to a keeper.
  • Add real references when accuracy is critical. Some models accept reference images alongside the prompt, which helps when the product appears mid scene rather than as the opening frame.

Model quirks are handled for you in the studio: pickers only offer combinations a model actually supports, and the exact credit price sits on the generate button before you click it. For background on how these models differ more broadly, see the video generation guide.

A realistic budget

Say you draft four motion ideas on the small model and finish two on the flagship. That is four drafts at 86 credits plus two finals at 250, about 844 credits, or roughly eight and a half dollars at face value. Failed or timed-out generations refund on their own, so a provider hiccup does not eat your budget. Full per-model numbers live on the pricing page, and the credit system itself is explained in the costs post.

Finishing touches

A silent clip is half an ad. Two additions carry most of the weight:

  • A voice line. One or two sentences of narration over the clip. Generating this costs a few cents, and choosing the right voice matters more than the script; the voice picking method takes about five minutes.
  • A music bed. A generated track matched to the product's energy, quiet enough to sit under the voice. See the music guide for style prompts that work.

Cut the three pieces together in any editor. For a six second ad slot, one strong motion beat, one line of voice, and a bed that ends clean is the whole recipe.

Common mistakes

  • Overlong clips. Attention dies fast. Two great five second clips beat one meandering fifteen second one, and cost less to iterate.
  • Busy first frames. The model animates what you give it. A cluttered photo produces cluttered motion. Shoot or generate a clean still first, in the image studio if needed.
  • Ignoring the vertical crop. If the ad runs 9:16, generate 9:16. Cropping a 16:9 clip afterwards usually amputates the product.