Generating Videos

Turn your images and references into short video clips using the in-editor assistant, which animates a still frame (image-to-video) or composes motion from several reference images and videos (reference-to-video). Every video generation run through the assistant asks for explicit confirmation first, because video costs scale with length. The chat bar's Video category is a different surface and does not — see Confirmation below.

Animate an image (image-to-video)

The assistant's generate_image_to_video tool animates one image into a clip, with an optional end frame.

To animate an image:

  1. Put the image on your canvas and select it so it becomes a reference for the assistant (it is numbered as @Image1, @Image2, …).
  2. Ask the assistant to animate it and describe the motion you want (e.g. "make @Image1 slowly pan, camera drifting left").
  3. Pick a model when you have a preference (see Models below). For a smooth start→end transition, also select a second image as the end frame — of the models the assistant can run, only Kling Video v3 and Seedance 2 take one.
  4. The assistant replies with a one-click confirmation showing the estimated credit cost. Generation only starts after you confirm (see Confirmation below).
  5. The clip appears in the conversation when it finishes — typically 30–120s.

Pointer: the resulting video lands in your video library and on the canvas, where it can be selected as a reference for reference-to-video or merged with other clips.

Use references (reference-to-video)

The generate_reference_to_video tool composes a clip from up to 9 reference images and up to 3 reference videos plus a prompt (Seedance 2 Reference).

To use references:

  1. Select the images and/or videos you want on the canvas. They become assistant references, numbered @Image1…@Image9 and @Video1…@Video3.
  2. Address them in your prompt by their tokens — e.g. "use the wardrobe from @Image1 and the camera motion from @Video1".
  3. Reference videos guide motion and style; reference images supply identity, wardrobe, environment, and atmosphere.
  4. Confirm the cost preview to start (see Confirmation).

Pointer: references are pulled from your current canvas selection, so anything you've generated or uploaded into the project can feed a reference-to-video clip.

Models, aspect ratio and duration

These are the models the assistant can run. The chat bar's Video category is a different surface with a wider catalogue; this page does not describe it.

Image-to-video models:

  • Seedance 2 — flagship, cinematic motion. Optional end frame, synced audio, 480p / 720p / 1080p.
  • Kling Video v3 — balanced quality/speed. Output ratio follows the start image. Optional end frame. Output is silent.
  • Wan v2.6 — fast, good for previews and quick iteration. 720p or 1080p. Single start frame only. Silent.
  • Veo 3.1 Lite — Google's high-quality image-to-video with native audio and lip-sync when you ask for sound. 720p or 1080p. Single start frame only.
  • Gemini Omni Flash — physically grounded motion, 16:9 or 9:16, fixed 720p. Always produces audio; there is no silent mode.

Reference-to-video models:

  • Seedance 2 Reference — up to 9 reference images and 3 reference videos.
  • Gemini Omni Flash Reference — up to 10 reference images, always with audio.

Choosing settings:

  • Aspect ratio — honored by Seedance 2, Veo 3.1 Lite and both Gemini Omni Flash entries. Kling's and Wan's output ratio instead follows the start image.
  • Duration — 4–15 seconds, which is what the assistant's tools accept. A model with a shorter ceiling snaps a longer ask down to its own maximum: Veo 3.1 Lite tops out at 8 seconds and Gemini Omni Flash at 10, so a 15-second ask on either costs what its ceiling costs, not three times the 5-second price.
  • Resolution — 480p / 720p / 1080p, honored by Seedance 2 (720p by default for reference-to-video), Veo 3.1 Lite and Wan.
  • Audio — off by default everywhere it is a choice: Seedance 2 and Veo 3.1 Lite render silent clips unless you ask for sound, and Kling is always silent. On Veo 3.1 Lite audio raises the rate, so the same clip with sound costs more. Gemini Omni Flash is always audible and cannot be silenced.

Just describe what you want (ratio, length, audio) in your message and the assistant maps it to the chosen model's supported settings.

Confirmation before every generation (cost)

Every video generation run through the assistant requires explicit confirmation. Before any clip is generated the assistant ends its message with a confirmation that renders a one-click Yes button plus an estimated credit cost and a per-call breakdown. The cost estimate accounts for the clip's duration, resolution, and whether audio is on. Nothing is generated until you confirm — you can also just type "yes". Cost is shown in credits.

The chat bar's Video category is a different surface and has no confirmation step: pressing Send there submits the clip straight away, the same way the image categories do. This page describes the assistant.

Cinematic Prompt (Seedance) assistant skill

For purpose-built Seedance 2.0 prompts there is a Cinematic Prompt — Seedance skill. It reads your uploaded reference images, videos, and audio (wardrobe, identity, voice, environment, atmosphere), routes them into Seedance's deep reference stack (up to 9 images / 3 videos / 3 audio per generation), and composes a Seedance-ready prompt. It works in five cinema modes — Narrative, Studio, Action, Performance, Atmospheric — favoring rhythmic prose over photography jargon, one primary camera move per shot, a locked subject anchor across multi-shot sequences, and inline Avoid X. constraints.

Ask for cinematic or Seedance video prompting to use it. For other video models (Veo, Sora, Kling), there is a sibling Cinematic Prompt skill with the same five-mode grammar. See The Agent (Skills Chat).

Video library and playback

Finished clips are saved to your video library and placed on the canvas as video nodes you can play back. Intrinsic width/height are captured client-side (from a generated thumbnail, or back-filled by loading video metadata) so clips size correctly in the grid; the fallback ratio is 16:9.

Pointer: from the library/canvas you can reuse a clip as a reference (reference-to-video) or combine clips (merge, below).

Merge clips into one video

You can stitch several clips end-to-end with the Merge Videos utility.

To merge:

  1. On the canvas, select 2 to 5 video nodes (no images mixed into the selection).
  2. In the multi-selection side toolbar, click the Merge button (tooltip "Merge N videos").
  3. In the Merge Videos dialog, reorder the clips (drag, or the up/down arrows) into the sequence you want and pick an output resolution: Auto (match inputs), Landscape 16:9, Portrait 9:16, Square HD, Landscape 4:3, or Portrait 3:4.
  4. Run the merge — the combined video is added back to your project.