Generate video
Generate an AI video from your composed shot. Pick a mode, pick a model, set duration, optionally enable audio, render. Plus how to choose between the Veo, Kling, Luma, Runway, and ByteDance options.
The core video-generation flow. Pick a video mode from the visualizer's mode dropdown, pick a model, write or adjust the prompt, click Generate Video (or Capture Animation in Video (from scene)). The result lands in the gallery and plays back inside the visualizer with full transport controls – play, pause, and a scrubber. Scrub to check a specific beat (how a cut lands, whether the model held your camera move at the midpoint) without exporting or watching the whole clip through.

Watch

What it does
Reads the 3D scene plus the active shot's camera, assembles the prompt, picks the chosen video model, and generates a clip. Duration is set by the Duration option and is capped by the chosen model. Some models support up to the shot's full length; others cap below it. Animated subjects and animated cameras are honored if the shot has them; static shots produce subtle camera motion baked in by the model.
The model dropdown lists the current working set of video models from multiple providers. See Models for the full list and guidance on when to reach for each.
The four visualizer modes
The Generate menu at the top of the visualizer panel picks what the panel produces. One still mode, three video modes, and the choice is really about what drives the motion.

Image. Still image generation from the scene plus prompt. Covered in Generate image.
Video (from keyframes). You supply the anchor frames and the description; the model invents the motion between them. Listed in the Generate menu as From Keyframes, "using description and references". Set a Start frame alone to let the prompt drive the motion, or a Start and an End frame to pin both ends and have the model interpolate. Covered in First and last frame.
Video (from scene). Your authored animation drives the motion, and the model renders against it rather than inventing its own. Listed as From scene animation. Direct Render is the no-AI option here, capturing the viewport exactly. Covered in Video (from scene).
Video (modify). Starts from an existing clip rather than the scene: select a source video, or upload one, and restyle or extend it. Listed as Modify or edit video, "edit existing or upload video". This is the video-to-video path.
Inside a video mode, two tabs split the inputs:
Keyframes holds the Start frame and End frame slots, and the source video on Video (modify).
Elements holds the reference images you name in the prompt. See Naming reference images in the prompt.
The model picker is scoped to the active mode and grouped by provider, so pick the mode that matches the job first and the dropdown narrows to the models that support it.
How to use it
Set the mode. Open the Generate menu, hover Video, and pick From Keyframes for the default flow, From scene animation to let your authored animation drive it, or Modify or edit video to work from an existing clip.
Pick a video model. Different strengths per model; see the table below for a starting heuristic.
Populate the inputs the mode requires. On the Keyframes tab: a Start frame for Video (from keyframes), optionally an End frame too, or a source video for Video (modify). Video (from scene) needs nothing extra, because the scene's own animation is the input.
Set the three video options: Duration (e.g. 4 seconds), Orientation (Landscape or Portrait), and Audio (No audio, or Audio enabled). The available options depend on the chosen model.
Read the auto-prompt. Same Scene / Style / Lighting tabs as image generation. Confirm the prompt reads right before committing. For audio-capable models, include the sounds and dialog you want in the prompt (see below).
Click Generate Video (or Capture Animation in Video (from scene)). Generation takes longer than image – often several minutes for complex shots. The job runs in the background; you can keep working in other parts of the project. A toast notifies when the result lands.

Choosing a video model
Quick heuristics. See Models for the full table.
Action-heavy with fast camera moves
Kling 2.6 Pro or Kling 3 Pro
A long, slow, atmospheric move
Veo 3.1
Faster generation at lower fidelity
Veo 3.1 Fast or Kling 2.5 Turbo
Honoring detailed 3D structure (like animated meshes)
Luma Ray 3.14
First-and-last frame interpolation
Veo 3.1, Kling 2.6 Pro or later, Seedance 2.0, or Omni Flash
A shot that has to follow specific reference imagery
Omni Flash, Kling o3 Pro, or Seedance 2.0
4K output
Kling 3 (4K) or Luma Ray 3.14
Quick experimentation across many variants
Kling 2.5 Turbo
Brand or product hero with strict consistency
Match the image model that worked, then pick its closest video sibling
Authoring audio with the prompt
A subset of the video models generate audio along with the video. Veo 3.1 supports it natively. Kling 2.6 Pro and later expose an optional audio track. Other models are video-only; the Audio toggle on the panel is disabled when an audio-incapable model is selected.
When you enable audio on a supporting model, the prompt drives both the video and the audio. Describe the sounds you want in the same prompt:
Dialog. Write the lines a character speaks. Models that generate dialog will lip-sync to it within their tolerance.
Sound effects. Name the SFX you want and where they happen in the scene, prose-style: "the squeak of shopping carts, low conversation, a distant intercom announcement."
Music or ambience. Describe the mood or genre: "low cinematic strings under the dialog", "wind moving through dry grass."
There's no separate audio prompt field. Whatever you write in the main prompt is what the model uses for both.
Naming reference images in the prompt
A video prompt can carry reference images from your media library and address them by name. Attach the images, then write them into the prompt as @image1, @image2, and so on, so the model knows which reference is which rather than blending them all into one impression.

Open the reference-image picker from the Elements tab. The dialog names the slot it is filling, so the header reads Select or upload one image for
@image1, and it offers your own My Media alongside Team Media.Toggle the images you want. The selection is staged: toggling only edits a local draft, and nothing joins the generation until you press Confirm. Cancelling or closing the modal discards the draft.
Read the strip. The attached references appear as a horizontal strip on the Media tab, each tile labelled with its token.
Write the tokens into the prompt. "The livery on
@image1applied to the aircraft, lit like@image2."
How many references a model accepts varies:
Omni Flash
Up to 8
Kling o3 Pro
Up to 8
Seedance 2.0
Up to 7
If an image is deleted from the media library after you attached it, its slot stays in the strip as an empty placeholder so the remaining labels keep their numbers. Renumbering mid-prompt would silently point your prompt at the wrong reference.

Cost note: audio adds materially to the per-clip cost on the models that support it. Run audio-off iterations to lock the visual, then add audio on the final pass.
Animation feeds the model
If your shot has authored camera animation (keyframes) or animated subjects, the visualizer feeds the per-frame state to the video model. The result tracks the motion you authored. Static shots produce subtle motion baked in by the model itself; if you want specific camera action, author it in Compose mode first.
Use Video (from scene) when your authored keyframes need to drive the motion. Use Direct Render when you want a wireframe preview of the scene with no AI involved.
When the video is wrong
The motion isn't what you authored. The video model is reinterpreting your animation. Try a different model: Luma Ray 3.14 honors authored 3D motion most strictly; some Kling variants take more liberties.
The subject drifts across frames. Image reference on the hero object helps, but video models honor references less tightly than image models. Multi-view references and consistent object names help most.
The render is too long or too short. Adjust the shot's duration in Shot details. The video duration matches the shot.
People come back frozen. Characters you don't describe as moving stay frozen in place. Prompt the motion explicitly – "people walking through the lobby", not just "people in the lobby".
Limits and known issues
Video-to-video isn't supported. A rendered video can't be fed back into the visualizer as a motion reference.
Scrubbing is for review. Pausing mid-clip and hitting Generate doesn't generate from that frame – the scrubber position has no effect on what gets generated next.
Audio doesn't sync to specific moments. The audio track is generated holistically; you can't specify "audio cue at frame 30". For precise sync, generate without audio and add the audio in your editing tool.
Orientation options depend on the model. Most video models default to landscape; Veo exposes landscape and portrait. Other models may offer only one. The Orientation dropdown reflects what the active model supports.
Related
First and last frame – the Video (from keyframes) mode in detail.
Video (from scene) – use your authored keyframes to drive the motion.
Direct Render – render the scene as wireframe video with no AI.
Last updated
Was this helpful?

