Control shot composition
How to control composition, lens, aspect ratio, and aim in your AI renders. The Compose-mode decisions the model can't override.
Pure prompt tools randomize composition. Ask for "a wide shot of a car on a wet street at night" and you get whatever wide shot the model trained on. Ask for "from a low angle, 35mm, 2.39:1 cinemascope, lead with negative space on the left" and you get back a generic wide shot anyway. Composition is spatial; text is bad at carrying it.
Compose mode is where you make composition non-negotiable. Camera position, lens, aspect, aim target – all spatial decisions, all passed to the model as structured input alongside the prompt. The model can't override what you set; it renders the framing you placed.
Steps
Add a shot. In Compose mode, click the camera + icon directly below the viewport. The new shot snapshots the current viewport framing as its starting point.
Position the camera. Orbit, pan, dolly, and zoom to find the framing. The viewport shows the camera's actual view. See Camera controls.
Set the lens. Open the lens picker and choose the focal length. 14mm ultra-wide for environmental establishment, 24mm or 35mm for documentary feel, 70mm for portraiture, 100mm or 200mm telephoto for compression. See Lenses.
Set the aspect ratio. 16:9 for digital delivery, 2.39:1 for cinemascope, 1:1 for social, 9:16 for vertical. The picker has the standard production aspects pre-loaded. See Aspect ratios and film gates.
Set an aim target if the camera should track an object. Aim Camera binds the camera's look-at point to a 3D object, so when the object moves the camera follows. The target appears as a crosshair anchored to the object. See Aim Camera and Target.
Frame the action. Use the PIP (picture-in-picture) overhead view in the top-left of the viewport to check spatial relationships – where the subject sits relative to the environment. The PIP is on by default in Compose.
Lock with Shot Details. Open Shot details. Name the shot from the brief and write a description. The shot is now ready for render.
Render. In Visualize mode, generate. The composition you set – position, lens, aspect, aim – is passed to the model as part of the scene. The render comes back framed exactly the way you composed it.
Why this works
The Compose mode camera is a real 3D camera, not a description of one. The model receives the camera's position, target, focal length, and aspect ratio as structured parameters – not as English. Lens choice changes perspective and field of view in the rendered output the way it would in a real camera, because the same math applies.
Prompt tools can't do this because they have no 3D camera state. "35mm" in a prompt is a token the model has loosely associated with certain visual qualities; it isn't actually a 35mm lens.
For the deeper version, see How the visualizer thinks.
Composition decisions in production order
Camera position
Compose viewport (orbit / pan / dolly / zoom)
Where the camera is in the scene
Lens / focal length
Compose lens picker
Field of view, perspective compression, depth of field
Aspect ratio
Compose aspect picker
Frame proportions, where to place subjects in frame
Aim target
Aim Camera tool
What the camera tracks if the subject moves
Subject placement
Build mode (move the object)
What's in frame at all
The order matters in production: build the world, then choose the lens, then position the camera, then aim. Reverse the order and you spend more time hunting for the framing.
Tips
Pick the lens before fine-tuning the position. Different focal lengths require different camera distances for the same composition. A 24mm at 5 feet and a 100mm at 20 feet frame the subject the same size, but with very different perspective. Choose the lens, then position.
Use the PIP overhead view for staging. When characters or vehicles need specific blocking relative to each other, the PIP shows the spatial layout that the camera view obscures.
Aim Camera for moving subjects. A racing car shot frames itself if you aim the camera at the car and let the car drive past. The camera tracks the subject across the shot's duration without manual keyframing.
Limits
No DOF / aperture control as a numeric setting. Depth of field is implied by lens choice and the model's interpretation. Telephoto lenses produce more compressed depth in renders than wide lenses, but you don't dial f-stop directly.
No camera shake or handheld simulation. The camera path is smooth unless you keyframe variation. For handheld feel, add it in your edit.
The first-and-last-frame video flow uses two stills. For controlled video composition with camera motion, see First and last frame.
Related
Last updated
Was this helpful?

