1. Choose an image that fits the session
For 9:16 portrait output, start with a portrait image of the same ratio; match a landscape image to landscape output. A chest-up to waist-up shot works well for conversation. A knee-up shot can give the hands more room, as long as the eyes and mouth stay clear. Keep the hands fully inside the frame rather than cropping them at the wrists. As a quick check, divide the width of the face box by the image’s shorter edge: image width in portrait, image height in landscape. This is a selection guide, not an upload limit.- Start here: 25% or more. A chest-up, waist-up, or close knee-up image keeps the face clear and leaves room for gestures.
- Try a closer image: below 25%. Full-body or distant shots can lose facial detail during conversation.
2. Compare two images
| Good | Less suitable |
| Face and hands are clear, with room for gestures. | Full outfit is visible, but the face is small for close conversation. |
3. Add an image description when useful
source_images[].description is optional. Add a short English description when you need to make the visible outfit, pose, or hand-object relationship explicit. Describe what is already in the image, not an action you want the avatar to perform later.
avatars[0].visual.source_images when creating a Session. This fragment shows only the image fields; see the Quickstart for a full request.
description; adapt it to the image you use, or omit the field.