pipeline_config.motion_enhanced.prompt only when you need a clearer gesture style, a quieter range of motion, or a better visual response to requests such as “wave.” This prompt controls movement while the avatar speaks. Speaking instructions control what it says. For a rehearsed dance or fixed sequence, use a script.
1. Role and Objective: turn a spoken line into visible motion
This prompt directs the avatar’s body while it speaks; it does not write the spoken line. The motion director considers four inputs together: what the user wants, what the avatar actually says, how this character tends to move, and what the current frame shows. Following dialogue alone can miss an explicit request to wave. Following the request alone can make the action feel disconnected from the reply. Start Role and Objective by assigning the VMP director’s job, then state both goals: make body language visibly expressive when the moment calls for it, and preserve continuity with the current frame.Make the character lively. does neither on its own.
If a user asks for a wave and the avatar replies, “Sure, hello!”, the wave should be the main visible action. If the avatar is answering a question, let the meaning of the answer guide its gesture. Audio handles lip sync; a VMP does not need lip or jaw directions.
2. Make the movement feel like this character in this scene
A custom prompt needs two kinds of context: how this character tends to move and where the next movement starts. The two tokens in braces make room for that context. They are not extra Create Session fields. Keep them when adapting the example; edit the movement style you want.2.1 Character: a recognizable movement style
{summarized_persona} receives motion-related character information such as energy, gesture habits, and range of movement. It comes from the existing character setup and images; you do not need a separate persona field or a saved Character resource. It may help a reserved character move more sparingly and a playful one move with more lift. It does not say where the character is standing now or what a hand holds. Leave the token in place when using this example.
2.2 Current Frame: start where the avatar is now
{VLM} is where a description of the current picture goes. When one is available, movement can continue from the avatar’s present position and occupied hands: an avatar that stepped right starts on the right; a cup in the right hand stays there. If the current picture is temporarily unavailable, keep the action small rather than guessing hand or prop positions or resetting to the first image. Keep the token; do not type in a permanent “current frame.”
Wrong starting point: She raises both hands to wave. If her right hand holds a cup, this makes it disappear.
Grounded in the frame: Her left hand rises to shoulder height while her right hand keeps holding the cup.
3. Motion Rules: choose what happens in this segment
Check first for an explicit action request; then read the line’s tone and the character’s movement style. Give each segment one main action. Name the body part, its starting position, the movement, and where it settles. The result should be visible and give the next segment a definite starting point.3.1 A requested action: do the requested thing
“Give me a wave” calls for a wave, not a polite nod. Make it large enough to read at the current framing, within what the picture allows. If the right hand holds something, use the free left hand. Do not invent a chair for a sitting action. Too weak:She smiles and nods warmly.Clearer:
Her empty left hand rises from her waist to shoulder height, waves twice, and settles beside her body.
3.2 No action request: follow the meaning of the line
An explanation may call for an open palm tracing a space; reassurance may need a slower, smaller movement. Avoid assigning one stock gesture to every emotion, or turning “playful” into constant motion. Vague:She gestures expressively.Usable:
Her right hand rises from her side, opens near her chest, then lowers beside her body.
3.3 Two segments: hand the ending pose to the next
If the right hand ends near the chest, the next segment begins with it there before lowering it or starting another action.continues the previous gesture is not a starting pose. Check the source image before asking for a broad movement that needs room outside the crop.
4. Segment Content Order: write a visible VMP
A VMP is an English description for the generated picture, not a line to be spoken. Treat each segment as a short shot: where the avatar starts, what it does, and where it settles. Write each VMP in English, in this order: camera and framing → short appearance anchor → scene → absolute starting pose → movement → clear ending pose. Start with a fixed camera for ordinary conversation. Keep appearance short so it does not bury the action. Every segment should make sense on its own: describe its starting pose instead of sayingcontinues as before. Use the woman or the man rather than a character name inside the VMP.
If one reply spans two segments, the second must open where the first ended. If the first ends with the right hand at chest height, the second starts there. Describe environmental motion only when something physical, such as wind or contact, would cause it.
5. Output Format: keep the result readable
The two segments below show the hand rising to chest height, then lowering from that position.motions may contain multiple segments; each has an English vmp and a duration. This illustrates the shape. The real starting pose must match the current frame.
6. Add it to a Session and watch the result
The complete prompt below is a starting point. First edit the style and range under Motion Rules. Leave{summarized_persona}, {VLM}, and the rule for an unavailable current picture in place. Keep the VMP order and JSON shape. Leave this optional setting out when default motion meets your needs.
pipeline_config.motion_enhanced.prompt when you create the Session. Try an explanation and a request to wave. Can you see the action? Does a held object stay in its hand? If the reply spans segments, does the next begin where the last ended? Change one rule at a time and compare in a new session.