pipeline_config.motion_planner.prompt only if it looks frozen, repeats a gesture too often, or snaps back to its first-image pose after moving. This prompt controls how it stays present while listening or waiting. Gestures during a reply belong to speaking motion.
1. Role and Objective: stay present between replies
This prompt directs one small continuation movement while the avatar is not speaking. It is not a quieter version of speaking motion or a place to start another performance. The avatar should appear attentive and at ease, beginning where its body actually came to rest. The result is one short VMP. Name the job in Role and Objective: an Idle-Moment Motion Planner that outputs exactly one short continuation segment.Keep the avatar alive and moving. is too open-ended; it can turn listening into constant nodding, waving, or a replay of the last performance. Ask for one motivated adjustment and a stable resting pose. Sometimes that adjustment should be barely noticeable.
2. Continue from the last visible moment
Finishing a reply does not return the avatar to its first image. Listening motion needs both the character’s usual way of moving and the pose it just reached. The two tokens in braces make room for that context at runtime. Keep them when copying this example; do not type in a fixed “current position.”2.1 Character: keep quiet moments in character
{summarized_persona} receives motion-related character information, such as a restrained or lively style. It is not the current pose. You do not need another persona field or a saved Character resource. Keep the token when adapting this example. To change the range or frequency of listening movements, edit Motion Rules.
2.2 Current Frame: pick up where the last action ended
{VLM} is where current-picture information goes. When a picture is available, an avatar that just stepped right begins listening on the right, and a cup in the right hand stays there. If no current picture is available, do not invent hand or prop positions or send the avatar back to its opening pose. Keep this token and the example’s small-motion fallback.
Causes a reset: Return to the original centered pose and fold both hands.
Continues the scene: She stays on the right side of the frame and lets her free left shoulder relax.
The first image establishes the opening, not a home position to revisit whenever the avatar goes quiet. Do not invent a seat when there is no visible support.
3. Motion Rules: one quiet continuation
Listening does not need a new routine every time. Look at the preceding action and the current positions of the body and hands, then choose one restrained movement. It may end in a new stable pose; there is no need to return to the source image or the previous starting pose.3.1 A speaking gesture just ended: let it settle
After a wave, let that hand come down. Do not add another wave, three nods, and a weight shift. A larger speaking beat often calls for a quieter listening beat. Still performing:She waves again, steps forward, and nods repeatedly.Listening:
Her raised left hand lowers beside her hip, and her shoulders settle.
3.2 A hand is occupied: keep it occupied
If the right hand holds a cup, it continues to hold it. A small adjustment can come from the free left side or the shoulders. Mentioning an object does not place it in the frame. Usable:Her right hand holds the cup near her waist while her free left shoulder relaxes.
3.3 The avatar just moved: stay where it arrived
If it stepped to the right side of the picture, its listening pose belongs there too. It does not need another step to seem alive, and it should not drift back to center. See Natural object interaction for sustained contact with props. Usable:On the right side of the frame, she shifts her weight gently and settles with both feet planted.
4. Segment Content Order: from the frame to a resting pose
A VMP is an English description for the generated picture, not spoken dialogue. For listening, state the current pose, one small action, and the resting pose. Write one English VMP in this order: fixed camera and current framing → short appearance anchor → scene anchor → absolute body and hand positions → one listening movement → clear ending pose.Stay natural is too vague, and continue the previous pose does not say where the hands and body actually are. Describe physical contact only where there is visible support, and environmental motion only when there is a physical cause.
5. Output Format: exactly one segment
Speaking motion can span several segments. Listening motion returns amotion array with exactly one item. That item has an English vmp and no duration. Return valid JSON alone, without commentary or extra fields.
6. Add it to a Session and watch several turns
The complete prompt below is a starting point. First edit the range and frequency under Motion Rules. Keep{summarized_persona}, {VLM}, and the small-motion fallback for an unavailable picture. Keep the VMP order and JSON shape. Leave this optional setting out when default listening motion already looks natural.
pipeline_config.motion_planner.prompt when you create the Session. Let the avatar finish a reply with a visible action, then wait a few seconds. Does it stay where it arrived, keep hold of any object, and avoid repeating the same gesture? Ask another question to check that speaking begins from the listening end pose.