Skip to main content
Sienna is a livestream singer standing center stage in a full-body shot, facing her audience: she sings, dances, takes song requests, and chats with viewers between numbers. This scenario covers the other core path of Streaming Avatar — script-driven rendering: you specify the lines and the motion directly, and no LLM generation is involved.
1c5ae7ed82cf69de

Full-Body Live Performance

Create full-body performers that move, interact with their surroundings, and respond live.

Creating the Session: Differences from the Claire Scenario

The session creation request has exactly the same structure as the Customer Support scenario — create the session with POST /v1/realtime-avatar/sessions, then use the response to join the TRTC media stream and connect to the WSS control channel. Only the material differs. To let the avatar perform full-body motion, the reference image must be a full-body shot — Sienna’s reference image is a centered full-body shot at center stage. Configure her session voice with the top-level pipeline_config.tts_config field; keep voice configuration out of the avatar object: Use the same media and Control connection order as the Claire scenario: enter TRTC and subscribe only to publisher_user_id first, then append token=<control.client_secret> to the returned control.url and open the browser Control WSS. Do not connect to the original URL without the token, and do not name the query parameter client_secret. Trigger response.script only after both channels are ready.
Request body excerpt

Core Capability: response.script

response.script renders avatar output that you specify directly. It bypasses the LLM: the model never rewrites your lines, calls tools, or consults response.instructions. Constraints and defaults:
  • Once script is set, input, instructions, tools, and max_output_tokens cannot be sent in the same request.
  • script defaults to the current session.active_avatar_id, the session voice from pipeline_config.tts_config, and the current source image.
  • If the performance should not be written into the default conversation, pass "conversation": "none".

Sing and Dance

script.vocal carries an audio clip (the full song), and script.visual.prompt describes the dance moves:
response.create — sing and dance

Fixed Line and Motion

Switch vocal.type to speech and pass the line as text; the session voice configured in pipeline_config.tts_config synthesizes it:
response.create — fixed line
For interstitial announcements and warm-up lines that should not pollute the conversation context, add "conversation": "none":
response.create — out-of-band line

Motion Only, No Speech

Provide only visual inside script and the avatar performs the motion in silence — useful for keeping Sienna moving to the music between numbers:
response.create — motion only

Writing Motion Prompts

Write visual.prompt as natural English describing continuous, concrete motion, one prompt per duration_ms. For long performances, split the choreography into shots: Sienna’s dance material uses one motion prompt per shot of roughly five seconds, and each segment repeats the key constraints — stay on a fixed mark, keep the steps small, never move toward or away from the camera — so the motion stays stable over long stretches. A shot description used in production:
Submit the split shots in order with scheduling_policy: "after_current_response" to queue them into one complete dance. Note that at most one response can be queued at a time — wait until the previous segment starts playing before submitting the next (see scheduling_policy in Client events). For more on writing motion prompts, see Script speech and performances.

Queueing and Interruption: scheduling_policy and response.cancel

Live shows constantly hit the case where the next request arrives before the current dance is over. Control the conflict behavior with response.scheduling_policy: When a viewer switches songs or calls off the performance, cancel the in-flight response with response.cancel — omit response_id to cancel the current response, or pass one to cancel a specific response:
response.cancel
A successful cancel still produces a terminal event, normally response.done with response.status set to "cancelled". If there is no cancellable response at the moment, the server returns an error.

Event Sequence

The send and receive order for one sing-and-dance performance. This example drives the vocal directly with audio (vocal.type = audio), so there are no response.output_text.* events; for speech scripts, the lines are still delivered via response.output_text.* events (see Script speech and performances):

Next Steps