Skip to main content
Make the character speak your exact text, perform to existing audio, or follow an authored movement. First connect the player, then send each JSON event with window.player.ws.send(JSON.stringify(event)).

1. Choose a script or a generated reply

Normal conversation and scripts solve different problems. Conversation lets the avatar respond to the user in the moment; a script gives it the words or movements for one specific performance. For an open-ended request such as “Give me a wave,” start with conversation. Use response.script when the line, recording, or movement is already prepared. The deciding question is whether the avatar can improvise. If it can, write the character instructions and use conversation. If it must perform material you supplied, send a script. Both paths can be used in one Session, but a script does not change the avatar’s standing instructions. Audio and movement in a script are not guaranteed to align word by word or beat by beat.

2. Speak a line

The current character voice synthesizes this text without a dialogue-model rewrite. Speech accepts plain text, not SSML. Use conversation when the character should compose its own answer.

3. Perform a movement

The visual description is a VMP (Visual Motion Prompt). This example supplies movement without speech. Write the camera, starting pose, movement, and ending pose in English. Keep arms visible for gestures and leave full-body space for dancing. “Raise the right hand to shoulder height, wave twice with an open palm, then lower it” is more useful than “be enthusiastic.” Start with one main action. Avoid contradictory poses, and begin each later movement where the previous one ends.

Write a VMP as visible movement

Organize each VMP as camera, a short character description, a scene anchor, starting pose, movement, and ending pose. Each segment must make sense on its own rather than saying “continue the previous action”. Place the main movement early in the action chain and keep appearance and finger detail proportionate. Separate incompatible poses into stages, and express amplitude using visible references such as shoulder height, shoulder width, or a number of steps. Too vague:
An observable version, assuming this matches the current scene:
This makes positions and repetitions easier to assess, without guaranteeing frame-exact execution. Match the script’s starting pose to the current picture. Audio drives mouth behavior; leave lip, jaw, and synchronized-mouth instructions out of the VMP.

4. Combine sound and movement

This example is a full-body dance. The Quickstart image does not show the feet, so first use an image with the entire body, visible feet, and space on both sides. Adding full-body to a prompt alone does not make the original framing suitable for footwork.
Replace the URL with an accessible MP3. Use audio/wav for WAV files. You can instead send Base64 file bytes in audio, without a data URL prefix; choose exactly one of url and audio. Supplied recordings are used directly, without character-voice synthesis. The example separates a right step, a left step, and a finish with [SHOT_SEP]. duration_ms is passed to each segment as its target duration, defaulting to 5000 milliseconds. Three segments with a 5000-millisecond target plan about 15 seconds of motion, rather than sharing five seconds. This does not change the supplied audio file’s duration. Actual timing also depends on the model and audio length; it is not a frame- or beat-accurate choreography timeline. Start with audio at a suitable pace, then adjust the movement and audio length after playback. A script can contain vocal, visual, or both. Sound and movement are submitted separately, so a gesture is not guaranteed to land on a particular word or beat. conversation: "none" keeps this performance out of default conversation history.

5. Perform an existing song

Prepare a WAV or MP3 that already contains singing. Set vocal.type to audio and provide the file URL and MIME type. This example uses an MP3:
Replace the URL with audio that includes vocals. Ordinary speech synthesizes spoken lines; supplying lyrics does not make it sing a chosen melody. This uses the recording’s existing vocals rather than resinging them in the selected TTS voice. Test a short clip, then adjust the accompanying movement. The movement prompt does not provide beat-accurate synchronization.

6. Let users request a named performance

For example, an application with prepared welcome_wave and dance_song routines can define its own play_performance function tool, accepting a performance_id from that allowlist. These names are application-defined, not built-in performances or APIs. Register and implement the function using the tool-calling guide, then merge these rules into avatars[].instructions.
After receiving the completed function call, validate the performance ID, deduplicate the call, and submit its prepared response.script as a separate request. Do not add tools or instructions to that same script request. A routine with speech or singing usually needs no extra model-generated introduction that could interrupt it. Report whether the script was submitted or playback actually finished; claim the latter only after observing the corresponding playback result. If conversation should continue afterward, follow the tool guide’s result-acknowledgement and continuation steps.

7. Reuse content and track completion

Register reusable content with asset.add under audio, speech_text, or visual_prompts. Reference it with audio_asset_id, speech_text_asset_id, or visual_prompt_asset_id. See client events for the complete fields. Do not combine script with input, instructions, tools, or max_output_tokens in the same request. New requests interrupt by default. Set response.scheduling_policy: "after_current_response" to play next; only one request can wait. Track events by response_id. Finished text does not mean finished playback: observe response.render.stopped and check the status in response.done. For missing sound, check browser playback permission, then the audio URL and MIME type. For unstable movement, test one action without music and check framing and the starting and ending poses.

8. FAQ

Is duration_ms the total performance length?

No. With [SHOT_SEP], it is the target duration for each segment. The same value applies to all segments in that request. For two segments accompanying 22.7 seconds of audio, start by testing 11350 milliseconds per segment rather than 22700 for each. Inspect the later segment too; model and audio constraints still affect the result.