> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Script speech and performances

Make the character speak your exact text, perform to existing audio, or follow an authored movement. First [connect the player](/streaming-avatar/integrate/trtc), then send each JSON event with `window.player.ws.send(JSON.stringify(event))`.

## 1. Choose a script or a generated reply

Normal conversation and scripts solve different problems. **Conversation lets the avatar respond to the user in the moment; a script gives it the words or movements for one specific performance.** For an open-ended request such as “Give me a wave,” start with conversation. Use `response.script` when the line, recording, or movement is already prepared.

| Choose | Use it when |
| - | - |
| Dialogue | **Users ask free-form questions.** Set `avatars[].instructions` to shape replies; each turn generates its own speech and matching motion. |
| Dialogue | **A user casually asks for a wave or a small step.** This is an immediate request, not necessarily a prepared show. Make sure the image leaves room to move. |
| Script | **A welcome line must be exact, or you already have a recording.** Use `response.script` to supply words or audio without dialogue rewriting them. |
| Script | **A dance or sequence of actions is already planned.** Use `response.script` to supply the VMP and, if needed, prepared audio. |
| Dialogue + script | **The user asks for a prepared show by name.** Dialogue selects it; an application tool retrieves and sends its `response.script`. |

The deciding question is whether the avatar can improvise. If it can, write the character instructions and use conversation. If it must perform material you supplied, send a script. Both paths can be used in one Session, but a script does not change the avatar’s standing instructions. Audio and movement in a script are not guaranteed to align word by word or beat by beat.

## 2. Speak a line

```json theme={null}
{
  "type": "response.create",
  "response": {
    "script": {
      "vocal": {
        "type": "speech",
        "text": "Welcome back. Let me show you the next item."
      }
    }
  }
}
```

The current character voice synthesizes this text without a dialogue-model rewrite. Speech accepts plain text, not SSML. Use [conversation](/streaming-avatar/interaction/conversation) when the character should compose its own answer.

## 3. Perform a movement

```json wrap theme={null}
{
  "type": "response.create",
  "response": {
    "script": {
      "visual": {
        "prompt": "Static medium shot at eye level. The character faces the camera with both hands at waist height. The right hand rises to shoulder height, waves twice with an open palm, then returns to the waist. The left hand stays at the waist.",
        "duration_ms": 5000
      }
    }
  }
}
```

The visual description is a VMP (Visual Motion Prompt). This example supplies movement without speech. Write the camera, starting pose, movement, and ending pose in English. Keep arms visible for gestures and leave full-body space for dancing.

“Raise the right hand to shoulder height, wave twice with an open palm, then lower it” is more useful than “be enthusiastic.” Start with one main action. Avoid contradictory poses, and begin each later movement where the previous one ends.

### Write a VMP as visible movement

Organize each VMP as camera, a short character description, a scene anchor, starting pose, movement, and ending pose. Each segment must make sense on its own rather than saying “continue the previous action”. Place the main movement early in the action chain and keep appearance and finger detail proportionate. Separate incompatible poses into stages, and express amplitude using visible references such as shoulder height, shoulder width, or a number of steps.

Too vague:

```text wrap theme={null}
Be welcoming and wave enthusiastically. Continue as before.
```

An observable version, assuming this matches the current scene:

```text wrap theme={null}
Static waist-up shot at eye level. A woman with shoulder-length dark hair stands before a
plain wall, facing the camera. Both empty hands begin at waist height. Her right hand rises to
shoulder height, opens with the palm facing the camera, and moves side to side twice. The
right hand lowers to waist height; both hands finish relaxed and visible.
```

This makes positions and repetitions easier to assess, without guaranteeing frame-exact execution. Match the script’s starting pose to the current picture. Audio drives mouth behavior; leave lip, jaw, and synchronized-mouth instructions out of the VMP.

## 4. Combine sound and movement

This example is a full-body dance. The Quickstart image does not show the feet, so first use an image with the entire body, visible feet, and space on both sides. Adding full-body to a prompt alone does not make the original framing suitable for footwork.

```json wrap theme={null}
{
  "type": "response.create",
  "response": {
    "conversation": "none",
    "script": {
      "vocal": {
        "type": "audio",
        "url": "https://YOUR_DOMAIN/dance.mp3",
        "mime_type": "audio/mpeg"
      },
      "visual": {
        "prompt": "Static full-body shot. The dancer stands with feet shoulder-width apart and hands at waist height, steps to the right and brings the left foot beside it, then opens both arms to shoulder height.[SHOT_SEP]Static full-body shot. With both arms at shoulder height, the dancer steps to the left, brings the right foot beside it, and lowers both hands to the waist.[SHOT_SEP]Static full-body shot. With hands at the waist, the dancer bends both knees slightly, straightens, and finishes facing the camera with both feet planted.",
        "duration_ms": 5000
      }
    }
  }
}
```

Replace the URL with an accessible MP3. Use `audio/wav` for WAV files. You can instead send Base64 file bytes in `audio`, without a data URL prefix; choose exactly one of `url` and `audio`. Supplied recordings are used directly, without character-voice synthesis.

The example separates a right step, a left step, and a finish with `[SHOT_SEP]`. `duration_ms` is passed to each segment as its target duration, defaulting to 5000 milliseconds. Three segments with a 5000-millisecond target plan about 15 seconds of motion, rather than sharing five seconds. This does not change the supplied audio file’s duration. Actual timing also depends on the model and audio length; it is not a frame- or beat-accurate choreography timeline. Start with audio at a suitable pace, then adjust the movement and audio length after playback.

A `script` can contain `vocal`, `visual`, or both. Sound and movement are submitted separately, so a gesture is not guaranteed to land on a particular word or beat. `conversation: "none"` keeps this performance out of default conversation history.

## 5. Perform an existing song

Prepare a WAV or MP3 that already contains singing. Set `vocal.type` to `audio` and provide the file URL and MIME type. This example uses an MP3:

```json wrap theme={null}
{
  "type": "response.create",
  "response": {
    "conversation": "none",
    "script": {
      "vocal": {
        "type": "audio",
        "url": "https://YOUR_DOMAIN/song-with-vocals.mp3",
        "mime_type": "audio/mpeg"
      },
      "visual": {
        "prompt": "Static medium shot. The character faces the camera with both hands relaxed at the waist. While performing, the right hand opens gently at chest height, then returns to the waist. The character keeps a calm, engaged expression.",
        "duration_ms": 5000
      }
    }
  }
}
```

Replace the URL with audio that includes vocals. Ordinary `speech` synthesizes spoken lines; supplying lyrics does not make it sing a chosen melody. This uses the recording’s existing vocals rather than resinging them in the selected TTS voice. Test a short clip, then adjust the accompanying movement. The movement prompt does not provide beat-accurate synchronization.

## 6. Let users request a named performance

For example, an application with prepared `welcome_wave` and `dance_song` routines can define its own `play_performance` function tool, accepting a `performance_id` from that allowlist. These names are application-defined, not built-in performances or APIs. Register and implement the function using the [tool-calling guide](/streaming-avatar/character/tools), then merge these rules into `avatars[].instructions`.

```text wrap theme={null}
## Conversation or a prepared performance
For ordinary questions and open-ended action requests, answer naturally using the character
instructions and the normal motion path.
When the user explicitly asks for one of the application's available named performances, call
play_performance with that performance_id instead of improvising a different routine.
If the requested performance is unavailable or the name is ambiguous, state the available
choices and ask which one they want. Do not invent a performance ID.
When the selected performance already contains speech or singing, return the tool call without
an extra spoken introduction. The prepared audio supplies that part of the experience.
Do not announce that playback has finished when the tool only reports that the script was
submitted. Use the actual application result.
```

After receiving the completed function call, validate the performance ID, deduplicate the call, and submit its prepared `response.script` as a separate request. Do not add tools or instructions to that same script request. A routine with speech or singing usually needs no extra model-generated introduction that could interrupt it. Report whether the script was submitted or playback actually finished; claim the latter only after observing the corresponding playback result. If conversation should continue afterward, follow the tool guide’s result-acknowledgement and continuation steps.

## 7. Reuse content and track completion

Register reusable content with `asset.add` under `audio`, `speech_text`, or `visual_prompts`. Reference it with `audio_asset_id`, `speech_text_asset_id`, or `visual_prompt_asset_id`. See [client events](/streaming-avatar/api-references/client-events) for the complete fields.

Do not combine `script` with `input`, `instructions`, `tools`, or `max_output_tokens` in the same request. New requests interrupt by default. Set `response.scheduling_policy: "after_current_response"` to play next; only one request can wait.

Track events by `response_id`. Finished text does not mean finished playback: observe `response.render.stopped` and check the status in `response.done`. For missing sound, check browser playback permission, then the audio URL and MIME type. For unstable movement, test one action without music and check framing and the starting and ending poses.

## 8. FAQ

#### Is duration\_ms the total performance length?

No. With \[SHOT\_SEP], it is the target duration for each segment. The same value applies to all segments in that request. For two segments accompanying 22.7 seconds of audio, start by testing 11350 milliseconds per segment rather than 22700 for each. Inspect the later segment too; model and audio constraints still affect the result.
