> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Sample Case: Live Performance

Sienna is a livestream singer standing center stage in a full-body shot, facing her audience: she sings, dances, takes song requests, and chats with viewers between numbers. This scenario covers the other core path of Streaming Avatar — **script-driven rendering**: you specify the lines and the motion directly, and no LLM generation is involved.

<Columns cols={2}>
  <Card title={"Full-Body Live Performance"} img={"https://static.vivi-x.ai/images/2026/07/08/1c5ae7ed82cf69de.png"}>
    Create full-body performers that move, interact with their surroundings, and respond live.
  </Card>
</Columns>

<h2 id="create-session-differences">
  Creating the Session: Differences from the Claire Scenario
</h2>

The session creation request has exactly the same structure as the [Customer Support](/overview/sample-cases/customer-support) scenario — create the session with `POST /v1/realtime-avatar/sessions`, then use the response to join the TRTC media stream and connect to the WSS control channel. Only the material differs. To let the avatar perform full-body motion, the reference image must be a **full-body shot** — Sienna's reference image is a centered full-body shot at center stage. Configure her session voice with the top-level `pipeline_config.tts_config` field; keep voice configuration out of the avatar object:

Use the same [media and Control connection order](/overview/sample-cases/customer-support#step-2-connect-media-and-control) as the Claire scenario: enter TRTC and subscribe only to publisher\_user\_id first, then append token=\<control.client\_secret> to the returned control.url and open the browser Control WSS. Do not connect to the original URL without the token, and do not name the query parameter client\_secret. Trigger response.script only after both channels are ready.

```json Request body excerpt theme={null}
{
  "pipeline_config": {
    "tts_config": {
      "tts_provider": "qwen-audio-3.0-tts-flash_ws",
      "tts_voice_id": "longanhuan_v3.6",
      "tts_model_id": "qwen-audio-3.0-tts-flash"
    }
  },
  "avatars": [
    {
      "avatar_id": "sienna",
      "instructions": "You are Sienna Blake, a charismatic live music host and singer. Keep replies concise, warm, and rhythmic, always moving the show forward. Output only natural spoken language - no narration, no stage directions, no emojis.",
      "visual": {
        "default_source_image_id": "stage",
        "source_images": [
          {
            "source_image_id": "stage",
            "url": "https://cdn.example.com/avatars/sienna-stage.png",
            "description": "A singer on stage facing the camera under bright stage lights.",
            "media_type": "image/png"
          }
        ]
      }
    }
  ]
}
```

<h2 id="response-script">
  Core Capability: response.script
</h2>

`response.script` renders avatar output that you specify directly. It **bypasses the LLM**: the model never rewrites your lines, calls tools, or consults `response.instructions`. Constraints and defaults:

* Once `script` is set, `input`, `instructions`, `tools`, and `max_output_tokens` cannot be sent in the same request.
* `script` defaults to the current `session.active_avatar_id`, the session voice from `pipeline_config.tts_config`, and the current source image.
* If the performance should not be written into the default conversation, pass `"conversation": "none"`.

<h3 id="script-sing-and-dance">
  Sing and Dance
</h3>

`script.vocal` carries an audio clip (the full song), and `script.visual.prompt` describes the dance moves:

```json response.create — sing and dance theme={null}
{
  "event_id": "evt_response_script_song_001",
  "type": "response.create",
  "response": {
    "scheduling_policy": "interrupt",
    "script": {
      "vocal": {
        "type": "audio",
        "url": "https://cdn.example.com/audio/performance.mp3",
        "mime_type": "audio/mpeg"
      },
      "visual": {
        "prompt": "follow the rhythm of the song, sway gently, step to the side, raise one hand, then smile toward the camera like a livestream host performing a short dance",
        "duration_ms": 10000
      }
    }
  }
}
```

<h3 id="script-fixed-line">
  Fixed Line and Motion
</h3>

Switch `vocal.type` to `speech` and pass the line as text; the session voice configured in `pipeline_config.tts_config` synthesizes it:

```json response.create — fixed line theme={null}
{
  "event_id": "evt_response_script_line_001",
  "type": "response.create",
  "response": {
    "script": {
      "vocal": {
        "type": "speech",
        "text": "I saw your request for Mirror Counting! This one is an original of mine - let us get into it."
      },
      "visual": {
        "prompt": "smile, give a small wave, then look toward the camera",
        "duration_ms": 3000
      }
    }
  }
}
```

For interstitial announcements and warm-up lines that should not pollute the conversation context, add `"conversation": "none"`:

```json response.create — out-of-band line theme={null}
{
  "event_id": "evt_oob_script_001",
  "type": "response.create",
  "response": {
    "conversation": "none",
    "script": {
      "vocal": {
        "type": "speech",
        "text": "This line is rendered on stage but never written into the default conversation."
      }
    }
  }
}
```

<h3 id="script-motion-only">
  Motion Only, No Speech
</h3>

Provide only `visual` inside `script` and the avatar performs the motion in silence — useful for keeping Sienna moving to the music between numbers:

```json response.create — motion only theme={null}
{
  "event_id": "evt_response_visual_only_001",
  "type": "response.create",
  "response": {
    "script": {
      "visual": {
        "prompt": "stay on one fixed center mark, slowly tilt your head, then softly open both arms from your waist to shoulder width",
        "duration_ms": 5000
      }
    }
  }
}
```

<h2 id="motion-prompts">
  Writing Motion Prompts
</h2>

Write `visual.prompt` as natural English describing continuous, concrete motion, one prompt per `duration_ms`. For long performances, split the choreography into shots: Sienna's dance material uses one motion prompt per shot of roughly five seconds, and each segment repeats the key constraints — stay on a fixed mark, keep the steps small, never move toward or away from the camera — so the motion stays stable over long stretches. A shot description used in production:

```text theme={null}
The dancer stays on one fixed center mark, facing the camera. She slowly tilts her head,
then softly opens both arms from her waist to shoulder width. Both feet remain low and
nearly still, and her body never moves toward or away from the camera.
```

Submit the split shots in order with `scheduling_policy: "after_current_response"` to queue them into one complete dance. Note that at most one response can be queued at a time — wait until the previous segment starts playing before submitting the next (see `scheduling_policy` in [Client events](/streaming-avatar/api-references/client-events)). For more on writing motion prompts, see [Script speech and performances](/streaming-avatar/interaction/scripted-performances).

<h2 id="scheduling-and-cancel">
  Queueing and Interruption: scheduling\_policy and response.cancel
</h2>

Live shows constantly hit the case where the next request arrives before the current dance is over. Control the conflict behavior with `response.scheduling_policy`:

| Value | Behavior |
| - | - |
| `interrupt` | Default. Interrupts the current response and executes the new response immediately. |
| `after_current_response` | If the avatar is busy, queues the new response until the current response finishes playing (in `video_avatar` mode, waits for the `response.done` terminal state). |
| `reject_if_busy` | Returns an `error` immediately when busy. |

When a viewer switches songs or calls off the performance, cancel the in-flight response with `response.cancel` — omit `response_id` to cancel the current response, or pass one to cancel a specific response:

```json response.cancel theme={null}
{
  "event_id": "evt_response_cancel_001",
  "type": "response.cancel"
}
```

A successful cancel still produces a terminal event, normally `response.done` with `response.status` set to `"cancelled"`. If there is no cancellable response at the moment, the server returns an `error`.

<h2 id="live-event-sequence">
  Event Sequence
</h2>

The send and receive order for one sing-and-dance performance. This example drives the vocal directly with audio (`vocal.type = audio`), so there are no `response.output_text.*` events; for `speech` scripts, the lines are still delivered via `response.output_text.*` events (see [Script speech and performances](/streaming-avatar/interaction/scripted-performances)):

```text theme={null}
-> response.create({ response.script.vocal = audio, response.script.visual = dance prompt })
<- response.created
<- response.render.started
<- response.render.stopped
<- response.done
```

<h2 id="live-next-steps">
  Next Steps
</h2>

* [Customer Support](/overview/sample-cases/customer-support) — the other sample case, focused on conversation, expression, and emotion: the full path from creating a session to writing user messages and having the avatar generate LLM replies rendered as realtime audio and video.
* [Script speech and performances](/streaming-avatar/interaction/scripted-performances) — scripted lines, performances driven by an audio file, and directed movement.
* [Change outfits and scenes](/streaming-avatar/interaction/outfits-and-scenes) — switch the avatar image during a session.
* [API References · Sessions](/streaming-avatar/api-references/sessions) — REST endpoints and parameter details.
* [API References · Client events](/streaming-avatar/api-references/client-events) / [Server events](/streaming-avatar/api-references/server-events) — the complete event mechanism.
* [Streaming World](/streaming-world/configuration) — realtime interactive video worlds driven by the W-series models, coming soon.
