> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Customer Support Avatar

A copy-ready sample case: Claire, a customer support avatar built around conversation, expression, and emotion. The default demo accepts both microphone speech and typed messages, generates replies with an LLM, and renders audio and video in real time — a complete walk through the Streaming Avatar flow, from creating to closing the session.

<Columns cols={2}>
  <Card title={"Customer Support"} img={"https://static.vivi-x.ai/images/2026/06/25/c90cf1b4832cb2d0.png"}>
    Engage users in natural, personalized conversations using your knowledge base.
  </Card>
</Columns>

Before you start, follow the [Quickstart in Introduction](/overview/introduction#quickstart) to create an account and generate an API key. All REST requests use `https://api.vivix.ai` as the base URL and carry the `Authorization: Bearer <API_KEY>` header. Every REST response is wrapped in `{"code": 0, "message": "success", "data": {...}}` — the business fields live in `data`.

Claire is a customer support avatar who sits in an office, facing the camera: she answers product questions, troubleshoots integration issues, and keeps a calm, professional expression and tone throughout the conversation. This scenario covers the most typical Streaming Avatar usage — **conversation-driven**: you write what the user says into the conversation context, the avatar generates a reply from that context, and the reply streams out in real time as speech, rendered video, and word-by-word captions.

<h2 id="core-concepts">
  Core Concepts
</h2>

A Streaming Avatar session is built from four objects:

| Concept | Description |
| - | - |
| **Session** | A single realtime session. Created over REST; the response returns `session_id`, the WSS control channel (`control`), and media stream info (`delivery.media`). |
| **Avatar** | A switchable character definition inside the session. This sample provides `visual` (reference images) and `instructions` (persona and tone), and configures the session voice separately with `pipeline_config.tts_config`. |
| **Conversation** | The context of this session. You write what the user says — and extra facts from your business systems — as conversation items. |
| **Response** | One round of avatar output. `response.create` makes the current active avatar generate a reply from the context; fixed lines, singing, and dancing use `response.script`. |

Three channels carry the flow, each with its own job:

1. **REST** (`/v1/realtime-avatar/sessions`) — the session lifecycle: create, retrieve, close.
2. **WSS control channel** (`control.url`, pointing at `/v1/realtime-avatar/control`) — sends and receives conversation / response / session events.
3. **Media stream** (chosen by `delivery.media.transport`; currently supports `trtc` and `agora`) — carries generated avatar audio and video downstream and user microphone audio upstream.

<h2 id="step-1-create-session">
  Step 1: Create a Session
</h2>

Use `POST /v1/realtime-avatar/sessions` to initialize the session's selectable avatars, default avatar, voice pipeline, and conversation defaults. `avatars[]` defines Claire's appearance and persona: `visual.source_images` provides the reference images (they decide what the avatar looks like and what scene she is in), `instructions` writes the persona, and the top-level `pipeline_config.tts_config` selects the voice used for generated speech.

```bash Request theme={null}
curl -X POST https://api.vivix.ai/v1/realtime-avatar/sessions \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vivix-a1-stream",
    "session": {
      "mode": "video_avatar",
      "active_avatar_id": "claire",
      "source_image_id": "default"
    },
    "output": {
      "aspect_ratio": "9:16",
      "resolution": "480p"
    },
    "pipeline_config": {
      "tts_config": {
        "tts_provider": "qwen-audio-3.0-tts-flash_ws",
        "tts_voice_id": "loongeva_v3.6",
        "tts_model_id": "qwen-audio-3.0-tts-flash"
      }
    },
    "conversation": {
      "instructions": "Help users understand Vivix products and API onboarding. Ask one clarifying question when the goal is unclear.",
      "response_defaults": {
        "temperature": 0.7,
        "max_output_tokens": 512
      },
      "turn_detection": {
        "type": "server_vad",
        "silence_duration_ms": 500
      },
      "input_audio_transcription": {
        "enabled": true,
        "language": "en"
      }
    },
    "avatars": [
      {
        "avatar_id": "claire",
        "instructions": "You are Claire Bennett, a calm, professional Vivix customer support specialist. Identify the user goal first, explain in plain language, and recommend the clearest next step. Output only natural spoken language - no speaker labels, no narration, no stage directions, no Markdown, no emojis.",
        "visual": {
          "default_source_image_id": "default",
          "source_images": [
            {
              "source_image_id": "default",
              "url": "https://cdn.example.com/avatars/claire.png",
              "description": "A customer support specialist facing the camera in a brightly lit studio, front view.",
              "media_type": "image/png"
            }
          ]
        }
      }
    ]
  }'
```

<Note>
  **Model choice.** The `model` field also accepts the lighter `vivix-a1-stream-lite` variant — see [Models](/overview/models) for the differences. For the full field tables, see the [Sessions API](/streaming-avatar/api-references/sessions).
</Note>

The response (some fields omitted for readability):

```json Response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "session_id": "avatar-session-123",
    "status": "active",
    "session": {
      "mode": "video_avatar",
      "active_avatar_id": "claire",
      "source_image_id": "default"
    },
    "control": {
      "url": "wss://api.vivix.ai/v1/realtime-avatar/control?session_id=avatar-session-123",
      "client_secret": "ctl_..."
    },
    "output": {
      "aspect_ratio": "9:16",
      "resolution": "480p",
      "fps": 24
    },
    "conversation": {
      "instructions": "Help users understand Vivix products and API onboarding. Ask one clarifying question when the goal is unclear.",
      "response_defaults": {
        "temperature": 0.7,
        "max_output_tokens": 512
      },
      "turn_detection": {
        "type": "server_vad",
        "silence_duration_ms": 500
      },
      "input_audio_transcription": {
        "enabled": true,
        "language": "en"
      }
    },
    "avatars": [
      {
        "avatar_id": "claire",
        "instructions": "You are Claire Bennett, ..."
      }
    ],
    "delivery": {
      "media": {
        "transport": "trtc",
        "trtc": {
          "sdk_app_id": "1400000000",
          "room_id": "avatar-session-123",
          "user_id": "viewer",
          "user_sig": "...",
          "publisher_user_id": "publisher_xxx"
        }
      }
    },
    "model": "vivix-a1-stream"
  }
}
```

From the response, save:

* `session_id`
* `control.url` and `control.client_secret`
* `delivery.media`
* the server-side effective `session.mode` / `active_avatar_id` / `source_image_id`

<h2 id="step-2-connect-media-and-control">
  Step 2: Connect Media, Then the Control Channel
</h2>

In `video_avatar` mode, join the TRTC media stream first. The default demo requests microphone permission after the user clicks Start and must run over HTTPS. Convert sdk\_app\_id to a number, pass the string room id as strRoomId, register remote video handlers before entering the room, and subscribe only to publisher\_user\_id:

```javascript TRTC setup theme={null}
import TRTC from "trtc-sdk-v5";

const media = session.delivery.media.trtc;
const publisherUserId = media.publisher_user_id;
const { microphone } = await TRTC.getPermissions({
  request: true,
  types: ["microphone"],
});
const trtc = TRTC.create();

trtc.on(TRTC.EVENT.REMOTE_VIDEO_AVAILABLE, ({ userId, streamType }) => {
  if (userId !== publisherUserId) return;
  void trtc.startRemoteVideo({ userId, streamType, view: "avatar-video" });
});

await trtc.enterRoom({
  sdkAppId: Number(media.sdk_app_id),
  userId: media.user_id,
  userSig: media.user_sig,
  strRoomId: String(media.room_id),
});

if (microphone !== "denied") {
  await trtc.startLocalAudio({ publish: true });
}
```

If the user denies microphone permission, disable the microphone button and keep typed input available. After entering the media room, connect to the returned `control.url` and append the session token as a token query parameter:

```javascript Browser Control WSS theme={null}
const controlUrl = new URL(session.control.url);
controlUrl.searchParams.set("token", session.control.client_secret);
const ws = new WebSocket(controlUrl.toString());
```

The query parameter name must be token and its value is the control.client\_secret returned by Create Session. Do not open the original control.url without adding it, and do not name the query parameter client\_secret. URL.searchParams preserves the existing session\_id and encodes the credential. Server-side clients may instead send Authorization: Bearer \<control.client\_secret>. Wait for both TRTC enterRoom and the WebSocket open event before enabling voice and text input.

<Note>
  `control.client_secret` is the browser-safe, session-scoped control credential (never your API key). Do not log it or the complete `control.url` out of logs as well.
</Note>

<h2 id="step-3-write-user-message">
  Step 3: Accept Voice and Text Input
</h2>

The default demo keeps both a microphone button and a text box. Both input paths share the same conversation and response output events.

### Voice Input

Step 1 enables `turn_detection.type: "server_vad"` and audio transcription. `startLocalAudio` publishes microphone audio to the same TRTC room. When the user finishes speaking, the server runs ASR, appends the user conversation item, and creates a response automatically. Voice turns do not send conversation.item.create or response.create from the client:

```text Voice turn theme={null}
-> microphone audio over TRTC
<- input_audio.speech_started
<- conversation.item.input_audio_transcription.delta
<- input_audio.speech_stopped
<- conversation.item.input_audio_transcription.completed
<- conversation.item.created
<- response.created
```

### Text Input

Use `conversation.item.create` to write a submitted text message into the conversation context:

```json conversation.item.create theme={null}
{
  "event_id": "evt_item_create_001",
  "type": "conversation.item.create",
  "item": {
    "type": "message",
    "role": "user",
    "content": [
      {
        "type": "input_text",
        "text": "How do I create an API key?"
      }
    ]
  }
}
```

The server confirms with `conversation.item.created`:

```json conversation.item.created theme={null}
{
  "event_id": "evt_item_created_001",
  "type": "conversation.item.created",
  "previous_item_id": null,
  "item": {
    "id": "item_001",
    "type": "message",
    "status": "completed",
    "role": "user",
    "content": [
      {
        "type": "input_text",
        "text": "How do I create an API key?"
      }
    ]
  }
}
```

Wait for this acknowledgment before sending `response.create`. Do not write conversation.item.create and response.create back-to-back in the same send loop; the response could otherwise start before the context item is confirmed.

Two things to note:

* `conversation.item.create` for typed input only adds an item to the context — it **does not automatically make the avatar speak**. After the acknowledgment, send `response.create`. In server\_vad mode, the server performs both steps for voice input.
* When `previous_item_id` is omitted, the item is appended to the end of the context; pass it only when you need to insert after a specific item.

<h2 id="step-4-make-claire-speak">
  Step 4: Trigger a Reply for Text Input
</h2>

After receiving conversation.item.created for typed input, send `response.create`. The current active avatar generates a reply from the conversation context:

```json response.create theme={null}
{
  "event_id": "evt_response_create_001",
  "type": "response.create",
  "response": {
    "scheduling_policy": "interrupt"
  }
}
```

`scheduling_policy` defaults to `interrupt` (interrupt the current response and run the new one immediately) — see the queueing and interruption section of the [Live Performance](/overview/sample-cases/live-performance) sample case. Server events then arrive in order on the control channel. For **captions**, use `response.output_text.delta` for word-by-word display and `response.output_text.done` for the final text:

```json response.output_text.delta theme={null}
{
  "event_id": "evt_text_delta_001",
  "type": "response.output_text.delta",
  "response_id": "resp_001",
  "item_id": "item_assistant_001",
  "output_index": 0,
  "delta": "Sure, "
}
```

```json response.output_text.done theme={null}
{
  "event_id": "evt_text_done_001",
  "type": "response.output_text.done",
  "response_id": "resp_001",
  "item_id": "item_assistant_001",
  "output_index": 0,
  "text": "Sure, open API Keys in the developer platform, create a key, and copy it to a secure location."
}
```

**Rendering and completion** hinge on these lifecycle events:

```json response.done theme={null}
{
  "event_id": "evt_response_done_001",
  "type": "response.done",
  "response": {
    "id": "resp_001",
    "status": "completed"
  }
}
```

* `response.render.started` / `response.render.stopped`: the avatar audio and video for this response start / stop rendering on RTC.
* `response.done`: the Response resource and its logical output stream are finished — the Vocal and Visual attached to the Response are all complete. It fires no matter which terminal state the response ends in; read the final status from `response.done.response.status`.
* In `video_avatar` mode, treat `response.done` as the signal that a round of avatar output has fully finished.

<h2 id="step-5-close-session">
  Step 5: Close the Session
</h2>

The default demo must provide an explicit End session button. If a response is still active, wait for its final response.done event, then close WSS, stop and destroy TRTC, and ask your own backend to close the session so remote resources are released. The Close Session REST call carries your VIVIX\_API\_KEY and must run on your backend — never from the browser:

```text Close session theme={null}
// Runs in the browser: tear down the local WebSocket and TRTC objects,
// then ask your own backend to close the session. The real Close Session
// call carries your VIVIX_API_KEY and must run on your backend - never
// ship the key to the browser.
async function closeAvatarSession({ session, trtc, ws }) {
  ws.close(1000, "client close");
  await trtc.stopLocalAudio().catch(() => {});
  await trtc.exitRoom().catch(() => {});
  trtc.destroy();

  // Your backend holds VIVIX_API_KEY and calls
  // POST https://api.vivix.ai/v1/realtime-avatar/sessions/{session_id}/close
  const response = await fetch("/your-backend/close-avatar-session", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ session_id: session.session_id }),
  });
  if (!response.ok) {
    throw new Error("Failed to close the session.");
  }
}
```

<h2 id="expression-and-emotion">
  Expression and Emotion: Writing Instructions
</h2>

Claire's expression and emotion are not driven by a dedicated interface — they come from how you write `avatars[].instructions`. Write instructions as "how the character speaks", not "how the system works":

* `avatars[].instructions` defines only the persona and tone — it grants no tools or business data access. Put the task goal of the current application in `conversation.instructions`, and one-off directives for a single generation in `response.instructions` (priority: avatar \< conversation \< response; see [Write spoken instructions](/streaming-avatar/character/spoken-instructions)).
* For production integrations, **explicitly constrain the output to plain spoken language**: only what the character would actually say out loud — no character names, no narration, no action descriptions, no stage directions, no Markdown, no JSON, no parenthetical asides, no emojis.
* When the user asks the avatar to perform an action, do not let the model describe the action in text — actions are driven by `response.script.visual.prompt` or session state (see the [Live Performance](/overview/sample-cases/live-performance) sample case).

Emotion is expressed through **approved audio tags**: let the model insert square-bracket tags into its lines. A tag describes vocal delivery — the emotion of the voice and non-verbal sounds — not actions. Approved tags fall into two groups:

| Category | Approved tags |
| - | - |
| direction (emotional direction) | `[happy]` `[sad]` `[excited]` `[angry]` `[annoyed]` `[appalled]` `[thoughtful]` `[surprised]` `[sarcastic]` `[curious]` `[mischievously]` `[frustrated]` `[dismissive]` `[sympathetic]` `[reassuring]` `[warmly]` `[nervously]` `[sheepishly]` `[desperately]` `[deadpan]` `[cautiously]` `[dramatically]` `[impressed]` `[delighted]` `[amazed]` |
| non-verbal (non-verbal sounds) | `[laughs]` `[laughing]` `[chuckles]` `[giggles]` `[sighs]` `[exhales]` `[exhales sharply]` `[inhales deeply]` `[whispers]` `[clears throat]` `[crying]` `[snorts]` `[short pause]` `[long pause]` |

The recommended constraint block (append it to the end of your instructions):

```text Constraint template theme={null}
Output only Claire's spoken words and approved square-bracket audio tags.
No speaker labels, no narration, no stage directions, no motion descriptions, no emojis, no parentheses.
Every response must include 1 or 2 approved non-verbal tags.
Use at most 1 approved direction tag, only when the emotion is clear.
Place tags immediately before the sentence or clause they modify, or at a natural pause point.
Do not cluster tags together. Tags describe vocal delivery only, not actions.
Never invent tags. If unsure, use a safe non-verbal tag such as [short pause] or [exhales].
```

Under these constraints, one of Claire's replies looks like this:

```json Sample reply theme={null}
[warmly] Let me make sure I understand. [short pause] What are you hoping to build with Vivix?
```

<h2 id="minimal-event-sequence">
  Minimal Runnable Event Sequence
</h2>

Putting the steps together, the default demo connects in this order:

```text Connection theme={null}
HTTP POST /v1/realtime-avatar/sessions
-> request microphone permission
-> TRTC enterRoom
-> TRTC startLocalAudio if permission is granted
-> append token=<control.client_secret> to control.url and connect WSS
```

For a typed turn:

```text Text input theme={null}
-> conversation.item.create(user message)
<- conversation.item.created
-> response.create
```

For a voice turn:

```text Voice input theme={null}
-> microphone audio over TRTC
<- input_audio.speech_started
<- conversation.item.input_audio_transcription.delta
<- input_audio.speech_stopped
<- conversation.item.input_audio_transcription.completed
<- conversation.item.created
```

Both input paths then use the same response output stream:

```html Response output theme={null}
<- response.created
<- response.render.started
<- response.output_text.delta
<- response.output_text.done
<- response.output_item.done
<- response.render.stopped
<- response.done
```

When the demo ends:

```text Cleanup theme={null}
-> WebSocket close
-> TRTC stopLocalAudio / exitRoom / destroy
HTTP POST /v1/realtime-avatar/sessions/{session_id}/close
```

Full field definitions for every event are in [Client events](/streaming-avatar/api-references/client-events) and [Server events](/streaming-avatar/api-references/server-events).

<h2 id="next-steps">
  Next Steps
</h2>

* [Live Performance](/overview/sample-cases/live-performance) — the other sample case, built around actions: use `response.script` to directly drive singing, dancing, and fixed lines.
* [Script speech and performances](/streaming-avatar/interaction/scripted-performances) — directed movement and scripted audio.
* [Change outfits and scenes](/streaming-avatar/interaction/outfits-and-scenes) — switch the avatar image during a session.
* [Add tools to the character](/streaming-avatar/character/tools) — function calls.
* [API References · Sessions](/streaming-avatar/api-references/sessions) — REST endpoints and parameter details.
* [API References · Client events](/streaming-avatar/api-references/client-events) / [Server events](/streaming-avatar/api-references/server-events) — the complete event mechanism.
* [Streaming World](/streaming-world/configuration) — realtime interactive video worlds driven by the W-series models; coming soon.
