control.url returned by Create Session (it points to /v1/realtime-avatar/control) and authenticate with control.client_secret when connecting: browsers append it to the returned URL as the token query parameter, and server-side clients may instead use a Bearer header. Do not open the original URL without the token, and do not name the query parameter client_secret. See Display the character on a webpage for a complete connection example. Each event carries a type field. The optional event_id field shown in each event schema (up to 512 characters) is recommended on every event so you can correlate server errors and acknowledgements with the client event that triggered them.
Forward compatibility. Vivix may add new fields over time; ignore any field you do not recognize. Unknown client events are rejected with an error event.Registryasset.addavatar.addSessionsession.updateConversationconversation.item.createconversation.item.retrieveconversation.item.deleteResponseresponse.createresponse.cancel
1. Session Events
session.update
Updates the mode, active avatar, or registered source image. Omitted fields stay unchanged. Only the session state listed below is mutable; this does not update voice, tools, or prompts. session.updated acknowledges acceptance. For visual changes, also listen for session.update.done.
event_id: optional string
Client-generated id for this event.
type: “session.update”
Event type. Must be session.update.
session: object
Partial session state to update.
session.mode: optional string enum
New session mode. Supported values are text_chat and video_avatar. Switching to text_chat clears the effective source_image_id to null. Switching to video_avatar requires a renderable active avatar; if source_image_id is omitted, Vivix uses the current source image or derives one from the active avatar’s visual.default_source_image_id.
session.active_avatar_id: optional string
Avatar whose instructions should apply by default for later responses. In video_avatar mode, this is also the avatar rendered in the video stream.
session.source_image_id: optional nullable string
Source image to use for avatar video. Used only in video_avatar mode. Use null to clear the selected source image; the effective value is also null in text_chat mode. Requests fail in video_avatar mode if no usable source image can be resolved.
session.update
2. Registry Events
asset.add
Registers additional reusable script assets while the session is live. Assets are additive: existing asset_id values cannot be changed or replaced in v1. If any asset in the event is invalid, none of the assets in that event are registered. Wait for asset.added before referencing the new ids from response.script.
event_id: optional string
Client-generated id for this event.
type: “asset.add”
Event type. Must be asset.add.
assets: object
Assets to register for later scripted responses. Provide at least one asset across images, audio, speech_text, or visual_prompts.
assets.images: optional array
Reusable images that can be registered to the underlying media session.
assets.images[].asset_id: string
Caller-defined image asset id, unique within the session.
assets.images[].url: string
HTTPS URL for the image file.
assets.images[].mime_type: optional string
Image MIME type, such as image/png or image/jpeg.
assets.audio: optional array
Reusable audio files for response.script.vocal.type: "audio_asset".
assets.audio[].asset_id: string
Caller-defined audio asset id, unique within the session.
assets.audio[].url: string
HTTPS URL for the audio file. The URL must be directly fetchable by Vivix without custom request headers.
assets.audio[].mime_type: optional string
Audio MIME type. Supported values are audio/wav and audio/mpeg.
assets.speech_text: optional array
Reusable exact speech text for response.script.vocal.type: "speech_asset".
assets.speech_text[].asset_id: string
Caller-defined speech text asset id, unique within the session.
assets.speech_text[].text: string
Exact text to synthesize. When used in response.script, it also streams through output text events.
assets.speech_text[].format: optional string
Text format, such as text or ssml.
assets.visual_prompts: optional array
Reusable visual instructions for response.script.visual.visual_prompt_asset_id.
assets.visual_prompts[].asset_id: string
Caller-defined visual prompt asset id, unique within the session.
assets.visual_prompts[].text: string
Reusable visual instruction text describing avatar movement, posture, gaze, gestures, or presentation.
assets.visual_prompts[].format: optional string
Text format, such as text.
asset.add
avatar.add
avatar.add registers new avatar IDs. It cannot replace an existing avatar or append source images to one. Wait for avatar.added before using a new avatar.
Registers additional avatars while the session is live. Avatars are additive: existing avatar_id values cannot be changed or replaced in v1. Added avatars do not become active automatically; after avatar.added, use session.update to select one. If any avatar in the event is invalid, none of the avatars in that event are registered.
event_id: optional string
Client-generated id for this event.
type: “avatar.add”
Event type. Must be avatar.add.
avatars: array
Additional avatar definitions to register. Uses the same avatar object shape as avatars[] in Create Session.
avatars[].avatar_id: string
Caller-defined avatar id, unique within the session.
avatars[].instructions: optional string
The avatar’s identity, speaking style, and response rules.
avatars[].voice: optional object
Optional per-avatar TTS configuration for sessions that omit pipeline_config.tts_config and require different voices for different avatars. Provide provider, provider_voice_id, and model. Otherwise omit this object; the added avatar uses the session TTS configuration.
avatars[].voice.provider: string
TTS provider, such as qwen-audio-3.0-tts-flash_ws, elevenlabs, or qwen3tts. Required with provider_voice_id and model.
avatars[].voice.provider_voice_id: string
Provider-side voice or style id. Required with provider and model.
avatars[].voice.model: string
Provider TTS model id. Required with provider and provider_voice_id.
avatars[].voice.speed: optional number
Speech speed multiplier. Valid range is 0.7 to 1.2.
avatars[].visual: optional object
Visual rendering settings for avatar video. Required when this avatar can be selected in video_avatar mode.
avatars[].visual.instructions: optional string
Default visual instructions for this avatar, describing movement, posture, gaze, gestures, and on-camera presentation.
avatars[].visual.source_images: optional array
Registered source images available for this avatar. Required when image-based avatar video rendering is used.
avatars[].visual.source_images[].source_image_id: string
Caller-defined source image id, unique within the avatar.
avatars[].visual.source_images[].url: string
Image URL for this avatar video source image.
avatars[].visual.source_images[].description: optional string
Optional description of the person, pose, and scene.
avatars[].visual.source_images[].media_type: required string
Image media type, such as image/png or image/jpeg.
avatars[].visual.default_source_image_id: optional string
With one usable source image, this field can be omitted. With multiple images and no explicit source_image_id, use it to select the initial image.
avatar.add
3. Conversation Events
conversation.item.create
conversation.item.create → response.create: wait for the matching conversation.item.created before requesting a reply that uses the new message. Use the same sequence when returning a tool result.
Adds a conversation item to the session. Use it to provide user messages, function calls, or function call outputs that should become part of the conversation context. The server confirms with conversation.item.created.
event_id: optional string
Client-generated id for this event.
type: “conversation.item.create”
Event type. Must be conversation.item.create.
previous_item_id: optional string
Item after which the new item should be inserted. If omitted, the item is appended. Use root to insert at the beginning of the conversation.
item: object
A single item within the realtime conversation. Provide one of the variants below.
id: optional string
Server-assigned or client-provided conversation item id.
type: string enum
One of message, function_call, or function_call_output.
status: optional string enum
One of completed, incomplete, or in_progress. Usually returned by the server.
Message variant
Message item with item.type: "message".
Message variant.role: string enum
One of system, user, or assistant. System messages provide additional conversation context or instructions. For assistant messages, generated audio and video are delivered over RTC media tracks, and the text for spoken output streams through response.output_text.* events.
Message variant.content: array
Message content parts. System messages typically use input_text; user messages typically use input_text, input_audio, and input_image; assistant messages typically use text parts.
Message variant.content[].type: string enum
One of input_text, input_audio, input_image, or text.
Message variant.content[].text: optional string
Text content. Used with input_text and text.
Message variant.content[].audio: optional string
Base64-encoded audio. Used with input_audio.
Message variant.content[].transcript: optional string
Optional transcript. Used with input_audio.
Message variant.content[].image_url: optional string
Image URL. Used with input_image.
Message variant.content[].detail: optional string enum
Image detail level of auto, low, or high. Used with input_image.
Function call variant
Function call item with item.type: "function_call".
Function call variant.call_id: optional string
Function call id.
Function call variant.name: string
Function name.
Function call variant.arguments: string
JSON-encoded function arguments.
Function call output variant
Function call output item with item.type: "function_call_output".
Function call output variant.call_id: string
Function call id this output belongs to.
Function call output variant.output: string
Function output as a string.
conversation.item.create
conversation.item.retrieve
Requests a conversation item by id. The server responds with conversation.item.retrieved if the item exists.
event_id: optional string
Client-generated id for this event.
type: “conversation.item.retrieve”
Event type. Must be conversation.item.retrieve.
item_id: string
Identifier of the conversation item to retrieve.
conversation.item.retrieve
conversation.item.delete
Removes a conversation item from the session context. The server confirms with conversation.item.deleted.
event_id: optional string
Client-generated id for this event.
type: “conversation.item.delete”
Event type. Must be conversation.item.delete.
item_id: string
Identifier of the conversation item to delete.
conversation.item.delete
4. Response Events
response.create
after_current_response allows only one pending response. Keep longer queues in your application and submit them in sequence.
response.input replaces the default context for this response. [] supplies no context. Complete inline items can be sent directly; an item_reference must already exist.
Asks the service to create a response. By default this triggers model inference; when response.script is supplied, the service renders the client-provided vocal and visual script directly instead. Avatar selection is handled by the session state, not by response.create.
event_id: optional string
Client-generated id for this event.
type: “response.create”
Event type. Must be response.create.
response: optional object
Per-response creation parameters. Fields set here apply only to this response.
response.conversation: optional string enum
Controls where the response is added. Use auto to write to the default conversation, or none for an out-of-band response that does not write output to the default conversation.
response.scheduling_policy: optional string enum
Controls what happens when the session is already busy. Use interrupt to cancel the active response or render and start this response, reject_if_busy to reject this event with an error, or after_current_response to run this response immediately after the current response finishes. Defaults to interrupt. Interrupted responses still emit their existing terminal lifecycle events, such as response.render.stopped followed by response.done.
response.input: optional array
Input items to include in the model context when asking the model to respond. Each array item is either an item_reference or a raw conversation_item. Providing this field creates a response-specific context instead of using the default conversation; an empty array clears context for this response. Invalid when script is supplied.
Variant: item_reference
References an existing item already present in the session.
response.input[].type: string
Use item_reference.
response.input[].id: string
Existing item id to include in the response context.
Variant: conversation_item
A raw conversation item, using the same message, function call, and function call output shapes documented for conversation.item.create.
response.instructions: optional string
Instructions for this response only. These steer model behavior; they are invalid when script is supplied. When these conflict with conversation or selected avatar instructions, these instructions take precedence for this response.
response.max_output_tokens: optional integer or string
Maximum output tokens for this response, inclusive of tool calls. Use an integer or inf. Invalid when script is supplied.
response.tools: optional array
Tools available for this response. When supplied, this overrides session conversation tools for this response only. Invalid when script is supplied.
response.tools[].function tool
Function tool with type, name, optional description, and optional JSON Schema parameters.
response.tools[].mcp tool
MCP tool support is planned and not yet available in this version. This version documents the function tool only.
response.script: optional object
Direct avatar output to render exactly as provided. When supplied, no LLM reply generation or tool calling occurs, and input, instructions, tools, and max_output_tokens are invalid.
response.script.vocal: optional object
Exact vocal output for the active avatar. Provide vocal, visual, or both.
Variant: speech
Synthesized speech from exact text.
response.script.vocal.type: “speech”
Use speech.
response.script.vocal.text: string
Exact spoken text. This text also streams through response.output_text.delta and response.output_text.done.
Variant: speech_asset
Synthesized speech from a reusable speech text asset.
response.script.vocal.type: “speech_asset”
Use speech_asset.
response.script.vocal.speech_text_asset_id: string
Previously registered asset from assets.speech_text. The asset text is exact spoken text and streams through response.output_text.delta and response.output_text.done.
Variant: audio
Supplied audio used to drive the avatar.
response.script.vocal.type: “audio”
Use audio.
response.script.vocal.mime_type: string
Required MIME type for the supplied audio. Supported values are audio/wav and audio/mpeg. The bytes served by url or encoded in audio must match this value.
response.script.vocal.url: optional string
HTTPS URL for an audio file to render through the avatar. The URL must be directly fetchable by Vivix without custom request headers. Provide exactly one of url or audio.
response.script.vocal.audio: optional string
Base64-encoded audio file bytes to render through the avatar. Do not include a data URL prefix. Provide exactly one of url or audio.
Variant: audio_asset
Reusable audio asset used to drive the avatar.
response.script.vocal.type: “audio_asset”
Use audio_asset.
response.script.vocal.audio_asset_id: string
Previously registered asset from assets.audio.
response.script.visual: optional object
Visual instruction for the active avatar. Used only in video_avatar mode.
response.script.visual.prompt: optional string
Inline instruction for avatar movement, posture, gaze, gestures, or presentation. Separate multiple ordered action descriptions with [SHOT_SEP].
response.script.visual.visual_prompt_asset_id: optional string
Previously registered asset from assets.visual_prompts. Provide exactly one of prompt or visual_prompt_asset_id.
response.script.visual.duration_ms: optional integer
Target duration of each visual action segment in milliseconds. Defaults to 5000 when omitted. When prompt contains [SHOT_SEP], this value applies to each segment, not the combined sequence. For example, three segments at 5000 milliseconds each plan about 15 seconds of motion. It does not change the duration of supplied audio or guarantee exact browser playback timing.
Script behavior.
Model bypass · Behavior: script renders direct avatar output. No LLM reply generation or tool calling occurs, and exact vocal text is not rewritten by model or avatar instructions.
video_avatar mode · Behavior: Vocal output drives speech or supplied audio, and visual text guides avatar motion over RTC media using the current active avatar, voice, source image, and visual defaults.
Conversation writing · Behavior: If conversation is auto or omitted, inline speech text or speech text asset content is written as an assistant message. If conversation is none, it is not.
Lower-level orchestration · Behavior: Lower-level explicit media orchestration, such as explicit shot queues, track item ids, segment arrays, delete and clear operations, timed audio gaps, and asset-backed media composition, is not publicly available yet.
Response scheduling.
The session is busy until the active response reaches terminal response.done. If rendering starts, response.render.stopped is emitted first and response.done follows.
interrupt · When busy: Cancels active output, clears any pending after_current_response, and starts the new response. Interrupted responses still emit terminal lifecycle events. · When idle: Starts immediately. This is the default.
reject_if_busy · When busy: Rejects with an error if the session is active, rendering, or has a pending response. · When idle: Starts immediately.
after_current_response · When busy: Accepts at most one pending response. A second pending request is rejected with an error. If the session is busy, the response is queued until the current response finishes playing (in video_avatar mode, until it reaches terminal response.done). · When idle: Starts immediately.
Pending after_current_response requests snapshot effective response inputs at submit time, including mode, active_avatar_id, source_image_id, selected avatar defaults and instructions, conversation.instructions, response fields, and referenced asset contents. Later session.update, avatar.add, or asset.add events do not silently change the pending response. Pending responses do not emit response.created until they start.
Default response.create
response.cancel
response.cancel targets the current active response. Omit response_id for the active response, or supply its matching ID. Successful cancellation also clears pending work. No active response or a mismatched ID produces an error; it cannot remove only a pending response.
Cancels an in-progress response. The server responds with response.done and a cancelled response status. If there is no response to cancel, the server responds with an error.
event_id: optional string
Client-generated id for this event.
type: “response.cancel”
Event type. Must be response.cancel.
response_id: optional string
When provided, this must match the active response ID. Omit it to cancel the active response. Use the ID from response.created.response.id.
response.cancel
Source image media type
avatars[].visual.source_images[].media_type is required and must start with image/, for example image/png or image/jpeg. Use the actual image format, not the video output format.