Skip to main content
Client events are JSON messages you send on the Streaming Avatar WSS connection. Connect to the control.url returned by Create Session (it points to /v1/realtime-avatar/control) and authenticate with control.client_secret when connecting: browsers append it to the returned URL as the token query parameter, and server-side clients may instead use a Bearer header. Do not open the original URL without the token, and do not name the query parameter client_secret. See Display the character on a webpage for a complete connection example. Each event carries a type field. The optional event_id field shown in each event schema (up to 512 characters) is recommended on every event so you can correlate server errors and acknowledgements with the client event that triggered them. Forward compatibility. Vivix may add new fields over time; ignore any field you do not recognize. Unknown client events are rejected with an error event.Registryasset.addavatar.addSessionsession.updateConversationconversation.item.createconversation.item.retrieveconversation.item.deleteResponseresponse.createresponse.cancel

1. Session Events

session.update

Updates the mode, active avatar, or registered source image. Omitted fields stay unchanged. Only the session state listed below is mutable; this does not update voice, tools, or prompts. session.updated acknowledges acceptance. For visual changes, also listen for session.update.done. event_id: optional string Client-generated id for this event. type: “session.update” Event type. Must be session.update. session: object Partial session state to update. session.mode: optional string enum New session mode. Supported values are text_chat and video_avatar. Switching to text_chat clears the effective source_image_id to null. Switching to video_avatar requires a renderable active avatar; if source_image_id is omitted, Vivix uses the current source image or derives one from the active avatar’s visual.default_source_image_id. session.active_avatar_id: optional string Avatar whose instructions should apply by default for later responses. In video_avatar mode, this is also the avatar rendered in the video stream. session.source_image_id: optional nullable string Source image to use for avatar video. Used only in video_avatar mode. Use null to clear the selected source image; the effective value is also null in text_chat mode. Requests fail in video_avatar mode if no usable source image can be resolved. session.update

2. Registry Events

asset.add

Registers additional reusable script assets while the session is live. Assets are additive: existing asset_id values cannot be changed or replaced in v1. If any asset in the event is invalid, none of the assets in that event are registered. Wait for asset.added before referencing the new ids from response.script. event_id: optional string Client-generated id for this event. type: “asset.add” Event type. Must be asset.add. assets: object Assets to register for later scripted responses. Provide at least one asset across images, audio, speech_text, or visual_prompts. assets.images: optional array Reusable images that can be registered to the underlying media session. assets.images[].asset_id: string Caller-defined image asset id, unique within the session. assets.images[].url: string HTTPS URL for the image file. assets.images[].mime_type: optional string Image MIME type, such as image/png or image/jpeg. assets.audio: optional array Reusable audio files for response.script.vocal.type: "audio_asset". assets.audio[].asset_id: string Caller-defined audio asset id, unique within the session. assets.audio[].url: string HTTPS URL for the audio file. The URL must be directly fetchable by Vivix without custom request headers. assets.audio[].mime_type: optional string Audio MIME type. Supported values are audio/wav and audio/mpeg. assets.speech_text: optional array Reusable exact speech text for response.script.vocal.type: "speech_asset". assets.speech_text[].asset_id: string Caller-defined speech text asset id, unique within the session. assets.speech_text[].text: string Exact text to synthesize. When used in response.script, it also streams through output text events. assets.speech_text[].format: optional string Text format, such as text or ssml. assets.visual_prompts: optional array Reusable visual instructions for response.script.visual.visual_prompt_asset_id. assets.visual_prompts[].asset_id: string Caller-defined visual prompt asset id, unique within the session. assets.visual_prompts[].text: string Reusable visual instruction text describing avatar movement, posture, gaze, gestures, or presentation. assets.visual_prompts[].format: optional string Text format, such as text. asset.add

avatar.add

avatar.add registers new avatar IDs. It cannot replace an existing avatar or append source images to one. Wait for avatar.added before using a new avatar. Registers additional avatars while the session is live. Avatars are additive: existing avatar_id values cannot be changed or replaced in v1. Added avatars do not become active automatically; after avatar.added, use session.update to select one. If any avatar in the event is invalid, none of the avatars in that event are registered. event_id: optional string Client-generated id for this event. type: “avatar.add” Event type. Must be avatar.add. avatars: array Additional avatar definitions to register. Uses the same avatar object shape as avatars[] in Create Session. avatars[].avatar_id: string Caller-defined avatar id, unique within the session. avatars[].instructions: optional string The avatar’s identity, speaking style, and response rules. avatars[].voice: optional object Optional per-avatar TTS configuration for sessions that omit pipeline_config.tts_config and require different voices for different avatars. Provide provider, provider_voice_id, and model. Otherwise omit this object; the added avatar uses the session TTS configuration. avatars[].voice.provider: string TTS provider, such as qwen-audio-3.0-tts-flash_ws, elevenlabs, or qwen3tts. Required with provider_voice_id and model. avatars[].voice.provider_voice_id: string Provider-side voice or style id. Required with provider and model. avatars[].voice.model: string Provider TTS model id. Required with provider and provider_voice_id. avatars[].voice.speed: optional number Speech speed multiplier. Valid range is 0.7 to 1.2. avatars[].visual: optional object Visual rendering settings for avatar video. Required when this avatar can be selected in video_avatar mode. avatars[].visual.instructions: optional string Default visual instructions for this avatar, describing movement, posture, gaze, gestures, and on-camera presentation. avatars[].visual.source_images: optional array Registered source images available for this avatar. Required when image-based avatar video rendering is used. avatars[].visual.source_images[].source_image_id: string Caller-defined source image id, unique within the avatar. avatars[].visual.source_images[].url: string Image URL for this avatar video source image. avatars[].visual.source_images[].description: optional string Optional description of the person, pose, and scene. avatars[].visual.source_images[].media_type: required string Image media type, such as image/png or image/jpeg. avatars[].visual.default_source_image_id: optional string With one usable source image, this field can be omitted. With multiple images and no explicit source_image_id, use it to select the initial image. avatar.add

3. Conversation Events

conversation.item.create

conversation.item.create → response.create: wait for the matching conversation.item.created before requesting a reply that uses the new message. Use the same sequence when returning a tool result. Adds a conversation item to the session. Use it to provide user messages, function calls, or function call outputs that should become part of the conversation context. The server confirms with conversation.item.created. event_id: optional string Client-generated id for this event. type: “conversation.item.create” Event type. Must be conversation.item.create. previous_item_id: optional string Item after which the new item should be inserted. If omitted, the item is appended. Use root to insert at the beginning of the conversation. item: object A single item within the realtime conversation. Provide one of the variants below. id: optional string Server-assigned or client-provided conversation item id. type: string enum One of message, function_call, or function_call_output. status: optional string enum One of completed, incomplete, or in_progress. Usually returned by the server. Message variant Message item with item.type: "message". Message variant.role: string enum One of system, user, or assistant. System messages provide additional conversation context or instructions. For assistant messages, generated audio and video are delivered over RTC media tracks, and the text for spoken output streams through response.output_text.* events. Message variant.content: array Message content parts. System messages typically use input_text; user messages typically use input_text, input_audio, and input_image; assistant messages typically use text parts. Message variant.content[].type: string enum One of input_text, input_audio, input_image, or text. Message variant.content[].text: optional string Text content. Used with input_text and text. Message variant.content[].audio: optional string Base64-encoded audio. Used with input_audio. Message variant.content[].transcript: optional string Optional transcript. Used with input_audio. Message variant.content[].image_url: optional string Image URL. Used with input_image. Message variant.content[].detail: optional string enum Image detail level of auto, low, or high. Used with input_image. Function call variant Function call item with item.type: "function_call". Function call variant.call_id: optional string Function call id. Function call variant.name: string Function name. Function call variant.arguments: string JSON-encoded function arguments. Function call output variant Function call output item with item.type: "function_call_output". Function call output variant.call_id: string Function call id this output belongs to. Function call output variant.output: string Function output as a string. conversation.item.create

conversation.item.retrieve

Requests a conversation item by id. The server responds with conversation.item.retrieved if the item exists. event_id: optional string Client-generated id for this event. type: “conversation.item.retrieve” Event type. Must be conversation.item.retrieve. item_id: string Identifier of the conversation item to retrieve. conversation.item.retrieve

conversation.item.delete

Removes a conversation item from the session context. The server confirms with conversation.item.deleted. event_id: optional string Client-generated id for this event. type: “conversation.item.delete” Event type. Must be conversation.item.delete. item_id: string Identifier of the conversation item to delete. conversation.item.delete

4. Response Events

response.create

after_current_response allows only one pending response. Keep longer queues in your application and submit them in sequence. response.input replaces the default context for this response. [] supplies no context. Complete inline items can be sent directly; an item_reference must already exist. Asks the service to create a response. By default this triggers model inference; when response.script is supplied, the service renders the client-provided vocal and visual script directly instead. Avatar selection is handled by the session state, not by response.create. event_id: optional string Client-generated id for this event. type: “response.create” Event type. Must be response.create. response: optional object Per-response creation parameters. Fields set here apply only to this response. response.conversation: optional string enum Controls where the response is added. Use auto to write to the default conversation, or none for an out-of-band response that does not write output to the default conversation. response.scheduling_policy: optional string enum Controls what happens when the session is already busy. Use interrupt to cancel the active response or render and start this response, reject_if_busy to reject this event with an error, or after_current_response to run this response immediately after the current response finishes. Defaults to interrupt. Interrupted responses still emit their existing terminal lifecycle events, such as response.render.stopped followed by response.done. response.input: optional array Input items to include in the model context when asking the model to respond. Each array item is either an item_reference or a raw conversation_item. Providing this field creates a response-specific context instead of using the default conversation; an empty array clears context for this response. Invalid when script is supplied. Variant: item_reference References an existing item already present in the session. response.input[].type: string Use item_reference. response.input[].id: string Existing item id to include in the response context. Variant: conversation_item A raw conversation item, using the same message, function call, and function call output shapes documented for conversation.item.create. response.instructions: optional string Instructions for this response only. These steer model behavior; they are invalid when script is supplied. When these conflict with conversation or selected avatar instructions, these instructions take precedence for this response. response.max_output_tokens: optional integer or string Maximum output tokens for this response, inclusive of tool calls. Use an integer or inf. Invalid when script is supplied. response.tools: optional array Tools available for this response. When supplied, this overrides session conversation tools for this response only. Invalid when script is supplied. response.tools[].function tool Function tool with type, name, optional description, and optional JSON Schema parameters. response.tools[].mcp tool MCP tool support is planned and not yet available in this version. This version documents the function tool only. response.script: optional object Direct avatar output to render exactly as provided. When supplied, no LLM reply generation or tool calling occurs, and input, instructions, tools, and max_output_tokens are invalid. response.script.vocal: optional object Exact vocal output for the active avatar. Provide vocal, visual, or both. Variant: speech Synthesized speech from exact text. response.script.vocal.type: “speech” Use speech. response.script.vocal.text: string Exact spoken text. This text also streams through response.output_text.delta and response.output_text.done. Variant: speech_asset Synthesized speech from a reusable speech text asset. response.script.vocal.type: “speech_asset” Use speech_asset. response.script.vocal.speech_text_asset_id: string Previously registered asset from assets.speech_text. The asset text is exact spoken text and streams through response.output_text.delta and response.output_text.done. Variant: audio Supplied audio used to drive the avatar. response.script.vocal.type: “audio” Use audio. response.script.vocal.mime_type: string Required MIME type for the supplied audio. Supported values are audio/wav and audio/mpeg. The bytes served by url or encoded in audio must match this value. response.script.vocal.url: optional string HTTPS URL for an audio file to render through the avatar. The URL must be directly fetchable by Vivix without custom request headers. Provide exactly one of url or audio. response.script.vocal.audio: optional string Base64-encoded audio file bytes to render through the avatar. Do not include a data URL prefix. Provide exactly one of url or audio. Variant: audio_asset Reusable audio asset used to drive the avatar. response.script.vocal.type: “audio_asset” Use audio_asset. response.script.vocal.audio_asset_id: string Previously registered asset from assets.audio. response.script.visual: optional object Visual instruction for the active avatar. Used only in video_avatar mode. response.script.visual.prompt: optional string Inline instruction for avatar movement, posture, gaze, gestures, or presentation. Separate multiple ordered action descriptions with [SHOT_SEP]. response.script.visual.visual_prompt_asset_id: optional string Previously registered asset from assets.visual_prompts. Provide exactly one of prompt or visual_prompt_asset_id. response.script.visual.duration_ms: optional integer Target duration of each visual action segment in milliseconds. Defaults to 5000 when omitted. When prompt contains [SHOT_SEP], this value applies to each segment, not the combined sequence. For example, three segments at 5000 milliseconds each plan about 15 seconds of motion. It does not change the duration of supplied audio or guarantee exact browser playback timing. Script behavior. Model bypass · Behavior: script renders direct avatar output. No LLM reply generation or tool calling occurs, and exact vocal text is not rewritten by model or avatar instructions. video_avatar mode · Behavior: Vocal output drives speech or supplied audio, and visual text guides avatar motion over RTC media using the current active avatar, voice, source image, and visual defaults. Conversation writing · Behavior: If conversation is auto or omitted, inline speech text or speech text asset content is written as an assistant message. If conversation is none, it is not. Lower-level orchestration · Behavior: Lower-level explicit media orchestration, such as explicit shot queues, track item ids, segment arrays, delete and clear operations, timed audio gaps, and asset-backed media composition, is not publicly available yet. Response scheduling. The session is busy until the active response reaches terminal response.done. If rendering starts, response.render.stopped is emitted first and response.done follows. interrupt · When busy: Cancels active output, clears any pending after_current_response, and starts the new response. Interrupted responses still emit terminal lifecycle events. · When idle: Starts immediately. This is the default. reject_if_busy · When busy: Rejects with an error if the session is active, rendering, or has a pending response. · When idle: Starts immediately. after_current_response · When busy: Accepts at most one pending response. A second pending request is rejected with an error. If the session is busy, the response is queued until the current response finishes playing (in video_avatar mode, until it reaches terminal response.done). · When idle: Starts immediately. Pending after_current_response requests snapshot effective response inputs at submit time, including mode, active_avatar_id, source_image_id, selected avatar defaults and instructions, conversation.instructions, response fields, and referenced asset contents. Later session.update, avatar.add, or asset.add events do not silently change the pending response. Pending responses do not emit response.created until they start. Default response.create
Reject if busy
After current response
Out-of-band response
Scripted speech and visual asset
Scripted audio asset and visual motion
Scripted visual-only action
Out-of-band scripted speech
Invalid scripted response (script with instructions and input is rejected)

response.cancel

response.cancel targets the current active response. Omit response_id for the active response, or supply its matching ID. Successful cancellation also clears pending work. No active response or a mismatched ID produces an error; it cannot remove only a pending response. Cancels an in-progress response. The server responds with response.done and a cancelled response status. If there is no response to cancel, the server responds with an error. event_id: optional string Client-generated id for this event. type: “response.cancel” Event type. Must be response.cancel. response_id: optional string When provided, this must match the active response ID. Omit it to cancel the active response. Use the ID from response.created.response.id. response.cancel

Source image media type

avatars[].visual.source_images[].media_type is required and must start with image/, for example image/png or image/jpeg. Use the actual image format, not the video output format.

Continue reading

Character instructions · Speech recognition · Text model · Text to speech · Opening acknowledgements