Skip to main content
Asks the service to create a response. By default this triggers model inference; when response.script is supplied, the service renders the client-provided vocal and visual script directly instead. Avatar selection is handled by the session state, not by response.create. after_current_response allows only one pending response. Keep longer queues in your application and submit them in sequence. response.input replaces the default context for this response. [] supplies no context. Complete inline items can be sent directly; an item_reference must already exist.

Event fields

string
Client-generated id for this event.
string
required
Event type. Must be response.create.
object
Per-response creation parameters. Fields set here apply only to this response.
string
Controls where the response is added. Use auto to write to the default conversation, or none for an out-of-band response that does not write output to the default conversation.
string
Controls what happens when the session is already busy. Use interrupt to cancel the active response or render and start this response, reject_if_busy to reject this event with an error, or after_current_response to run this response immediately after the current response finishes. Defaults to interrupt. Interrupted responses still emit their existing terminal lifecycle events, such as response.render.stopped followed by response.done.
array
Input items to include in the model context when asking the model to respond. Each array item is either an item_reference or a raw conversation_item. Providing this field creates a response-specific context instead of using the default conversation; an empty array clears context for this response. Invalid when script is supplied.

Variant: item_reference

References an existing item already present in the session.
string
required
Use item_reference.
string
required
Existing item id to include in the response context.

Variant: conversation_item

A raw conversation item, using the same message, function call, and function call output shapes documented for conversation.item.create.
string
Instructions for this response only. These steer model behavior; they are invalid when script is supplied. When these conflict with conversation or selected avatar instructions, these instructions take precedence for this response.
integer | string
Maximum output tokens for this response, inclusive of tool calls. Use an integer or inf. Invalid when script is supplied.
array
Tools available for this response. When supplied, this overrides session conversation tools for this response only. Invalid when script is supplied.
response.tools[].function tool Function tool with type, name, optional description, and optional JSON Schema parameters. response.tools[].mcp tool MCP tool support is planned and not yet available in this version. This version documents the function tool only.
object
Direct avatar output to render exactly as provided. When supplied, no LLM reply generation or tool calling occurs, and input, instructions, tools, and max_output_tokens are invalid.
object
Exact vocal output for the active avatar. Provide vocal, visual, or both.

Variant: speech

Synthesized speech from exact text.
string
required
Use speech.
string
required
Exact spoken text. This text also streams through response.output_text.delta and response.output_text.done.

Variant: speech_asset

Synthesized speech from a reusable speech text asset.
string
required
Use speech_asset.
string
required
Previously registered asset from assets.speech_text. The asset text is exact spoken text and streams through response.output_text.delta and response.output_text.done.

Variant: audio

Supplied audio used to drive the avatar.
string
required
Use audio.
string
required
Required MIME type for the supplied audio. Supported values are audio/wav and audio/mpeg. The bytes served by url or encoded in audio must match this value.
string
HTTPS URL for an audio file to render through the avatar. The URL must be directly fetchable by Vivix without custom request headers. Provide exactly one of url or audio.
string
Base64-encoded audio file bytes to render through the avatar. Do not include a data URL prefix. Provide exactly one of url or audio.

Variant: audio_asset

Reusable audio asset used to drive the avatar.
string
required
Use audio_asset.
string
required
Previously registered asset from assets.audio.
object
Visual instruction for the active avatar. Used only in video_avatar mode.
string
Inline instruction for avatar movement, posture, gaze, gestures, or presentation. Separate multiple ordered action descriptions with [SHOT_SEP].
string
Previously registered asset from assets.visual_prompts. Provide exactly one of prompt or visual_prompt_asset_id.
integer
Target duration of each visual action segment in milliseconds. Defaults to 5000 when omitted. When prompt contains [SHOT_SEP], this value applies to each segment, not the combined sequence. For example, three segments at 5000 milliseconds each plan about 15 seconds of motion. It does not change the duration of supplied audio or guarantee exact browser playback timing.
Script behavior. Response scheduling. The session is busy until the active response reaches terminal response.done. If rendering starts, response.render.stopped is emitted first and response.done follows.
When busy
required
Cancels active output, clears any pending after_current_response, and starts the new response. Interrupted responses still emit terminal lifecycle events. · When idle: Starts immediately. This is the default.
When busy
required
Rejects with an error if the session is active, rendering, or has a pending response. · When idle: Starts immediately.
When busy
required
Accepts at most one pending response. A second pending request is rejected with an error. If the session is busy, the response is queued until the current response finishes playing (in video_avatar mode, until it reaches terminal response.done). · When idle: Starts immediately.
Pending after_current_response requests snapshot effective response inputs at submit time, including mode, active_avatar_id, source_image_id, selected avatar defaults and instructions, conversation.instructions, response fields, and referenced asset contents. Later session.update, avatar.add, or asset.add events do not silently change the pending response. Pending responses do not emit response.created until they start.