- Create a session with
POST /v1/realtime-avatar/sessions, or start from a saved character withPOST /v1/realtime-avatar/character-sessions. The full-body request defines the initial mutable session state, output settings, conversation defaults, available avatars, media transport, and the maximum session duration. - The server validates the mode, avatars, active avatar, and source image, provisions media when required, and returns
session_id, control channel connection details, expiry, the effective output and session settings, and media connection details. - Connect to
control.urland sendcontrol.client_secretduring the handshake: browsers append it to the returned URL as the token query parameter, and server-side clients may instead use Authorization: Bearer CONTROL_CLIENT_SECRET. After the connection opens, send control events such asasset.add,avatar.add,conversation.item.create,response.create,response.cancel, andsession.update. - Close the session with
POST /v1/realtime-avatar/sessions/{session_id}/closewhen the client is done. - Poll
GET /v1/realtime-avatar/sessions/{session_id}when you need to confirm cleanup. Stop onclosedor handlefailed; use a bounded timeout. - To inspect or reclaim concurrency, list the workspace’s open sessions with
GET /v1/realtime-avatar/sessions, then close any that already ended but still hold quota.
{code, message, data}: code = 0 means success and any non-zero code means failure. Field descriptions on this page refer to fields inside data. JSON fields use snake_case, and 64-bit integers may be returned as strings.
1. Create Session from Character
Creates a Streaming Avatar session from a saved character in the current workspace. Sendcharacter_id in the JSON body, not in the URL path. The response matches Create Session.
POST /v1/realtime-avatar/character-sessions
character_id is a required string. Optional fields are idempotency_key, model, max_duration_seconds, recording_mode, output, and delivery. output and delivery are shallow-merged with the saved top-level objects. Do not send avatars in the same request.
Save changes to pipeline_config, conversation, or auto_close in the Character before starting a new session. They are not overrides on this endpoint.
Request
2. Create Session
Body Parameters
recording_mode · optional string: on or off. Send off to disable recording.
model: string
Use vivix-a1-stream. A lightweight variant vivix-a1-stream-lite is also available.
idempotency_key: optional string
Client-generated key used to prevent duplicate session allocation when a create request is retried. On a 429/5xx/timeout, retry the create call with the same idempotency_key; it is safe and will not allocate a duplicate session.
region: optional string
Optional deployment region hint. Returned back on the session object; may be empty if not provided. It does not currently affect scheduling or guarantee data residency.
session: optional object
Initial mutable session state for mode, active avatar, and optional video source image. Use session.update to change these values after the session starts.
session.mode: optional string enum
Initial session mode. Supported values are text_chat and video_avatar. Defaults to video_avatar. text_chat returns text only, and video_avatar returns avatar video with generated speech and text events.
session.active_avatar_id: optional string
Avatar whose instructions are active when the session starts. In video_avatar mode, this is also the avatar rendered in the video stream. If omitted and exactly one avatar is provided, that avatar is used; if omitted with multiple avatars, the request fails.
session.source_image_id: optional string
Source image to use for avatar video. Used only when mode is video_avatar. If omitted in video_avatar, uses the active avatar’s visual.default_source_image_id. If no default is set and that avatar has exactly one source image, it is selected automatically. In video mode, the request fails if no usable image can be resolved. In text_chat, the field is omitted from the REST response.
pipeline_config: optional object
Session-scoped pipeline configuration. For normal generated speech, provide tts_config here with a non-empty tts_voice_id; tts_provider and tts_model_id may be omitted and then default to the Qwen Audio pair. It applies to every avatar and cannot be changed after session creation. See Speech recognition configuration.
pipeline_config.tts_config: object
Session-wide TTS configuration. It takes precedence over every avatars[].voice.
output: object
Create-time avatar video output settings for the session. Required for video_avatar mode and applied when session.mode is video_avatar; ignored in text_chat mode.
output.aspect_ratio: optional string
Requested avatar video aspect ratio, such as 9:16, 1:1, or 16:9.
output.resolution: optional string
Requested avatar video resolution. Supported values are 480p and 720p.
assets: optional object
Session-scoped reusable assets for response.script. Asset ids must be unique across audio, speech_text, and visual_prompts. Avatar source images stay under avatars[].visual.source_images. On the REST side the visual prompt text field is prompt; on the WSS asset.add side the same asset uses text — the two shapes differ.
assets.audio: optional array
Reusable audio files that can be referenced by scripted vocal output.
assets.audio[]: object
One reusable audio asset.
assets.audio[].asset_id: string
Caller-defined audio asset id.
assets.audio[].url: string
HTTPS URL for the audio file. The URL must be directly fetchable by Vivix without custom request headers.
assets.audio[].mime_type: optional string
Audio MIME type. Supported values are audio/wav and audio/mpeg.
assets.speech_text: optional array
Reusable exact speech text that can be referenced by scripted vocal output.
assets.speech_text[]: object
One reusable speech text asset.
assets.speech_text[].asset_id: string
Caller-defined speech text asset id.
assets.speech_text[].text: string
Exact text to synthesize and stream through output text events when used in response.script.
assets.visual_prompts: optional array
Reusable visual instructions that can be referenced by scripted visual output.
assets.visual_prompts[]: object
One reusable visual prompt asset.
assets.visual_prompts[].asset_id: string
Caller-defined visual prompt asset id.
assets.visual_prompts[].prompt: string
Instruction for avatar movement, posture, gaze, gesture, or presentation.
conversation: optional object
Initial conversational defaults such as instructions, tools, turn handling, and response defaults.
conversation.instructions: optional string
Default session and task instructions for generated responses. Effective instructions are composed from the selected avatar instructions, these conversation instructions, and any per-response instructions; conflicts are resolved in the order response.instructions, conversation.instructions, then avatars[].instructions.
conversation.tools: optional array
Tools available to responses in the session.
conversation.tools[]: object
One tool definition. Only function tools are documented for v1.
conversation.tools[].type: string
Use function.
conversation.tools[].name: string
Function name available to the response planner.
conversation.tools[].description: optional string
Human-readable description of when the tool should be used.
conversation.tools[].parameters: optional object
JSON Schema object describing function arguments.
conversation.response_defaults: optional object
Default response-generation settings for later turns.
conversation.response_defaults.temperature: optional number
Sampling temperature for generated response text. Supported range is 0 to 2.
conversation.response_defaults.max_output_tokens: optional integer
Maximum generated text tokens for a response.
conversation.turn_detection: optional object
Default turn handling for realtime user input.
conversation.turn_detection.type: string
manual means the client starts responses explicitly. server_vad means the server detects turn boundaries from incoming user audio.
conversation.turn_detection.silence_duration_ms: optional integer
Compatibility field. Do not use it to tune actual end-of-utterance timing. See Speech recognition configuration for recognition and turn-detection settings.
conversation.input_audio_transcription: optional object
Default transcription settings for user audio input.
conversation.input_audio_transcription.enabled: optional boolean
Whether user audio should be transcribed.
conversation.input_audio_transcription.language: optional string
BCP-47 language code, or auto for automatic language detection.
avatars: array
Session-local avatars available for conversation and rendering. A full create request requires a nonempty list with unique avatar_id values. An avatar used for video rendering needs a visual definition with a usable source image; a text-only avatar does not. Configure the session-wide voice in pipeline_config.tts_config.
avatars[].avatar_id: string
Caller-defined avatar identifier, unique within the session.
avatars[].instructions: optional string
The avatar’s identity, speaking style, and response rules. These instructions apply only when this avatar is selected by the current session.active_avatar_id, and they do not grant tools or data access.
avatars[].voice: optional object
Optional per-avatar TTS configuration for multi-avatar sessions that require different voices. Provide provider, provider_voice_id, and model. For the normal session-wide voice path, omit this object and configure pipeline_config.tts_config. A session TTS configuration takes precedence over every per-avatar voice.
avatars[].voice.provider: string
TTS backend, such as qwen-audio-3.0-tts-flash_ws, elevenlabs, or qwen3tts. Required with provider_voice_id and model.
avatars[].voice.provider_voice_id: string
Provider-side voice or style id. Required with provider and model.
avatars[].voice.model: string
Provider TTS model id. Required with provider and provider_voice_id.
avatars[].voice.speed: optional number
Speech speed multiplier. The valid range is 0.7 to 1.2.
avatars[].visual: object
Visual rendering settings. Required when this avatar is used for video rendering.
avatars[].visual.source_images: array
Registered source images available for this avatar. Required when image-based avatar video rendering is used.
avatars[].visual.source_images[]: object
One source image that can be selected for avatar video rendering.
avatars[].visual.source_images[].source_image_id: string
Caller-defined source image identifier, unique within the avatar.
avatars[].visual.source_images[].url: optional string
Accessible image URL. Required for a usable video source image and for each source image registered through avatar.add. A text-only avatar does not need a video source image.
avatars[].visual.source_images[].description: optional string
Optional description of the person, pose, and scene.
avatars[].visual.source_images[].media_type: required string
Image media type, such as image/png or image/jpeg.
avatars[].visual.default_source_image_id: optional string
With one usable source image, this field can be omitted. With multiple images and no explicit source_image_id, use it to select the initial image.
delivery: optional object
Media delivery settings used when session.mode is video_avatar. If omitted, the media transport defaults to TRTC. It is not used for text_chat.
delivery.media: optional object
Selects how the client sends and receives realtime media.
delivery.media.transport: optional string enum
Media transport requested for the session. Use trtc for TRTC or agora for Agora. If omitted, it defaults to TRTC. The response returns the corresponding provider-specific join details under delivery.media.trtc or delivery.media.agora.
max_duration_seconds: optional integer
Requested maximum session duration in seconds. If omitted, the service-configured default applies. An explicit value must be greater than 3 (minimum 4); 0 is rejected. The service may still close the session for account limits, idle timeout, failures, or maintenance.
auto_close: optional object
Automatic close policy fixed when the session is created. The response returns the effective values after platform defaults are applied.
auto_close.disconnected_timeout_seconds: optional integer
Closes after all control WebSocket connections have been absent for this period, including when none was ever opened. Omission uses the platform setting; read the effective value in the response. Use 0 to disable, or an integer from 10 to 3600 to enable.
auto_close.interaction_idle_timeout_seconds: optional integer
Closes the session after this duration without a successfully accepted client interaction or detected user speech. Accepted business events and user speech reset the timeout; WSS Ping/Pong and server events do not. Omitted or 0 disables this policy; explicit non-zero values must be 30–3600.
An explicit non-zero timeout cannot exceed the effective maximum session duration.
Automatic closure fields
Direct Create Session accepts an optional top-levelauto_close object. The response returns the effective rules under the same name.
auto_close.disconnected_timeout_seconds · optional integer, seconds. 0 disables the rule; nonzero values must be 10–3600. Omission uses the platform setting.
auto_close.interaction_idle_timeout_seconds · optional integer, seconds. Omitted or 0 disables the rule; nonzero values must be 30–3600.
An explicit nonzero value cannot exceed the effective maximum session duration. See Automatic session closure for setup and closing conditions.
Creates a Streaming Avatar session, registers reusable script assets and the avatars that can appear in the session, and returns the WSS control channel plus media connection details.
After the call returns, save session_id, control.url, control.client_secret, delivery.media, and the effective session.mode / active_avatar_id / source_image_id resolved by the server.
A browser must not open control.url without a token. Append the credential as the token query parameter:
URL.searchParams preserves the existing session_id and encodes the credential. The query parameter name must be token, not client_secret. Server-side WebSocket clients may instead send Authorization: Bearer CONTROL_CLIENT_SECRET. Keep the credential and the complete URL out of logs.
POST /v1/realtime-avatar/sessions
Start from a saved character. If you already have a character_id, use Create Session from Character instead of sending the full configuration again. Save characters in Characters.
Returns
session_id: string
Identifier for the created avatar session.
status: string
Current session status, such as active, closing, closed, or failed.
auto_close: object
The effective policy snapshotted when the session was first created.
auto_close.disconnected_timeout_seconds: integer
The effective no-control-WSS timeout in seconds. 0 means disabled.
auto_close.interaction_idle_timeout_seconds: integer
The effective interaction idle timeout in seconds. 0 means disabled.
region: string
Region selected for the session.
session: object
Effective mutable session state after defaults and validation are applied.
session.mode: string
Effective session mode, one of text_chat or video_avatar.
session.active_avatar_id: string
Avatar whose instructions are active for responses when no per-response avatar is selected.
session.source_image_id: optional string
Effective source image for avatar video. Omitted in text_chat mode.
control: object
Connection details for the session control channel.
control.url: string
Absolute WSS URL for the control channel, carrying client control events and server state events. The path is /v1/realtime-avatar/control and the session is identified by the session_id query parameter; always connect to the exact URL returned in this field.
control.client_secret: string
Session-scoped credential that authenticates the control channel. It must be sent during the WebSocket handshake: browsers append it to control.url as the token query parameter, and server-side clients may instead use a Bearer header. The query parameter name must be token, not client_secret. Safe to use from the browser (unlike your API key); it stays valid only until the session ends and cannot be refreshed. It is not single-use — you can reconnect the control channel with it during the same session; it stops working as soon as the session closes or expires, and there is no separate way to revoke it. Do not log it or reuse it across sessions.
output: optional object
Effective avatar video output settings after defaults and limits are applied. Present when session.mode is video_avatar.
output.aspect_ratio: string
Effective avatar video aspect ratio.
output.resolution: string
Effective avatar video resolution, either 480p or 720p.
output.fps: integer
Effective avatar video frame rate applied by the server. Defaults to 24.
assets: optional object
Effective reusable script assets available in the session.
assets.audio: optional array
Reusable audio assets for response.script.vocal.type: "audio_asset".
assets.speech_text: optional array
Reusable exact speech text assets for response.script.vocal.type: "speech_asset".
assets.visual_prompts: optional array
Reusable visual prompt assets for response.script.visual.visual_prompt_asset_id.
conversation: object
Effective conversation defaults after server limits and defaults are applied. Same structure and meaning as the conversation request field.
avatars: array
Effective avatar definitions in the session. Same structure and meaning as the avatars request field.
delivery: optional object
Media connection details returned by the server. Present when session.mode is video_avatar.
delivery.media: optional oneOf object
Temporary media connection details. Present when avatar video media is available. Exactly one media variant is returned, selected by transport. These details are scoped to this session and are not provider account credentials.
delivery.media.transport: string enum
Discriminator that selects which media variant is returned. Supported values are trtc and agora.
delivery.media.trtc: required object
TRTC room join details for the trtc media variant.
delivery.media.trtc.sdk_app_id: string
TRTC application id used by the client SDK. Returned as a string, consistent with the convention that 64-bit integers may be returned as strings.
delivery.media.trtc.room_id: string
TRTC room id for this session.
delivery.media.trtc.user_id: string
TRTC user id assigned to the client.
delivery.media.trtc.user_sig: string
Session-scoped signature for joining the TRTC room.
delivery.media.trtc.publisher_user_id: string
User id of the avatar video publisher in the TRTC room. Subscribe only to this user id to receive the avatar stream.
delivery.media.agora: required object
Agora channel join details for the agora media variant.
delivery.media.agora.app_id: string
Agora application id used by the client SDK.
delivery.media.agora.channel_name: string
Agora channel name for this session.
delivery.media.agora.token: string
Session-scoped token for joining the Agora channel.
delivery.media.agora.user_id: string
Agora user id assigned to the client.
delivery.media.agora.publisher_user_id: string
User id of the avatar video publisher in the Agora channel. Subscribe only to this user id to receive the avatar stream.
model: string
The model from the request.
expires_at: optional timestamp
Time at which the session and its control credential expire. Present when the server returns a fixed expiration; otherwise omitted.
Request
3. Get Session
Closure and recording results
close_reason · string. Closure reason, such as disconnected_timeout or interaction_idle_timeout. Check it with status and closed_at. failed is a terminal failure; do not keep waiting for closed.
recording_mode · string. The effective recording mode for this session: on or off.
Retrieves the current status and effective configuration for a Streaming Avatar session. Use this endpoint to poll after a close request is accepted when the client needs confirmation that cleanup has completed.
GET /v1/realtime-avatar/sessions/{session_id}
Path Parameters
session_id: string
Identifier of the Streaming Avatar session to retrieve.
Returns
Returns the same effective fields as Create Session. The auto-close snapshot and fields that differ from Create Session are detailed below.auto_close: object
The effective automatic close policy fixed when the session was created.
auto_close.disconnected_timeout_seconds: integer
The effective no-control-WSS timeout in seconds. 0 means disabled.
auto_close.interaction_idle_timeout_seconds: integer
The effective interaction idle timeout in seconds. 0 means disabled.
closed_at: optional timestamp
Cleanup completion timestamp. Present after the session reaches a terminal state.
close_reason: optional string
The first reason that initiated shutdown, including disconnected_timeout and interaction_idle_timeout.
delivery: optional object
Present when session.mode is video_avatar and delivery details are still available; delivery.media is returned only while media connection details are still available.
control.client_secret: string
Returned only in the Create Session response; Get Session returns an empty value. Save the control channel credential when the session is created.
Request
4. List Sessions
Lists the open (not yet closed) Streaming Avatar session ids for your workspace, newest first. Session id is the only field returned; use it withGET /v1/realtime-avatar/sessions/{session_id} to fetch full state, or POST /v1/realtime-avatar/sessions/{session_id}/close to close a session that already ended but still holds a workspace concurrency slot.
GET /v1/realtime-avatar/sessions
Query Parameters
page_size: optional integer Maximum number of session ids to return. Defaults to100; the maximum accepted value is 200. Values outside the range fall back to the default.
Returns
session_ids: array of string Open session ids for the current workspace, ordered by most recently created first. An empty array is returned when there are no open sessions.page_size only caps the number of ids returned; it does not paginate, so a single call returns at most page_size ids (200 at most). Pass a larger page_size to raise the cap when a workspace holds many open sessions.
Request
5. Close Session
Requests session closure. To let current media finish, wait forresponse.render.stopped before closing. The endpoint usually returns closing; retrieve the Session to confirm closed or handle failed.
POST /v1/realtime-avatar/sessions/{session_id}/close
Path Parameters
session_id: string
Identifier of the Streaming Avatar session to close.
Body Parameters
close_reason: optional string
Caller-defined reason stored as the session close reason.
Returns
session_id: string
Identifier for the closing avatar session.
status: string
Close status. Returns closing while shutdown is in progress, or closed when it has completed.
Request
6. Session Error Cases
Create request has no avatars, duplicate avatar ids, incomplete avatar definitions, or an invalid initial session state. · API error code:30004 invalid argument
The requested model, output, or maximum duration is not supported. · API error code: 30004 invalid argument
An auto-close timeout is outside its allowed range or exceeds the maximum session duration. · API error code: 30004 invalid argument
The session id does not exist or the realtime URL has expired. · API error code: 20005 not found
The API key is missing or invalid. · API error code: 10001 missing api key or 10003 invalid api key
Workspace billing is inactive or the balance is insufficient. · API error code: 10004 workspace billing is not active or 10005 insufficient workspace balance
The request exceeded the API key rate limit (HTTP 429; 60 requests/min per key by default, configurable; the response includes a Retry-After header). Back off exponentially and honor Retry-After. · API error code: 10008 API rate limit exceeded
The number of concurrently open sessions for this workspace reached its limit (HTTP 429, no Retry-After). Call List Sessions to find unclosed sessions and Close the ones that already ended but still hold quota, then retry; if your business needs a higher concurrency limit, contact sales to adjust the workspace quota. · API error code: 30006 workspace concurrency limit exceeded
Platform-side model capacity is currently insufficient (HTTP 429, no Retry-After). This is not a quota issue: retry after a short backoff, and contact us if it persists so capacity can be extended. · API error code: 30008 model resource capacity exceeded
This API key has not been granted access to the requested model or API. Confirm in the console that the app has the corresponding model/API permission enabled, ask your account owner to enable access, then retry. · API error code: 10009 api access denied
Error responses use the same {code, message, data} structure. A non-zero code indicates failure, and message is an English error description.
Pipeline configuration fields
These fields belong in the creation request’spipeline_config object. Merge them with the existing configuration when creating a new session. They cannot be changed through session.update.
tag_audio_enabled · boolean · Optional; defaults to false. Enables faster audio output for an eligible first short speech segment of the response, not emotion tags. Supply a boolean, not a string. See opening acknowledgements.
asr_config
provider · string · Required when asr_config is supplied. Common values are doubao and nova-3. The platform supplies a default service when the whole object is omitted, but not when only language is provided.
language · string · Optional. Examples: zh-CN or en-US for Doubao; en for Nova. Controls recognition language, not the reply language.
eou_timeout_ms · string · Optional end-of-utterance judge timeout, for example “1500”. Not a fixed silence length or total response latency.
llm_config
llm_backend · string · Optional; only openai is supported.
llm_base_url · string · Optional OpenAI-compatible endpoint. Supply with llm_model for a custom service; omit when only tuning generation settings.
llm_model · string · Optional custom model name.
llm_api_key · string · Optional service credential. Keep private credentials out of browser code and public configuration.
llm_temperature · number · Optional generation temperature; use the range supported by the model.
llm_max_output_tokens · integer · Optional generated-token limit, not a word count.
llm_extra_payload · string · Optional JSON-encoded object string; do not supply a nested object.
tts_config
tts_voice_id · string · Required nonempty provider voice ID or Vivix cloned voice_id.
tts_provider · string · Optional. For built-in Qwen Audio voices, defaults to qwen-audio-3.0-tts-flash_ws. A Vivix cloned voice_id resolves to its associated speech service. Set the matching value for other services.
tts_model_id · string · Optional. For built-in Qwen Audio voices, defaults to qwen-audio-3.0-tts-flash. A Vivix cloned voice_id resolves to its associated model. Set the matching value for other services.
tts_endpoint · string · Optional endpoint override when nonempty. Not needed for built-in voices.
tts_api_key · string · Optional provider credential when nonempty. Built-in voices need no separate key.
payload · object · Optional speech-service extensions; arbitrary provider parameters are not guaranteed to take effect. Do not duplicate reserved authentication or routing fields here.
Source image media type
avatars[].visual.source_images[].media_type is required and must start with image/, for example image/png or image/jpeg. Use the actual image format, not the video output format.