> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create Session

> Creates a Streaming Avatar session, registers reusable script assets and the avatars that can appear in the session, and returns the WSS control channel plus media connection details.

After the call returns, save `session_id`, `control.url`, `control.client_secret`, `delivery.media`, and the effective `session.mode` / `active_avatar_id` / `source_image_id` resolved by the server.

**Start from a saved character.** If you already have a `character_id`, use Create Session from Character instead of sending the full configuration again. Save characters in [Characters](/streaming-avatar/api-references/characters).

<AccordionGroup>
  <Accordion title="Connect to the control channel">
    A browser must not open `control.url` without a token. Append the credential as the token query parameter:

    ```javascript theme={null}
    const controlUrl = new URL(session.control.url);
    controlUrl.searchParams.set("token", session.control.client_secret);
    const ws = new WebSocket(controlUrl.toString());
    ```

    `URL.searchParams` preserves the existing `session_id` and encodes the credential. The query parameter name must be token, not client\_secret. Server-side WebSocket clients may instead send Authorization: Bearer CONTROL\_CLIENT\_SECRET. Keep the credential and the complete URL out of logs.
  </Accordion>

  <Accordion title="Automatic closure">
    Direct Create Session accepts an optional top-level `auto_close` object. The response returns the effective rules under the same name.

    An explicit nonzero value cannot exceed the effective maximum session duration. See [Automatic session closure](/streaming-avatar/integrate/sessions) for setup and closing conditions.
  </Accordion>

  <Accordion title="Errors">
    Error responses use the same `{code, message, data}` envelope; a non-zero `code` identifies the error.

    | Case | Error code |
    | - | - |
    | Create request has no avatars, duplicate avatar ids, incomplete avatar definitions, or an invalid initial session state | `30004 invalid argument` |
    | The requested model, output, or maximum duration is not supported | `30004 invalid argument` |
    | An auto-close timeout is outside its allowed range or exceeds the maximum session duration | `30004 invalid argument` |
    | The session id does not exist or the realtime URL has expired | `20005 not found` |
    | The API key is missing or invalid | `10001 missing api key` or `10003 invalid api key` |
    | Workspace billing is inactive or the balance is insufficient | `10004 workspace billing is not active` or `10005 insufficient workspace balance` |
    | The request exceeded the API key rate limit (HTTP 429; 60 requests/min per key by default, configurable; the response includes a Retry-After header). Back off exponentially and honor Retry-After | `10008 API rate limit exceeded` |
    | The number of concurrently open sessions for this workspace reached its limit (HTTP 429, no Retry-After). Call List Sessions to find unclosed sessions and Close the ones that already ended but still hold quota, then retry; if your business needs a higher concurrency limit, contact sales to adjust the workspace quota | `30006 workspace concurrency limit exceeded` |
    | Platform-side model capacity is currently insufficient (HTTP 429, no Retry-After). This is not a quota issue: retry after a short backoff, and contact us if it persists so capacity can be extended | `30008 model resource capacity exceeded` |
    | This API key has not been granted access to the requested model or API. Confirm in the console that the app has the corresponding model/API permission enabled, ask your account owner to enable access, then retry | `10009 api access denied` |
  </Accordion>
</AccordionGroup>


## OpenAPI

````yaml streaming-avatar/api-references/openapi.json POST /v1/realtime-avatar/sessions
openapi: 3.1.0
info:
  title: Vivix Streaming Avatar API
  version: 1.0.0
  description: REST endpoints for Streaming Avatar sessions, voices, and characters.
servers:
  - url: https://api.vivix.ai
security:
  - bearerAuth: []
paths:
  /v1/realtime-avatar/sessions:
    post:
      summary: Create Session
      description: >-
        Creates a Streaming Avatar session, registers reusable script assets and
        the avatars that can appear in the session, and returns the WSS control
        channel plus media connection details.


        After the call returns, save `session_id`, `control.url`,
        `control.client_secret`, `delivery.media`, and the effective
        `session.mode` / `active_avatar_id` / `source_image_id` resolved by the
        server.


        **Start from a saved character.** If you already have a `character_id`,
        use Create Session from Character instead of sending the full
        configuration again. Save characters in
        [Characters](/streaming-avatar/api-references/characters).
      operationId: createSession
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                model:
                  type: string
                  description: >-
                    Use `vivix-a1-stream`. A lightweight variant
                    `vivix-a1-stream-lite` is also available.
                idempotency_key:
                  type: string
                  description: >-
                    Client-generated key used to prevent duplicate session
                    allocation when a create request is retried. On a
                    429/5xx/timeout, retry the create call with the same
                    idempotency_key; it is safe and will not allocate a
                    duplicate session.
                region:
                  type: string
                  description: >-
                    Optional deployment region hint. Returned back on the
                    session object; may be empty if not provided. It does not
                    currently affect scheduling or guarantee data residency.
                session:
                  type: object
                  description: >-
                    Initial mutable session state for mode, active avatar, and
                    optional video source image. Use `session.update` to change
                    these values after the session starts.
                  properties:
                    mode:
                      type: string
                      description: >-
                        Initial session mode. Supported values are `text_chat`
                        and `video_avatar`. Defaults to `video_avatar`.
                        `text_chat` returns text only, and `video_avatar`
                        returns avatar video with generated speech and text
                        events.
                    active_avatar_id:
                      type: string
                      description: >-
                        Avatar whose instructions are active when the session
                        starts. In `video_avatar` mode, this is also the avatar
                        rendered in the video stream. If omitted and exactly one
                        avatar is provided, that avatar is used; if omitted with
                        multiple avatars, the request fails.
                    source_image_id:
                      type: string
                      description: >-
                        Source image to use for avatar video. Used only when
                        `mode` is `video_avatar`. If omitted in `video_avatar`,
                        uses the active avatar’s
                        `visual.default_source_image_id`. If no default is set
                        and that avatar has exactly one source image, it is
                        selected automatically. In video mode, the request fails
                        if no usable image can be resolved. In `text_chat`, the
                        field is omitted from the REST response.
                output:
                  type: object
                  description: >-
                    Create-time avatar video output settings for the session.
                    Required for `video_avatar` mode and applied when
                    `session.mode` is `video_avatar`; ignored in `text_chat`
                    mode.
                  properties:
                    aspect_ratio:
                      type: string
                      description: >-
                        Requested avatar video aspect ratio, such as `9:16`,
                        `1:1`, or `16:9`.
                    resolution:
                      type: string
                      description: >-
                        Requested avatar video resolution. Supported values are
                        `480p` and `720p`.
                avatars:
                  type: array
                  description: >-
                    Session-local avatars available for conversation and
                    rendering. A full create request requires a nonempty list
                    with unique `avatar_id` values. An avatar used for video
                    rendering needs a visual definition with a usable source
                    image; a text-only avatar does not. Configure the
                    session-wide voice in `pipeline_config.tts_config`.
                  items:
                    type: object
                    properties:
                      avatar_id:
                        type: string
                        description: >-
                          Caller-defined avatar identifier, unique within the
                          session.
                      instructions:
                        type: string
                        description: >-
                          The avatar’s identity, speaking style, and response
                          rules. These instructions apply only when this avatar
                          is selected by the current `session.active_avatar_id`,
                          and they do not grant tools or data access.
                      voice:
                        type: object
                        description: >-
                          Optional per-avatar TTS configuration for multi-avatar
                          sessions that require different voices. Provide
                          `provider`, `provider_voice_id`, and `model`. For the
                          normal session-wide voice path, omit this object and
                          configure `pipeline_config.tts_config`. A session TTS
                          configuration takes precedence over every per-avatar
                          voice.
                        properties:
                          provider:
                            type: string
                            description: >-
                              TTS backend, such as
                              `qwen-audio-3.0-tts-flash_ws`, `elevenlabs`, or
                              `qwen3tts`. Required with `provider_voice_id` and
                              `model`.
                          provider_voice_id:
                            type: string
                            description: >-
                              Provider-side voice or style id. Required with
                              `provider` and `model`.
                          model:
                            type: string
                            description: >-
                              Provider TTS model id. Required with `provider`
                              and provider_voice_id.
                          speed:
                            type: number
                            description: >-
                              Speech speed multiplier. The valid range is `0.7`
                              to `1.2`.
                        required:
                          - provider
                          - provider_voice_id
                          - model
                      visual:
                        type: object
                        description: >-
                          Visual rendering settings. Required when this avatar
                          is used for video rendering.
                        properties:
                          source_images:
                            type: array
                            description: >-
                              Registered source images available for this
                              avatar. Required when image-based avatar video
                              rendering is used.
                            items:
                              type: object
                              description: >-
                                One source image that can be selected for avatar
                                video rendering.
                              properties:
                                source_image_id:
                                  type: string
                                  description: >-
                                    Caller-defined source image identifier,
                                    unique within the avatar.
                                url:
                                  type: string
                                  description: >-
                                    Accessible image URL. Required for a usable
                                    video source image and for each source image
                                    registered through `avatar.add`. A text-only
                                    avatar does not need a video source image.
                                description:
                                  type: string
                                  description: >-
                                    Optional description of the person, pose,
                                    and scene.
                                media_type:
                                  type: string
                                  description: >-
                                    Image media type, such as `image/png` or
                                    `image/jpeg`.
                              required:
                                - source_image_id
                                - media_type
                          default_source_image_id:
                            type: string
                            description: >-
                              With one usable source image, this field can be
                              omitted. With multiple images and no explicit
                              `source_image_id`, use it to select the initial
                              image.
                        required:
                          - source_images
                    required:
                      - avatar_id
                      - visual
                pipeline_config:
                  type: object
                  description: >-
                    Session-scoped pipeline configuration. For normal generated
                    speech, provide `tts_config` here with a non-empty
                    `tts_voice_id`; `tts_provider` and `tts_model_id` may be
                    omitted and then default to the Qwen Audio pair. It applies
                    to every avatar and cannot be changed after session
                    creation. See [Speech recognition
                    configuration](/streaming-avatar/integrate/speech-recognition).
                  properties:
                    tag_audio_enabled:
                      type: boolean
                      description: >-
                        Optional; defaults to false. Enables faster audio output
                        for an eligible first short speech segment of the
                        response, not emotion tags. Supply a boolean, not a
                        string. See [opening
                        acknowledgements](/streaming-avatar/interaction/acknowledgement).
                    tts_config:
                      type: object
                      properties:
                        tts_voice_id:
                          type: string
                          description: >-
                            Required nonempty `provider` voice ID or Vivix
                            cloned voice_id.
                        tts_provider:
                          type: string
                          description: >-
                            Optional. For built-in Qwen Audio voices, defaults
                            to `qwen-audio-3.0-tts-flash_ws`. A Vivix cloned
                            `voice_id` resolves to its associated speech
                            service. Set the matching value for other services.
                        tts_model_id:
                          type: string
                          description: >-
                            Optional. For built-in Qwen Audio voices, defaults
                            to `qwen-audio-3.0-tts-flash`. A Vivix cloned
                            `voice_id` resolves to its associated model. Set the
                            matching value for other services.
                        tts_endpoint:
                          type: string
                          description: >-
                            Optional endpoint override when nonempty. Not needed
                            for built-in voices.
                        tts_api_key:
                          type: string
                          description: >-
                            Optional `provider` credential when nonempty.
                            Built-in voices need no separate key.
                        payload:
                          type: object
                          description: >-
                            Optional speech-service extensions; arbitrary
                            `provider` parameters are not guaranteed to take
                            effect. Do not duplicate reserved authentication or
                            routing fields here.
                      required:
                        - tts_voice_id
                      description: >-
                        Session-wide TTS configuration. It takes precedence over
                        every `avatars[].voice`.
                    asr_config:
                      type: object
                      properties:
                        provider:
                          type: string
                          description: >-
                            Required when `asr_config` is supplied. Common
                            values are `doubao` and nova-3. The platform
                            supplies a default service when the whole object is
                            omitted, but not when only `language` is provided.
                        language:
                          type: string
                          description: >-
                            Optional. Examples: `zh-CN` or `en-US` for Doubao;
                            en for Nova. Controls recognition `language`, not
                            the reply language.
                        eou_timeout_ms:
                          type: string
                          description: >-
                            Optional end-of-utterance judge timeout, for example
                            "1500". Not a fixed silence length or total response
                            latency.
                      required:
                        - provider
                      description: >-
                        Advanced override. Omit it to use the platform default,
                        `doubao`. Set `provider` (`doubao` for multilingual, or
                        `nova-3` for English-only) and optional `language`. See
                        [`asr_config`](/streaming-avatar/integrate/speech-recognition).
                    llm_config:
                      type: object
                      properties:
                        llm_backend:
                          type: string
                          description: Optional; only `openai` is supported.
                        llm_base_url:
                          type: string
                          description: >-
                            Optional OpenAI-compatible endpoint. Supply with
                            llm_model for a custom service; omit when only
                            tuning generation settings.
                        llm_model:
                          type: string
                          description: Optional custom model name.
                        llm_api_key:
                          type: string
                          description: >-
                            Optional service credential. Keep private
                            credentials out of browser code and public
                            configuration.
                        llm_temperature:
                          type: number
                          description: >-
                            Optional generation temperature; use the range
                            supported by the model.
                        llm_max_output_tokens:
                          type: integer
                          description: Optional generated-token limit, not a word count.
                        llm_extra_payload:
                          type: string
                          description: >-
                            Optional JSON-encoded object string; do not supply a
                            nested object.
                      description: >-
                        `llm_config` · optional object. Set `llm_base_url` and
                        `llm_model` for a custom service, or override only
                        generation parameters. See [Text model
                        configuration](/streaming-avatar/integrate/dialogue-model)
                        for field types and examples.
                    motion_enhanced:
                      type: object
                      description: >-
                        `prompt` drives the motion generated while the avatar is
                        speaking. See
                        [motion_enhanced](/streaming-avatar/character/speaking-motion).
                      properties:
                        prompt:
                          type: string
                          description: >-
                            Motion prompt. See the linked guide for how to write
                            it.
                    motion_planner:
                      type: object
                      description: >-
                        `prompt` drives listening (idle) motion. The key is
                        `motion_planner`, not `idle_motion`. See
                        [motion_planner](/streaming-avatar/character/listening-motion).
                      properties:
                        prompt:
                          type: string
                          description: >-
                            Motion prompt. See the linked guide for how to write
                            it.
                  required:
                    - tts_config
                conversation:
                  type: object
                  description: >-
                    Initial conversational defaults such as instructions, tools,
                    turn handling, and response defaults.
                  properties:
                    instructions:
                      type: string
                      description: >-
                        Default session and task instructions for generated
                        responses. Effective instructions are composed from the
                        selected avatar instructions, these conversation
                        instructions, and any per-response instructions;
                        conflicts are resolved in the order
                        `response.instructions`, `conversation.instructions`,
                        then `avatars[].instructions`.
                    tools:
                      type: array
                      description: Tools available to responses in the session.
                      items:
                        type: object
                        description: >-
                          One tool definition. Only function tools are
                          documented for v1.
                        properties:
                          type:
                            type: string
                            description: Use `function`.
                          name:
                            type: string
                            description: Function name available to the response planner.
                          description:
                            type: string
                            description: >-
                              Human-readable description of when the tool should
                              be used.
                          parameters:
                            type: object
                            description: JSON Schema object describing function arguments.
                        required:
                          - type
                          - name
                    response_defaults:
                      type: object
                      description: Default response-generation settings for later turns.
                      properties:
                        temperature:
                          type: number
                          description: >-
                            Sampling temperature for generated response text.
                            Supported range is `0` to `2`.
                        max_output_tokens:
                          type: integer
                          description: Maximum generated text tokens for a response.
                    turn_detection:
                      type: object
                      description: Default turn handling for realtime user input.
                      properties:
                        type:
                          type: string
                          description: >-
                            `manual` means the client starts responses
                            explicitly. `server_vad` means the server detects
                            turn boundaries from incoming user audio.
                        silence_duration_ms:
                          type: integer
                          description: >-
                            Compatibility field. Do not use it to tune actual
                            end-of-utterance timing. See [Speech recognition
                            configuration](/streaming-avatar/integrate/speech-recognition)
                            for recognition and turn-detection settings.
                      required:
                        - type
                    input_audio_transcription:
                      type: object
                      description: Default transcription settings for user audio input.
                      properties:
                        enabled:
                          type: boolean
                          description: Whether user audio should be transcribed.
                        language:
                          type: string
                          description: >-
                            BCP-47 `language` code, or `auto` for automatic
                            `language` detection.
                assets:
                  type: object
                  description: >-
                    Session-scoped reusable assets for `response.script`. Asset
                    ids must be unique across `audio`, `speech_text`, and
                    `visual_prompts`. Avatar source images stay under
                    `avatars[].visual.source_images`. On the REST side the
                    visual prompt text field is `prompt`; on the WSS `asset.add`
                    side the same asset uses `text` — the two shapes differ.
                  properties:
                    audio:
                      type: array
                      description: >-
                        Reusable audio files that can be referenced by scripted
                        vocal output.
                      items:
                        type: object
                        description: One reusable audio asset.
                        properties:
                          asset_id:
                            type: string
                            description: Caller-defined audio asset id.
                          url:
                            type: string
                            description: >-
                              HTTPS URL for the audio file. The URL must be
                              directly fetchable by Vivix without custom request
                              headers.
                          mime_type:
                            type: string
                            description: >-
                              Audio MIME type. Supported values are `audio/wav`
                              and `audio/mpeg`.
                        required:
                          - asset_id
                          - url
                    speech_text:
                      type: array
                      description: >-
                        Reusable exact speech text that can be referenced by
                        scripted vocal output.
                      items:
                        type: object
                        description: One reusable speech text asset.
                        properties:
                          asset_id:
                            type: string
                            description: Caller-defined speech text asset id.
                          text:
                            type: string
                            description: >-
                              Exact text to synthesize and stream through output
                              text events when used in `response.script`.
                        required:
                          - asset_id
                          - text
                    visual_prompts:
                      type: array
                      description: >-
                        Reusable visual instructions that can be referenced by
                        scripted visual output.
                      items:
                        type: object
                        description: One reusable visual prompt asset.
                        properties:
                          asset_id:
                            type: string
                            description: Caller-defined visual prompt asset id.
                          prompt:
                            type: string
                            description: >-
                              Instruction for avatar movement, posture, gaze,
                              gesture, or presentation.
                        required:
                          - asset_id
                          - prompt
                delivery:
                  type: object
                  description: >-
                    Media delivery settings used when `session.mode` is
                    `video_avatar`. If omitted, the media transport defaults to
                    TRTC. It is not used for `text_chat`.
                  properties:
                    media:
                      type: object
                      description: >-
                        Selects how the client sends and receives realtime
                        media.
                      properties:
                        transport:
                          type: string
                          description: >-
                            Media transport requested for the session. Use
                            `trtc` for TRTC or `agora` for Agora. If omitted, it
                            defaults to TRTC. The response returns the
                            corresponding provider-specific join details under
                            `delivery.media.trtc` or `delivery.media.agora`.
                max_duration_seconds:
                  type: integer
                  description: >-
                    Requested maximum session duration in seconds. If omitted,
                    the service-configured default applies. An explicit value
                    must be greater than `3` (minimum 4); 0 is rejected. The
                    service may still close the session for account limits, idle
                    timeout, failures, or maintenance.
                auto_close:
                  type: object
                  description: >-
                    Automatic close policy fixed when the session is created.
                    The response returns the effective values after platform
                    defaults are applied.
                  properties:
                    disconnected_timeout_seconds:
                      type: integer
                      description: >-
                        Closes after all control WebSocket connections have been
                        absent for this period, including when none was ever
                        opened. Omission uses the platform setting; read the
                        effective value in the response. Use `0` to disable, or
                        an integer from `10` to `3600` to enable.


                        seconds. `0` disables the rule; nonzero values must be
                        `10`–`3600`. Omission uses the platform setting.
                    interaction_idle_timeout_seconds:
                      type: integer
                      description: >-
                        Closes the session after this duration without a
                        successfully accepted client interaction or detected
                        user speech. Accepted business events and user speech
                        reset the timeout; WSS Ping/Pong and server events do
                        not. Omitted or 0 disables this policy; explicit
                        non-zero values must be 30–3600.


                        An explicit non-zero timeout cannot exceed the effective
                        maximum session duration.


                        seconds. Omitted or `0` disables the rule; nonzero
                        values must be `30`–`3600`.
                recording_mode:
                  type: string
                  description: '`on` or `off`. Send `off` to disable recording.'
              required:
                - model
                - output
                - avatars
            example:
              model: vivix-a1-stream
              idempotency_key: avatar-create-123
              region: us-west
              session:
                mode: video_avatar
                active_avatar_id: host_a
                source_image_id: front
              output:
                aspect_ratio: '9:16'
                resolution: 480p
              pipeline_config:
                tts_config:
                  tts_provider: qwen-audio-3.0-tts-flash_ws
                  tts_voice_id: longanhuan_v3.6
                  tts_model_id: qwen-audio-3.0-tts-flash
              assets:
                audio:
                  - asset_id: welcome_jingle
                    url: https://cdn.example.com/audio/welcome-jingle.wav
                    mime_type: audio/wav
                speech_text:
                  - asset_id: welcome_line
                    text: Welcome back. I saved your place.
                visual_prompts:
                  - asset_id: wave_small
                    prompt: smile and give a small wave
              conversation:
                instructions: Answer as a concise livestream host.
                tools:
                  - type: function
                    name: lookup_product
                    description: Look up product details by product id.
                    parameters:
                      type: object
                      properties:
                        product_id:
                          type: string
                      required:
                        - product_id
                response_defaults:
                  temperature: 0.7
                  max_output_tokens: 512
                turn_detection:
                  type: server_vad
                  silence_duration_ms: 500
                input_audio_transcription:
                  enabled: true
                  language: auto
              avatars:
                - avatar_id: host_a
                  instructions: >-
                    You are Host A, a warm livestream host who gives concise
                    product-focused replies.
                  visual:
                    source_images:
                      - source_image_id: front
                        url: https://cdn.example.com/avatars/host-a-front.png
                        description: >-
                          A woman facing the camera in a brightly lit studio,
                          front view.
                        media_type: image/png
                      - source_image_id: side
                        url: https://cdn.example.com/avatars/host-a-side.png
                        description: The same woman in the same studio, side view.
                        media_type: image/png
                    default_source_image_id: front
                - avatar_id: host_b
                  instructions: >-
                    You are Host B, a calm expert who gives brief technical
                    explanations.
                  visual:
                    source_images:
                      - source_image_id: front
                        url: https://cdn.example.com/avatars/host-b-front.png
                        description: >-
                          A man facing the camera in a brightly lit studio,
                          front view.
                        media_type: image/png
                    default_source_image_id: front
              auto_close:
                disconnected_timeout_seconds: 60
                interaction_idle_timeout_seconds: 300
              delivery:
                media:
                  transport: trtc
              max_duration_seconds: 300
      responses:
        '200':
          description: Success
          content:
            application/json:
              schema:
                type: object
                required:
                  - code
                  - message
                  - data
                properties:
                  code:
                    type: integer
                    description: '`0` on success; any other value is an error code.'
                  message:
                    type: string
                    description: '`success`, or an English error description.'
                  data:
                    type: object
                    properties:
                      session_id:
                        type: string
                        description: Identifier for the created avatar session.
                      status:
                        type: string
                        description: >-
                          Current session status, such as `active`, `closing`,
                          `closed`, or `failed`.
                      auto_close:
                        type: object
                        description: >-
                          The effective policy snapshotted when the session was
                          first created.
                        properties:
                          disconnected_timeout_seconds:
                            type: integer
                            description: >-
                              The effective no-control-WSS timeout in seconds. 0
                              means disabled.
                          interaction_idle_timeout_seconds:
                            type: integer
                            description: >-
                              The effective interaction idle timeout in seconds.
                              0 means disabled.
                        required:
                          - disconnected_timeout_seconds
                          - interaction_idle_timeout_seconds
                      region:
                        type: string
                        description: Region selected for the session.
                      session:
                        type: object
                        description: >-
                          Effective mutable session state after defaults and
                          validation are applied.
                        properties:
                          mode:
                            type: string
                            description: >-
                              Effective session mode, one of `text_chat` or
                              `video_avatar`.
                          active_avatar_id:
                            type: string
                            description: >-
                              Avatar whose instructions are active for responses
                              when no per-response avatar is selected.
                          source_image_id:
                            type: string
                            description: >-
                              Effective source image for avatar video. Omitted
                              in `text_chat` mode.
                        required:
                          - mode
                          - active_avatar_id
                      control:
                        type: object
                        description: Connection details for the session control channel.
                        properties:
                          url:
                            type: string
                            description: >-
                              Absolute WSS URL for the control channel, carrying
                              client control events and server state events. The
                              path is `/v1/realtime-avatar/control` and the
                              session is identified by the `session_id` query
                              parameter; always connect to the exact URL
                              returned in this field.
                          client_secret:
                            type: string
                            description: >-
                              Session-scoped credential that authenticates the
                              control channel. It must be sent during the
                              WebSocket handshake: browsers append it to
                              `control.url` as the token query parameter, and
                              server-side clients may instead use a Bearer
                              header. The query parameter name must be token,
                              not client_secret. Safe to use from the browser
                              (unlike your API key); it stays valid only until
                              the session ends and cannot be refreshed. It is
                              not single-use — you can reconnect the control
                              channel with it during the same session; it stops
                              working as soon as the session closes or expires,
                              and there is no separate way to revoke it. Do not
                              log it or reuse it across sessions.
                        required:
                          - url
                          - client_secret
                      output:
                        type: object
                        description: >-
                          Effective avatar video output settings after defaults
                          and limits are applied. Present when `session.mode` is
                          `video_avatar`.
                        properties:
                          aspect_ratio:
                            type: string
                            description: Effective avatar video aspect ratio.
                          resolution:
                            type: string
                            description: >-
                              Effective avatar video resolution, either `480p`
                              or `720p`.
                          fps:
                            type: integer
                            description: >-
                              Effective avatar video frame rate applied by the
                              server. Defaults to `24`.
                        required:
                          - aspect_ratio
                          - resolution
                          - fps
                      assets:
                        type: object
                        description: >-
                          Effective reusable script assets available in the
                          session.
                        properties:
                          audio:
                            type: array
                            description: >-
                              Reusable audio assets for
                              `response.script.vocal.type: "audio_asset"`.
                            items:
                              type: object
                              description: One reusable audio asset.
                              properties:
                                asset_id:
                                  type: string
                                  description: Caller-defined audio asset id.
                                url:
                                  type: string
                                  description: >-
                                    HTTPS URL for the audio file. The URL must
                                    be directly fetchable by Vivix without
                                    custom request headers.
                                mime_type:
                                  type: string
                                  description: >-
                                    Audio MIME type. Supported values are
                                    `audio/wav` and `audio/mpeg`.
                              required:
                                - asset_id
                                - url
                          speech_text:
                            type: array
                            description: >-
                              Reusable exact speech text assets for
                              `response.script.vocal.type: "speech_asset"`.
                            items:
                              type: object
                              description: One reusable speech text asset.
                              properties:
                                asset_id:
                                  type: string
                                  description: Caller-defined speech text asset id.
                                text:
                                  type: string
                                  description: >-
                                    Exact text to synthesize and stream through
                                    output text events when used in
                                    `response.script`.
                              required:
                                - asset_id
                                - text
                          visual_prompts:
                            type: array
                            description: >-
                              Reusable visual prompt assets for
                              `response.script.visual.visual_prompt_asset_id`.
                            items:
                              type: object
                              description: One reusable visual prompt asset.
                              properties:
                                asset_id:
                                  type: string
                                  description: Caller-defined visual prompt asset id.
                                prompt:
                                  type: string
                                  description: >-
                                    Instruction for avatar movement, posture,
                                    gaze, gesture, or presentation.
                              required:
                                - asset_id
                                - prompt
                      conversation:
                        type: object
                        description: >-
                          Effective conversation defaults after server limits
                          and defaults are applied. Same structure and meaning
                          as the `conversation` request field.
                        properties:
                          instructions:
                            type: string
                            description: >-
                              Default session and task instructions for
                              generated responses. Effective instructions are
                              composed from the selected avatar instructions,
                              these conversation instructions, and any
                              per-response instructions; conflicts are resolved
                              in the order `response.instructions`,
                              `conversation.instructions`, then
                              `avatars[].instructions`.
                          tools:
                            type: array
                            description: Tools available to responses in the session.
                            items:
                              type: object
                              description: >-
                                One tool definition. Only function tools are
                                documented for v1.
                              properties:
                                type:
                                  type: string
                                  description: Use `function`.
                                name:
                                  type: string
                                  description: >-
                                    Function name available to the response
                                    planner.
                                description:
                                  type: string
                                  description: >-
                                    Human-readable description of when the tool
                                    should be used.
                                parameters:
                                  type: object
                                  description: >-
                                    JSON Schema object describing function
                                    arguments.
                              required:
                                - type
                                - name
                          response_defaults:
                            type: object
                            description: >-
                              Default response-generation settings for later
                              turns.
                            properties:
                              temperature:
                                type: number
                                description: >-
                                  Sampling temperature for generated response
                                  text. Supported range is `0` to `2`.
                              max_output_tokens:
                                type: integer
                                description: Maximum generated text tokens for a response.
                          turn_detection:
                            type: object
                            description: Default turn handling for realtime user input.
                            properties:
                              type:
                                type: string
                                description: >-
                                  `manual` means the client starts responses
                                  explicitly. `server_vad` means the server
                                  detects turn boundaries from incoming user
                                  audio.
                              silence_duration_ms:
                                type: integer
                                description: >-
                                  Compatibility field. Do not use it to tune
                                  actual end-of-utterance timing. See [Speech
                                  recognition
                                  configuration](/streaming-avatar/integrate/speech-recognition)
                                  for recognition and turn-detection settings.
                            required:
                              - type
                          input_audio_transcription:
                            type: object
                            description: >-
                              Default transcription settings for user audio
                              input.
                            properties:
                              enabled:
                                type: boolean
                                description: Whether user audio should be transcribed.
                              language:
                                type: string
                                description: >-
                                  BCP-47 `language` code, or `auto` for
                                  automatic `language` detection.
                      avatars:
                        type: array
                        description: >-
                          Effective avatar definitions in the session. Same
                          structure and meaning as the `avatars` request field.
                        items:
                          type: object
                          properties:
                            avatar_id:
                              type: string
                              description: >-
                                Caller-defined avatar identifier, unique within
                                the session.
                            instructions:
                              type: string
                              description: >-
                                The avatar’s identity, speaking style, and
                                response rules. These instructions apply only
                                when this avatar is selected by the current
                                `session.active_avatar_id`, and they do not
                                grant tools or data access.
                            voice:
                              type: object
                              description: >-
                                Optional per-avatar TTS configuration for
                                multi-avatar sessions that require different
                                voices. Provide `provider`, `provider_voice_id`,
                                and `model`. For the normal session-wide voice
                                path, omit this object and configure
                                `pipeline_config.tts_config`. A session TTS
                                configuration takes precedence over every
                                per-avatar voice.
                              properties:
                                provider:
                                  type: string
                                  description: >-
                                    TTS backend, such as
                                    `qwen-audio-3.0-tts-flash_ws`, `elevenlabs`,
                                    or `qwen3tts`. Required with
                                    `provider_voice_id` and `model`.
                                provider_voice_id:
                                  type: string
                                  description: >-
                                    Provider-side voice or style id. Required
                                    with `provider` and `model`.
                                model:
                                  type: string
                                  description: >-
                                    Provider TTS model id. Required with
                                    `provider` and provider_voice_id.
                                speed:
                                  type: number
                                  description: >-
                                    Speech speed multiplier. The valid range is
                                    `0.7` to `1.2`.
                              required:
                                - provider
                                - provider_voice_id
                                - model
                            visual:
                              type: object
                              description: >-
                                Visual rendering settings. Required when this
                                avatar is used for video rendering.
                              properties:
                                source_images:
                                  type: array
                                  description: >-
                                    Registered source images available for this
                                    avatar. Required when image-based avatar
                                    video rendering is used.
                                  items:
                                    type: object
                                    description: >-
                                      One source image that can be selected for
                                      avatar video rendering.
                                    properties:
                                      source_image_id:
                                        type: string
                                        description: >-
                                          Caller-defined source image identifier,
                                          unique within the avatar.
                                      url:
                                        type: string
                                        description: >-
                                          Accessible image URL. Required for a
                                          usable video source image and for each
                                          source image registered through
                                          `avatar.add`. A text-only avatar does
                                          not need a video source image.
                                      description:
                                        type: string
                                        description: >-
                                          Optional description of the person,
                                          pose, and scene.
                                      media_type:
                                        type: string
                                        description: >-
                                          Image media type, such as `image/png` or
                                          `image/jpeg`.
                                    required:
                                      - source_image_id
                                      - media_type
                                default_source_image_id:
                                  type: string
                                  description: >-
                                    With one usable source image, this field can
                                    be omitted. With multiple images and no
                                    explicit `source_image_id`, use it to select
                                    the initial image.
                              required:
                                - source_images
                          required:
                            - avatar_id
                            - visual
                      delivery:
                        type: object
                        description: >-
                          Media connection details returned by the server.
                          Present when `session.mode` is `video_avatar`.
                        properties:
                          media:
                            type: object
                            description: >-
                              Temporary media connection details. Present when
                              avatar video media is available. Exactly one media
                              variant is returned, selected by `transport`.
                              These details are scoped to this session and are
                              not `provider` account credentials.
                            properties:
                              transport:
                                type: string
                                description: >-
                                  Discriminator that selects which media variant
                                  is returned. Supported values are `trtc` and
                                  `agora`.
                              trtc:
                                type: object
                                description: >-
                                  TRTC room join details for the `trtc` media
                                  variant.
                                properties:
                                  sdk_app_id:
                                    type: string
                                    description: >-
                                      TRTC application id used by the client
                                      SDK. Returned as a string, consistent with
                                      the convention that 64-bit integers may be
                                      returned as strings.
                                  room_id:
                                    type: string
                                    description: TRTC room id for this session.
                                  user_id:
                                    type: string
                                    description: TRTC user id assigned to the client.
                                  user_sig:
                                    type: string
                                    description: >-
                                      Session-scoped signature for joining the
                                      TRTC room.
                                  publisher_user_id:
                                    type: string
                                    description: >-
                                      User id of the avatar video publisher in
                                      the TRTC room. Subscribe only to this user
                                      id to receive the avatar stream.
                                required:
                                  - sdk_app_id
                                  - room_id
                                  - user_id
                                  - user_sig
                                  - publisher_user_id
                              agora:
                                type: object
                                description: >-
                                  Agora channel join details for the `agora`
                                  media variant.
                                properties:
                                  app_id:
                                    type: string
                                    description: >-
                                      Agora application id used by the client
                                      SDK.
                                  channel_name:
                                    type: string
                                    description: Agora channel name for this session.
                                  token:
                                    type: string
                                    description: >-
                                      Session-scoped token for joining the Agora
                                      channel.
                                  user_id:
                                    type: string
                                    description: Agora user id assigned to the client.
                                  publisher_user_id:
                                    type: string
                                    description: >-
                                      User id of the avatar video publisher in
                                      the Agora channel. Subscribe only to this
                                      user id to receive the avatar stream.
                                required:
                                  - app_id
                                  - channel_name
                                  - token
                                  - user_id
                                  - publisher_user_id
                            required:
                              - transport
                      model:
                        type: string
                        description: The model from the request.
                      expires_at:
                        type: string
                        format: date-time
                        description: >-
                          Time at which the session and its control credential
                          expire. Present when the server returns a fixed
                          expiration; otherwise omitted.
                    required:
                      - session_id
                      - status
                      - auto_close
                      - region
                      - session
                      - control
                      - conversation
                      - avatars
                      - model
              example:
                code: 0
                message: success
                data:
                  session_id: avatar-session-123
                  status: active
                  auto_close:
                    disconnected_timeout_seconds: 60
                    interaction_idle_timeout_seconds: 300
                  region: us-west
                  session:
                    mode: video_avatar
                    active_avatar_id: host_a
                    source_image_id: front
                  control:
                    url: >-
                      wss://api.vivix.ai/v1/realtime-avatar/control?session_id=avatar-session-123
                    client_secret: ctl_...
                  output:
                    aspect_ratio: '9:16'
                    resolution: 480p
                    fps: 24
                  assets:
                    audio:
                      - asset_id: welcome_jingle
                        url: https://cdn.example.com/audio/welcome-jingle.wav
                        mime_type: audio/wav
                    speech_text:
                      - asset_id: welcome_line
                        text: Welcome back. I saved your place.
                    visual_prompts:
                      - asset_id: wave_small
                        prompt: smile and give a small wave
                  conversation:
                    instructions: Answer as a concise livestream host.
                    tools:
                      - type: function
                        name: lookup_product
                        description: Look up product details by product id.
                        parameters:
                          type: object
                          properties:
                            product_id:
                              type: string
                          required:
                            - product_id
                    response_defaults:
                      temperature: 0.7
                      max_output_tokens: 512
                    turn_detection:
                      type: server_vad
                      silence_duration_ms: 500
                    input_audio_transcription:
                      enabled: true
                      language: auto
                  avatars:
                    - avatar_id: host_a
                      instructions: >-
                        You are Host A, a warm livestream host who gives concise
                        product-focused replies.
                      visual:
                        source_images:
                          - source_image_id: front
                            url: https://cdn.example.com/avatars/host-a-front.png
                            description: >-
                              A woman facing the camera in a brightly lit
                              studio, front view.
                            media_type: image/png
                          - source_image_id: side
                            url: https://cdn.example.com/avatars/host-a-side.png
                            description: The same woman in the same studio, side view.
                            media_type: image/png
                        default_source_image_id: front
                    - avatar_id: host_b
                      instructions: >-
                        You are Host B, a calm expert who gives brief technical
                        explanations.
                      visual:
                        source_images:
                          - source_image_id: front
                            url: https://cdn.example.com/avatars/host-b-front.png
                            description: >-
                              A man facing the camera in a brightly lit studio,
                              front view.
                            media_type: image/png
                        default_source_image_id: front
                  delivery:
                    media:
                      transport: trtc
                      trtc:
                        sdk_app_id: '1400000000'
                        room_id: avatar-session-123
                        user_id: viewer
                        user_sig: ...
                        publisher_user_id: publisher_xxx
                  model: vivix-a1-stream
                  expires_at: '2026-07-16T10:05:00Z'
components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: Your Vivix API key. Keep it on your server.

````