> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Update Character

> **PUT replaces the whole record; it is not PATCH.** The body matches create (including `name`). Keys omitted from the request are deleted from the stored config. The success `data` is the character after the write, so you can check what was stored.

<AccordionGroup>
  <Accordion title="Errors">
    Error responses use the same `{code, message, data}` envelope; a non-zero `code` identifies the error.

    | Case | Error code |
    | - | - |
    | The character does not exist, is deleted, or is not in the current workspace | `20005 not found` |
    | Other validation failures | Same as Create Character. |
  </Accordion>
</AccordionGroup>


## OpenAPI

````yaml streaming-avatar/api-references/openapi.json PUT /v1/characters/{character_id}
openapi: 3.1.0
info:
  title: Vivix Streaming Avatar API
  version: 1.0.0
  description: REST endpoints for Streaming Avatar sessions, voices, and characters.
servers:
  - url: https://api.vivix.ai
security:
  - bearerAuth: []
paths:
  /v1/characters/{character_id}:
    put:
      summary: Update Character
      description: >-
        **PUT replaces the whole record; it is not PATCH.** The body matches
        create (including `name`). Keys omitted from the request are deleted
        from the stored config. The success `data` is the character after the
        write, so you can check what was stored.
      operationId: updateCharacter
      parameters:
        - name: character_id
          in: path
          required: true
          schema:
            type: string
          description: >-
            Character id to replace.


            Body parameters match Create Character. Validation and moderation
            rules are the same.
          example: chr_a1b2c3d4e5f6789012345678abcdef01
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                name:
                  type: string
                  description: Display name for lists.
                model:
                  type: string
                  description: >-
                    Model id. Use `vivix-a1-stream`, or `vivix-a1-stream-lite`
                    for the lighter variant. See
                    [Models](https://platform.vivix.ai/doc/overview/models).
                region:
                  type: string
                  description: >-
                    Optional deployment region hint. Returned back on the
                    session object; may be empty if not provided. It does not
                    currently affect scheduling or guarantee data residency.
                session:
                  type: object
                  description: >-
                    Initial mutable session state: mode, active avatar, and the
                    source image used for video. Optional — with a single avatar
                    the defaults below already resolve. Change these values
                    after the session starts with `session.update` on the
                    control channel.
                  properties:
                    mode:
                      type: string
                      description: >-
                        Initial session mode. Supported values are `text_chat`
                        and `video_avatar`. Defaults to `video_avatar`.
                        `text_chat` returns text only, and `video_avatar`
                        returns avatar video with generated speech and text
                        events.
                    active_avatar_id:
                      type: string
                      description: >-
                        Avatar whose instructions are active when the session
                        starts. In `video_avatar` mode, this is also the avatar
                        rendered in the video stream. If omitted and exactly one
                        avatar is provided, that avatar is used; if omitted with
                        multiple avatars, the request fails.
                    source_image_id:
                      type: string
                      description: >-
                        Source image to use for avatar video. Used only when
                        `mode` is `video_avatar`. If omitted in `video_avatar`,
                        uses the active avatar’s
                        `visual.default_source_image_id`. If no default is set
                        and that avatar has exactly one source image, it is
                        selected automatically. In video mode, the request fails
                        if no usable image can be resolved. In `text_chat`, the
                        field is omitted from the REST response.
                output:
                  type: object
                  description: >-
                    Create-time avatar video output settings. Required for
                    `video_avatar` mode and ignored in `text_chat` mode. See
                    [output](/streaming-avatar/character/spoken-instructions).
                  properties:
                    aspect_ratio:
                      type: string
                      description: >-
                        Requested avatar video aspect ratio, such as `9:16`,
                        `1:1`, or `16:9`.
                    resolution:
                      type: string
                      description: >-
                        Requested avatar video resolution. Supported values are
                        `480p` and `720p`.
                avatars:
                  type: array
                  description: >-
                    Avatars available in the session. Must not be empty, and
                    `avatar_id` values must not repeat.
                  items:
                    type: object
                    properties:
                      avatar_id:
                        type: string
                        description: >-
                          Caller-defined avatar identifier, unique within the
                          session.
                      instructions:
                        type: string
                        description: >-
                          The avatar’s identity, speaking style, and response
                          rules. These instructions apply only when this avatar
                          is selected by the current `session.active_avatar_id`,
                          and they do not grant tools or data access.
                      voice:
                        type: object
                        description: >-
                          Optional per-avatar TTS configuration for multi-avatar
                          sessions that require different voices. Provide
                          `provider`, `provider_voice_id`, and `model`. For the
                          normal session-wide voice path, omit this object and
                          configure `pipeline_config.tts_config`. A session TTS
                          configuration takes precedence over every per-avatar
                          voice.
                        properties:
                          provider:
                            type: string
                            description: >-
                              TTS backend, such as
                              `qwen-audio-3.0-tts-flash_ws`, `elevenlabs`, or
                              `qwen3tts`. Required with `provider_voice_id` and
                              `model`.
                          provider_voice_id:
                            type: string
                            description: >-
                              Provider-side voice or style id. Required with
                              `provider` and `model`.
                          model:
                            type: string
                            description: >-
                              Provider TTS model id. Required with `provider`
                              and provider_voice_id.
                          speed:
                            type: number
                            description: >-
                              Speech speed multiplier. The valid range is `0.7`
                              to `1.2`.
                        required:
                          - provider
                          - provider_voice_id
                          - model
                      visual:
                        type: object
                        description: >-
                          Visual rendering settings. Required when this avatar
                          is used for video rendering.
                        properties:
                          source_images:
                            type: array
                            description: >-
                              Registered source images available for this
                              avatar. Required when image-based avatar video
                              rendering is used.
                            items:
                              type: object
                              description: >-
                                One source image that can be selected for avatar
                                video rendering.
                              properties:
                                source_image_id:
                                  type: string
                                  description: >-
                                    Caller-defined source image identifier,
                                    unique within the avatar.
                                url:
                                  type: string
                                  description: >-
                                    Accessible image URL. Required for a usable
                                    video source image and for each source image
                                    registered through `avatar.add`. A text-only
                                    avatar does not need a video source image.
                                description:
                                  type: string
                                  description: >-
                                    Optional description of the person, pose,
                                    and scene.
                                media_type:
                                  type: string
                                  description: >-
                                    Image media type, such as `image/png` or
                                    `image/jpeg`.
                              required:
                                - source_image_id
                                - media_type
                          default_source_image_id:
                            type: string
                            description: >-
                              With one usable source image, this field can be
                              omitted. With multiple images and no explicit
                              `source_image_id`, use it to select the initial
                              image.
                        required:
                          - source_images
                    required:
                      - avatar_id
                      - visual
                pipeline_config:
                  type: object
                  description: >-
                    Session-wide ASR, TTS, LLM, and motion settings. Omit
                    overrides to use platform behavior. Use `tts_config` for the
                    voice and `avatars[].instructions` for identity and response
                    style. `motion_enhanced` and `motion_planner` are optional
                    advanced overrides. These settings are fixed when a session
                    is created.
                  properties:
                    tag_audio_enabled:
                      type: boolean
                      description: >-
                        Optional; defaults to false. Enables faster audio output
                        for an eligible first short speech segment of the
                        response, not emotion tags. Supply a boolean, not a
                        string. See [opening
                        acknowledgements](/streaming-avatar/interaction/acknowledgement).
                    tts_config:
                      type: object
                      properties:
                        tts_voice_id:
                          type: string
                          description: >-
                            Required nonempty `provider` voice ID or Vivix
                            cloned voice_id.
                        tts_provider:
                          type: string
                          description: >-
                            Optional. For built-in Qwen Audio voices, defaults
                            to `qwen-audio-3.0-tts-flash_ws`. A Vivix cloned
                            `voice_id` resolves to its associated speech
                            service. Set the matching value for other services.
                        tts_model_id:
                          type: string
                          description: >-
                            Optional. For built-in Qwen Audio voices, defaults
                            to `qwen-audio-3.0-tts-flash`. A Vivix cloned
                            `voice_id` resolves to its associated model. Set the
                            matching value for other services.
                        tts_endpoint:
                          type: string
                          description: >-
                            Optional endpoint override when nonempty. Not needed
                            for built-in voices.
                        tts_api_key:
                          type: string
                          description: >-
                            Optional `provider` credential when nonempty.
                            Built-in voices need no separate key.
                        payload:
                          type: object
                          description: >-
                            Optional speech-service extensions; arbitrary
                            `provider` parameters are not guaranteed to take
                            effect. Do not duplicate reserved authentication or
                            routing fields here.
                      required:
                        - tts_voice_id
                      description: >-
                        Session-wide TTS configuration. It takes precedence over
                        every `avatars[].voice`.
                    asr_config:
                      type: object
                      properties:
                        provider:
                          type: string
                          description: >-
                            Required when `asr_config` is supplied. Common
                            values are `doubao` and nova-3. The platform
                            supplies a default service when the whole object is
                            omitted, but not when only `language` is provided.
                        language:
                          type: string
                          description: >-
                            Optional. Examples: `zh-CN` or `en-US` for Doubao;
                            en for Nova. Controls recognition `language`, not
                            the reply language.
                        eou_timeout_ms:
                          type: string
                          description: >-
                            Optional end-of-utterance judge timeout, for example
                            "1500". Not a fixed silence length or total response
                            latency.
                      required:
                        - provider
                      description: >-
                        Advanced override. Omit it to use the platform default,
                        `doubao`. Set `provider` (`doubao` for multilingual, or
                        `nova-3` for English-only) and optional `language`. See
                        [`asr_config`](/streaming-avatar/integrate/speech-recognition).
                    llm_config:
                      type: object
                      properties:
                        llm_backend:
                          type: string
                          description: Optional; only `openai` is supported.
                        llm_base_url:
                          type: string
                          description: >-
                            Optional OpenAI-compatible endpoint. Supply with
                            llm_model for a custom service; omit when only
                            tuning generation settings.
                        llm_model:
                          type: string
                          description: Optional custom model name.
                        llm_api_key:
                          type: string
                          description: >-
                            Optional service credential. Keep private
                            credentials out of browser code and public
                            configuration.
                        llm_temperature:
                          type: number
                          description: >-
                            Optional generation temperature; use the range
                            supported by the model.
                        llm_max_output_tokens:
                          type: integer
                          description: Optional generated-token limit, not a word count.
                        llm_extra_payload:
                          type: string
                          description: >-
                            Optional JSON-encoded object string; do not supply a
                            nested object.
                      description: >-
                        `llm_config` · optional object. Set `llm_base_url` and
                        `llm_model` for a custom service, or override only
                        generation parameters. See [Text model
                        configuration](/streaming-avatar/integrate/dialogue-model)
                        for field types and examples.
                    motion_enhanced:
                      type: object
                      description: >-
                        `prompt` drives the motion generated while the avatar is
                        speaking. See
                        [motion_enhanced](/streaming-avatar/character/speaking-motion).
                      properties:
                        prompt:
                          type: string
                          description: >-
                            Motion prompt. See the linked guide for how to write
                            it.
                    motion_planner:
                      type: object
                      description: >-
                        `prompt` drives listening (idle) motion. The key is
                        `motion_planner`, not `idle_motion`. See
                        [motion_planner](/streaming-avatar/character/listening-motion).
                      properties:
                        prompt:
                          type: string
                          description: >-
                            Motion prompt. See the linked guide for how to write
                            it.
                  required:
                    - tts_config
                conversation:
                  type: object
                  description: >-
                    Session-level conversation defaults: `tools`,
                    `turn_detection`, and `input_audio_transcription`. See
                    [conversation](/streaming-avatar/character/spoken-instructions).
                  properties:
                    instructions:
                      type: string
                      description: >-
                        Default session and task instructions for generated
                        responses. Effective instructions are composed from the
                        selected avatar instructions, these conversation
                        instructions, and any per-response instructions;
                        conflicts are resolved in the order
                        `response.instructions`, `conversation.instructions`,
                        then `avatars[].instructions`.
                    tools:
                      type: array
                      description: Tools available to responses in the session.
                      items:
                        type: object
                        description: >-
                          One tool definition. Only function tools are
                          documented for v1.
                        properties:
                          type:
                            type: string
                            description: Use `function`.
                          name:
                            type: string
                            description: Function name available to the response planner.
                          description:
                            type: string
                            description: >-
                              Human-readable description of when the tool should
                              be used.
                          parameters:
                            type: object
                            description: JSON Schema object describing function arguments.
                        required:
                          - type
                          - name
                    response_defaults:
                      type: object
                      description: Default response-generation settings for later turns.
                      properties:
                        temperature:
                          type: number
                          description: >-
                            Sampling temperature for generated response text.
                            Supported range is `0` to `2`.
                        max_output_tokens:
                          type: integer
                          description: Maximum generated text tokens for a response.
                    turn_detection:
                      type: object
                      description: Default turn handling for realtime user input.
                      properties:
                        type:
                          type: string
                          description: >-
                            `manual` means the client starts responses
                            explicitly. `server_vad` means the server detects
                            turn boundaries from incoming user audio.
                        silence_duration_ms:
                          type: integer
                          description: >-
                            Compatibility field. Do not use it to tune actual
                            end-of-utterance timing. See [Speech recognition
                            configuration](/streaming-avatar/integrate/speech-recognition)
                            for recognition and turn-detection settings.
                      required:
                        - type
                    input_audio_transcription:
                      type: object
                      description: Default transcription settings for user audio input.
                      properties:
                        enabled:
                          type: boolean
                          description: Whether user audio should be transcribed.
                        language:
                          type: string
                          description: >-
                            BCP-47 `language` code, or `auto` for automatic
                            `language` detection.
                assets:
                  type: object
                  description: >-
                    Session-scoped reusable assets for `response.script`. Asset
                    ids must be unique across `audio`, `speech_text`, and
                    `visual_prompts`. Avatar source images stay under
                    `avatars[].visual.source_images`. On the REST side the
                    visual prompt text field is `prompt`; on the WSS `asset.add`
                    side the same asset uses `text` — the two shapes differ.
                  properties:
                    audio:
                      type: array
                      description: >-
                        Reusable audio files that can be referenced by scripted
                        vocal output.
                      items:
                        type: object
                        description: One reusable audio asset.
                        properties:
                          asset_id:
                            type: string
                            description: Caller-defined audio asset id.
                          url:
                            type: string
                            description: >-
                              HTTPS URL for the audio file. The URL must be
                              directly fetchable by Vivix without custom request
                              headers.
                          mime_type:
                            type: string
                            description: >-
                              Audio MIME type. Supported values are `audio/wav`
                              and `audio/mpeg`.
                        required:
                          - asset_id
                          - url
                    speech_text:
                      type: array
                      description: >-
                        Reusable exact speech text that can be referenced by
                        scripted vocal output.
                      items:
                        type: object
                        description: One reusable speech text asset.
                        properties:
                          asset_id:
                            type: string
                            description: Caller-defined speech text asset id.
                          text:
                            type: string
                            description: >-
                              Exact text to synthesize and stream through output
                              text events when used in `response.script`.
                        required:
                          - asset_id
                          - text
                    visual_prompts:
                      type: array
                      description: >-
                        Reusable visual instructions that can be referenced by
                        scripted visual output.
                      items:
                        type: object
                        description: One reusable visual prompt asset.
                        properties:
                          asset_id:
                            type: string
                            description: Caller-defined visual prompt asset id.
                          prompt:
                            type: string
                            description: >-
                              Instruction for avatar movement, posture, gaze,
                              gesture, or presentation.
                        required:
                          - asset_id
                          - prompt
                delivery:
                  type: object
                  description: >-
                    Media delivery settings used when `session.mode` is
                    `video_avatar`. If omitted, the media transport defaults to
                    TRTC. It is not used for `text_chat`.
                  properties:
                    media:
                      type: object
                      description: >-
                        Selects how the client sends and receives realtime
                        media.
                      properties:
                        transport:
                          type: string
                          description: >-
                            Media transport requested for the session. Use
                            `trtc` for TRTC or `agora` for Agora. If omitted, it
                            defaults to TRTC. The response returns the
                            corresponding provider-specific join details under
                            `delivery.media.trtc` or `delivery.media.agora`.
                max_duration_seconds:
                  type: integer
                  description: >-
                    Maximum session duration in seconds. If omitted, the default
                    session length applies. An explicit value must be greater
                    than `3` (minimum 4).
                auto_close:
                  type: object
                  description: >-
                    Automatic close policy fixed when the session is created.
                    The response returns the effective values after platform
                    defaults are applied.
                  properties:
                    disconnected_timeout_seconds:
                      type: integer
                      description: >-
                        Closes after all control WebSocket connections have been
                        absent for this period, including when none was ever
                        opened. Omission uses the platform setting; read the
                        effective value in the response. Use `0` to disable, or
                        an integer from `10` to `3600` to enable.


                        seconds. `0` disables the rule; nonzero values must be
                        `10`–`3600`. Omission uses the platform setting.
                    interaction_idle_timeout_seconds:
                      type: integer
                      description: >-
                        Closes the session after this duration without a
                        successfully accepted client interaction or detected
                        user speech. Accepted business events and user speech
                        reset the timeout; WSS Ping/Pong and server events do
                        not. Omitted or 0 disables this policy; explicit
                        non-zero values must be 30–3600.


                        An explicit non-zero timeout cannot exceed the effective
                        maximum session duration.


                        seconds. Omitted or `0` disables the rule; nonzero
                        values must be `30`–`3600`.
                recording_mode:
                  type: string
                  description: '`on` or `off`. Send `off` to disable recording.'
              required:
                - name
                - model
                - output
                - avatars
            example:
              name: Studio Host v2
              model: vivix-a1-stream
              output:
                aspect_ratio: '9:16'
                resolution: 720p
              pipeline_config:
                tts_config:
                  tts_provider: qwen-audio-3.0-tts-flash_ws
                  tts_voice_id: longanhuan_v3.6
                  tts_model_id: qwen-audio-3.0-tts-flash
              avatars:
                - avatar_id: host_a
                  instructions: You are Host A, a warm livestream host.
                  visual:
                    default_source_image_id: front
                    source_images:
                      - source_image_id: front
                        url: https://cdn.example.com/avatars/host-a-front.png
                        description: >-
                          A woman facing the camera in a brightly lit studio,
                          front view.
                        media_type: image/png
      responses:
        '200':
          description: Success
          content:
            application/json:
              schema:
                type: object
                required:
                  - code
                  - message
                  - data
                properties:
                  code:
                    type: integer
                    description: '`0` on success; any other value is an error code.'
                  message:
                    type: string
                    description: '`success`, or an English error description.'
                  data:
                    type: object
                    description: >-
                      The saved character: metadata plus the entire stored
                      configuration. Source image URLs may be rewritten during
                      saving.
                    properties:
                      character_id:
                        type: string
                        description: >-
                          Character id, shaped like `chr_…`. Use it to start a
                          session from [Sessions
                          API](/streaming-avatar/api-references/sessions).
                      name:
                        type: string
                        description: Display name.
                      created_at:
                        type: string
                        format: date-time
                        description: UTC, formatted as `YYYY-MM-DDTHH:MM:SSZ`.
                      updated_at:
                        type: string
                        format: date-time
                        description: >-
                          UTC, formatted as `YYYY-MM-DDTHH:MM:SSZ`.


                          **stored config: object**


                          The saved configuration fields are returned alongside
                          metadata under `data`. `name` is stored as metadata;
                          `character_id` and `idempotency_key` are not saved as
                          session configuration. Source image URLs may be
                          rewritten during saving.
                      model:
                        type: string
                        description: >-
                          Model id. Use `vivix-a1-stream`, or
                          `vivix-a1-stream-lite` for the lighter variant. See
                          [Models](https://platform.vivix.ai/doc/overview/models).
                      region:
                        type: string
                        description: >-
                          Optional deployment region hint. Returned back on the
                          session object; may be empty if not provided. It does
                          not currently affect scheduling or guarantee data
                          residency.
                      session:
                        type: object
                        description: >-
                          Initial mutable session state: mode, active avatar,
                          and the source image used for video. Optional — with a
                          single avatar the defaults below already resolve.
                          Change these values after the session starts with
                          `session.update` on the control channel.
                        properties:
                          mode:
                            type: string
                            description: >-
                              Initial session mode. Supported values are
                              `text_chat` and `video_avatar`. Defaults to
                              `video_avatar`. `text_chat` returns text only, and
                              `video_avatar` returns avatar video with generated
                              speech and text events.
                          active_avatar_id:
                            type: string
                            description: >-
                              Avatar whose instructions are active when the
                              session starts. In `video_avatar` mode, this is
                              also the avatar rendered in the video stream. If
                              omitted and exactly one avatar is provided, that
                              avatar is used; if omitted with multiple avatars,
                              the request fails.
                          source_image_id:
                            type: string
                            description: >-
                              Source image to use for avatar video. Used only
                              when `mode` is `video_avatar`. If omitted in
                              `video_avatar`, uses the active avatar’s
                              `visual.default_source_image_id`. If no default is
                              set and that avatar has exactly one source image,
                              it is selected automatically. In video mode, the
                              request fails if no usable image can be resolved.
                              In `text_chat`, the field is omitted from the REST
                              response.
                      output:
                        type: object
                        description: >-
                          Create-time avatar video output settings. Required for
                          `video_avatar` mode and ignored in `text_chat` mode.
                          See
                          [output](/streaming-avatar/character/spoken-instructions).
                        properties:
                          aspect_ratio:
                            type: string
                            description: >-
                              Requested avatar video aspect ratio, such as
                              `9:16`, `1:1`, or `16:9`.
                          resolution:
                            type: string
                            description: >-
                              Requested avatar video resolution. Supported
                              values are `480p` and `720p`.
                      avatars:
                        type: array
                        description: >-
                          Avatars available in the session. Must not be empty,
                          and `avatar_id` values must not repeat.
                        items:
                          type: object
                          properties:
                            avatar_id:
                              type: string
                              description: >-
                                Caller-defined avatar identifier, unique within
                                the session.
                            instructions:
                              type: string
                              description: >-
                                The avatar’s identity, speaking style, and
                                response rules. These instructions apply only
                                when this avatar is selected by the current
                                `session.active_avatar_id`, and they do not
                                grant tools or data access.
                            voice:
                              type: object
                              description: >-
                                Optional per-avatar TTS configuration for
                                multi-avatar sessions that require different
                                voices. Provide `provider`, `provider_voice_id`,
                                and `model`. For the normal session-wide voice
                                path, omit this object and configure
                                `pipeline_config.tts_config`. A session TTS
                                configuration takes precedence over every
                                per-avatar voice.
                              properties:
                                provider:
                                  type: string
                                  description: >-
                                    TTS backend, such as
                                    `qwen-audio-3.0-tts-flash_ws`, `elevenlabs`,
                                    or `qwen3tts`. Required with
                                    `provider_voice_id` and `model`.
                                provider_voice_id:
                                  type: string
                                  description: >-
                                    Provider-side voice or style id. Required
                                    with `provider` and `model`.
                                model:
                                  type: string
                                  description: >-
                                    Provider TTS model id. Required with
                                    `provider` and provider_voice_id.
                                speed:
                                  type: number
                                  description: >-
                                    Speech speed multiplier. The valid range is
                                    `0.7` to `1.2`.
                              required:
                                - provider
                                - provider_voice_id
                                - model
                            visual:
                              type: object
                              description: >-
                                Visual rendering settings. Required when this
                                avatar is used for video rendering.
                              properties:
                                source_images:
                                  type: array
                                  description: >-
                                    Registered source images available for this
                                    avatar. Required when image-based avatar
                                    video rendering is used.
                                  items:
                                    type: object
                                    description: >-
                                      One source image that can be selected for
                                      avatar video rendering.
                                    properties:
                                      source_image_id:
                                        type: string
                                        description: >-
                                          Caller-defined source image identifier,
                                          unique within the avatar.
                                      url:
                                        type: string
                                        description: >-
                                          Accessible image URL. Required for a
                                          usable video source image and for each
                                          source image registered through
                                          `avatar.add`. A text-only avatar does
                                          not need a video source image.
                                      description:
                                        type: string
                                        description: >-
                                          Optional description of the person,
                                          pose, and scene.
                                      media_type:
                                        type: string
                                        description: >-
                                          Image media type, such as `image/png` or
                                          `image/jpeg`.
                                    required:
                                      - source_image_id
                                      - media_type
                                default_source_image_id:
                                  type: string
                                  description: >-
                                    With one usable source image, this field can
                                    be omitted. With multiple images and no
                                    explicit `source_image_id`, use it to select
                                    the initial image.
                              required:
                                - source_images
                          required:
                            - avatar_id
                            - visual
                      pipeline_config:
                        type: object
                        description: >-
                          Session-wide ASR, TTS, LLM, and motion settings. Omit
                          overrides to use platform behavior. Use `tts_config`
                          for the voice and `avatars[].instructions` for
                          identity and response style. `motion_enhanced` and
                          `motion_planner` are optional advanced overrides.
                          These settings are fixed when a session is created.
                        properties:
                          tag_audio_enabled:
                            type: boolean
                            description: >-
                              Optional; defaults to false. Enables faster audio
                              output for an eligible first short speech segment
                              of the response, not emotion tags. Supply a
                              boolean, not a string. See [opening
                              acknowledgements](/streaming-avatar/interaction/acknowledgement).
                          tts_config:
                            type: object
                            properties:
                              tts_voice_id:
                                type: string
                                description: >-
                                  Required nonempty `provider` voice ID or Vivix
                                  cloned voice_id.
                              tts_provider:
                                type: string
                                description: >-
                                  Optional. For built-in Qwen Audio voices,
                                  defaults to `qwen-audio-3.0-tts-flash_ws`. A
                                  Vivix cloned `voice_id` resolves to its
                                  associated speech service. Set the matching
                                  value for other services.
                              tts_model_id:
                                type: string
                                description: >-
                                  Optional. For built-in Qwen Audio voices,
                                  defaults to `qwen-audio-3.0-tts-flash`. A
                                  Vivix cloned `voice_id` resolves to its
                                  associated model. Set the matching value for
                                  other services.
                              tts_endpoint:
                                type: string
                                description: >-
                                  Optional endpoint override when nonempty. Not
                                  needed for built-in voices.
                              tts_api_key:
                                type: string
                                description: >-
                                  Optional `provider` credential when nonempty.
                                  Built-in voices need no separate key.
                              payload:
                                type: object
                                description: >-
                                  Optional speech-service extensions; arbitrary
                                  `provider` parameters are not guaranteed to
                                  take effect. Do not duplicate reserved
                                  authentication or routing fields here.
                            required:
                              - tts_voice_id
                            description: >-
                              Session-wide TTS configuration. It takes
                              precedence over every `avatars[].voice`.
                          asr_config:
                            type: object
                            properties:
                              provider:
                                type: string
                                description: >-
                                  Required when `asr_config` is supplied. Common
                                  values are `doubao` and nova-3. The platform
                                  supplies a default service when the whole
                                  object is omitted, but not when only
                                  `language` is provided.
                              language:
                                type: string
                                description: >-
                                  Optional. Examples: `zh-CN` or `en-US` for
                                  Doubao; en for Nova. Controls recognition
                                  `language`, not the reply language.
                              eou_timeout_ms:
                                type: string
                                description: >-
                                  Optional end-of-utterance judge timeout, for
                                  example "1500". Not a fixed silence length or
                                  total response latency.
                            required:
                              - provider
                            description: >-
                              Advanced override. Omit it to use the platform
                              default, `doubao`. Set `provider` (`doubao` for
                              multilingual, or `nova-3` for English-only) and
                              optional `language`. See
                              [`asr_config`](/streaming-avatar/integrate/speech-recognition).
                          llm_config:
                            type: object
                            properties:
                              llm_backend:
                                type: string
                                description: Optional; only `openai` is supported.
                              llm_base_url:
                                type: string
                                description: >-
                                  Optional OpenAI-compatible endpoint. Supply
                                  with llm_model for a custom service; omit when
                                  only tuning generation settings.
                              llm_model:
                                type: string
                                description: Optional custom model name.
                              llm_api_key:
                                type: string
                                description: >-
                                  Optional service credential. Keep private
                                  credentials out of browser code and public
                                  configuration.
                              llm_temperature:
                                type: number
                                description: >-
                                  Optional generation temperature; use the range
                                  supported by the model.
                              llm_max_output_tokens:
                                type: integer
                                description: >-
                                  Optional generated-token limit, not a word
                                  count.
                              llm_extra_payload:
                                type: string
                                description: >-
                                  Optional JSON-encoded object string; do not
                                  supply a nested object.
                            description: >-
                              `llm_config` · optional object. Set `llm_base_url`
                              and `llm_model` for a custom service, or override
                              only generation parameters. See [Text model
                              configuration](/streaming-avatar/integrate/dialogue-model)
                              for field types and examples.
                          motion_enhanced:
                            type: object
                            description: >-
                              `prompt` drives the motion generated while the
                              avatar is speaking. See
                              [motion_enhanced](/streaming-avatar/character/speaking-motion).
                            properties:
                              prompt:
                                type: string
                                description: >-
                                  Motion prompt. See the linked guide for how to
                                  write it.
                          motion_planner:
                            type: object
                            description: >-
                              `prompt` drives listening (idle) motion. The key
                              is `motion_planner`, not `idle_motion`. See
                              [motion_planner](/streaming-avatar/character/listening-motion).
                            properties:
                              prompt:
                                type: string
                                description: >-
                                  Motion prompt. See the linked guide for how to
                                  write it.
                        required:
                          - tts_config
                      conversation:
                        type: object
                        description: >-
                          Session-level conversation defaults: `tools`,
                          `turn_detection`, and `input_audio_transcription`. See
                          [conversation](/streaming-avatar/character/spoken-instructions).
                        properties:
                          instructions:
                            type: string
                            description: >-
                              Default session and task instructions for
                              generated responses. Effective instructions are
                              composed from the selected avatar instructions,
                              these conversation instructions, and any
                              per-response instructions; conflicts are resolved
                              in the order `response.instructions`,
                              `conversation.instructions`, then
                              `avatars[].instructions`.
                          tools:
                            type: array
                            description: Tools available to responses in the session.
                            items:
                              type: object
                              description: >-
                                One tool definition. Only function tools are
                                documented for v1.
                              properties:
                                type:
                                  type: string
                                  description: Use `function`.
                                name:
                                  type: string
                                  description: >-
                                    Function name available to the response
                                    planner.
                                description:
                                  type: string
                                  description: >-
                                    Human-readable description of when the tool
                                    should be used.
                                parameters:
                                  type: object
                                  description: >-
                                    JSON Schema object describing function
                                    arguments.
                              required:
                                - type
                                - name
                          response_defaults:
                            type: object
                            description: >-
                              Default response-generation settings for later
                              turns.
                            properties:
                              temperature:
                                type: number
                                description: >-
                                  Sampling temperature for generated response
                                  text. Supported range is `0` to `2`.
                              max_output_tokens:
                                type: integer
                                description: Maximum generated text tokens for a response.
                          turn_detection:
                            type: object
                            description: Default turn handling for realtime user input.
                            properties:
                              type:
                                type: string
                                description: >-
                                  `manual` means the client starts responses
                                  explicitly. `server_vad` means the server
                                  detects turn boundaries from incoming user
                                  audio.
                              silence_duration_ms:
                                type: integer
                                description: >-
                                  Compatibility field. Do not use it to tune
                                  actual end-of-utterance timing. See [Speech
                                  recognition
                                  configuration](/streaming-avatar/integrate/speech-recognition)
                                  for recognition and turn-detection settings.
                            required:
                              - type
                          input_audio_transcription:
                            type: object
                            description: >-
                              Default transcription settings for user audio
                              input.
                            properties:
                              enabled:
                                type: boolean
                                description: Whether user audio should be transcribed.
                              language:
                                type: string
                                description: >-
                                  BCP-47 `language` code, or `auto` for
                                  automatic `language` detection.
                      assets:
                        type: object
                        description: >-
                          Session-scoped reusable assets for `response.script`.
                          Asset ids must be unique across `audio`,
                          `speech_text`, and `visual_prompts`. Avatar source
                          images stay under `avatars[].visual.source_images`. On
                          the REST side the visual prompt text field is
                          `prompt`; on the WSS `asset.add` side the same asset
                          uses `text` — the two shapes differ.
                        properties:
                          audio:
                            type: array
                            description: >-
                              Reusable audio files that can be referenced by
                              scripted vocal output.
                            items:
                              type: object
                              description: One reusable audio asset.
                              properties:
                                asset_id:
                                  type: string
                                  description: Caller-defined audio asset id.
                                url:
                                  type: string
                                  description: >-
                                    HTTPS URL for the audio file. The URL must
                                    be directly fetchable by Vivix without
                                    custom request headers.
                                mime_type:
                                  type: string
                                  description: >-
                                    Audio MIME type. Supported values are
                                    `audio/wav` and `audio/mpeg`.
                              required:
                                - asset_id
                                - url
                          speech_text:
                            type: array
                            description: >-
                              Reusable exact speech text that can be referenced
                              by scripted vocal output.
                            items:
                              type: object
                              description: One reusable speech text asset.
                              properties:
                                asset_id:
                                  type: string
                                  description: Caller-defined speech text asset id.
                                text:
                                  type: string
                                  description: >-
                                    Exact text to synthesize and stream through
                                    output text events when used in
                                    `response.script`.
                              required:
                                - asset_id
                                - text
                          visual_prompts:
                            type: array
                            description: >-
                              Reusable visual instructions that can be
                              referenced by scripted visual output.
                            items:
                              type: object
                              description: One reusable visual prompt asset.
                              properties:
                                asset_id:
                                  type: string
                                  description: Caller-defined visual prompt asset id.
                                prompt:
                                  type: string
                                  description: >-
                                    Instruction for avatar movement, posture,
                                    gaze, gesture, or presentation.
                              required:
                                - asset_id
                                - prompt
                      delivery:
                        type: object
                        description: >-
                          Media delivery settings used when `session.mode` is
                          `video_avatar`. If omitted, the media transport
                          defaults to TRTC. It is not used for `text_chat`.
                        properties:
                          media:
                            type: object
                            description: >-
                              Selects how the client sends and receives realtime
                              media.
                            properties:
                              transport:
                                type: string
                                description: >-
                                  Media transport requested for the session. Use
                                  `trtc` for TRTC or `agora` for Agora. If
                                  omitted, it defaults to TRTC. The response
                                  returns the corresponding provider-specific
                                  join details under `delivery.media.trtc` or
                                  `delivery.media.agora`.
                      max_duration_seconds:
                        type: integer
                        description: >-
                          Maximum session duration in seconds. If omitted, the
                          default session length applies. An explicit value must
                          be greater than `3` (minimum 4).
                      auto_close:
                        type: object
                        description: >-
                          Automatic close policy fixed when the session is
                          created. The response returns the effective values
                          after platform defaults are applied.
                        properties:
                          disconnected_timeout_seconds:
                            type: integer
                            description: >-
                              Closes after all control WebSocket connections
                              have been absent for this period, including when
                              none was ever opened. Omission uses the platform
                              setting; read the effective value in the response.
                              Use `0` to disable, or an integer from `10` to
                              `3600` to enable.


                              seconds. `0` disables the rule; nonzero values
                              must be `10`–`3600`. Omission uses the platform
                              setting.
                          interaction_idle_timeout_seconds:
                            type: integer
                            description: >-
                              Closes the session after this duration without a
                              successfully accepted client interaction or
                              detected user speech. Accepted business events and
                              user speech reset the timeout; WSS Ping/Pong and
                              server events do not. Omitted or 0 disables this
                              policy; explicit non-zero values must be 30–3600.


                              An explicit non-zero timeout cannot exceed the
                              effective maximum session duration.


                              seconds. Omitted or `0` disables the rule; nonzero
                              values must be `30`–`3600`.
                      recording_mode:
                        type: string
                        description: '`on` or `off`. Send `off` to disable recording.'
                    required:
                      - character_id
                      - name
                      - created_at
                      - updated_at
              example:
                code: 0
                message: success
                data:
                  character_id: chr_a1b2c3d4e5f6789012345678abcdef01
                  name: Studio Host v2
                  created_at: '2026-08-19T10:00:00Z'
                  updated_at: '2026-08-19T11:00:00Z'
                  model: vivix-a1-stream
                  output:
                    aspect_ratio: '9:16'
                    resolution: 720p
                  pipeline_config:
                    tts_config:
                      tts_provider: qwen-audio-3.0-tts-flash_ws
                      tts_voice_id: longanhuan_v3.6
                      tts_model_id: qwen-audio-3.0-tts-flash
                  avatars:
                    - avatar_id: host_a
                      instructions: You are Host A, a warm livestream host.
                      visual:
                        default_source_image_id: front
                        source_images:
                          - source_image_id: front
                            url: https://cdn.example.com/avatars/host-a-front.png
                            description: >-
                              A woman facing the camera in a brightly lit
                              studio, front view.
                            media_type: image/png
components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: Your Vivix API key. Keep it on your server.

````