> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# response.create

Asks the service to create a response. By default this triggers model inference; when `response.script` is supplied, the service renders the client-provided vocal and visual script directly instead. Avatar selection is handled by the session state, not by `response.create`.

**`after_current_response`** allows only one pending response. Keep longer queues in your application and submit them in sequence.

**`response.input`** replaces the default context for this response. `[]` supplies no context. Complete inline items can be sent directly; an `item_reference` must already exist.

## Event fields

<ParamField body="event_id" type="string">
  Client-generated id for this event.
</ParamField>

<ParamField body="type" type="string" required>
  Event type. Must be `response.create`.
</ParamField>

<ParamField body="response" type="object">
  Per-response creation parameters. Fields set here apply only to this response.
</ParamField>

<ParamField body="response.conversation" type="string">
  Controls where the response is added. Use `auto` to write to the default conversation, or `none` for an out-of-band response that does not write output to the default conversation.
</ParamField>

<ParamField body="response.scheduling_policy" type="string">
  Controls what happens when the session is already busy. Use `interrupt` to cancel the active response or render and start this response, `reject_if_busy` to reject this event with an `error`, or `after_current_response` to run this response immediately after the current response finishes. Defaults to `interrupt`. Interrupted responses still emit their existing terminal lifecycle events, such as `response.render.stopped` followed by `response.done`.
</ParamField>

<ParamField body="response.input" type="array">
  Input items to include in the model context when asking the model to respond. Each array item is either an `item_reference` or a raw `conversation_item`. Providing this field creates a response-specific context instead of using the default conversation; an empty array clears context for this response. Invalid when `script` is supplied.
</ParamField>

#### Variant: `item_reference`

References an existing item already present in the session.

<ParamField body="response.input[].type" type="string" required>
  Use `item_reference`.
</ParamField>

<ParamField body="response.input[].id" type="string" required>
  Existing item id to include in the response context.
</ParamField>

#### Variant: `conversation_item`

A raw conversation item, using the same message, function call, and function call output shapes documented for `conversation.item.create`.

<ParamField body="response.instructions" type="string">
  Instructions for this response only. These steer model behavior; they are invalid when `script` is supplied. When these conflict with conversation or selected avatar instructions, these instructions take precedence for this response.
</ParamField>

<ParamField body="response.max_output_tokens" type="integer | string">
  Maximum output tokens for this response, inclusive of tool calls. Use an integer or `inf`. Invalid when `script` is supplied.
</ParamField>

<ParamField body="response.tools" type="array">
  Tools available for this response. When supplied, this overrides session conversation tools for this response only. Invalid when `script` is supplied.
</ParamField>

**`response.tools[].function` tool**

Function tool with `type`, `name`, optional `description`, and optional JSON Schema `parameters`.

**`response.tools[].mcp` tool**

MCP tool support is planned and not yet available in this version. This version documents the `function` tool only.

<ParamField body="response.script" type="object">
  Direct avatar output to render exactly as provided. When supplied, no LLM reply generation or tool calling occurs, and `input`, `instructions`, `tools`, and `max_output_tokens` are invalid.
</ParamField>

<ParamField body="response.script.vocal" type="object">
  Exact vocal output for the active avatar. Provide `vocal`, `visual`, or both.
</ParamField>

#### Variant: `speech`

Synthesized speech from exact text.

<ParamField body="response.script.vocal.type" type="string" required>
  Use `speech`.
</ParamField>

<ParamField body="response.script.vocal.text" type="string" required>
  Exact spoken text. This text also streams through `response.output_text.delta` and `response.output_text.done`.
</ParamField>

#### Variant: `speech_asset`

Synthesized speech from a reusable speech text asset.

<ParamField body="response.script.vocal.type" type="string" required>
  Use `speech_asset`.
</ParamField>

<ParamField body="response.script.vocal.speech_text_asset_id" type="string" required>
  Previously registered asset from `assets.speech_text`. The asset text is exact spoken text and streams through `response.output_text.delta` and `response.output_text.done`.
</ParamField>

#### Variant: `audio`

Supplied audio used to drive the avatar.

<ParamField body="response.script.vocal.type" type="string" required>
  Use `audio`.
</ParamField>

<ParamField body="response.script.vocal.mime_type" type="string" required>
  Required MIME type for the supplied audio. Supported values are `audio/wav` and `audio/mpeg`. The bytes served by `url` or encoded in `audio` must match this value.
</ParamField>

<ParamField body="response.script.vocal.url" type="string">
  HTTPS URL for an audio file to render through the avatar. The URL must be directly fetchable by Vivix without custom request headers. Provide exactly one of `url` or `audio`.
</ParamField>

<ParamField body="response.script.vocal.audio" type="string">
  Base64-encoded audio file bytes to render through the avatar. Do not include a data URL prefix. Provide exactly one of `url` or `audio`.
</ParamField>

#### Variant: `audio_asset`

Reusable audio asset used to drive the avatar.

<ParamField body="response.script.vocal.type" type="string" required>
  Use `audio_asset`.
</ParamField>

<ParamField body="response.script.vocal.audio_asset_id" type="string" required>
  Previously registered asset from `assets.audio`.
</ParamField>

<ParamField body="response.script.visual" type="object">
  Visual instruction for the active avatar. Used only in `video_avatar` mode.
</ParamField>

<ParamField body="response.script.visual.prompt" type="string">
  Inline instruction for avatar movement, posture, gaze, gestures, or presentation. Separate multiple ordered action descriptions with `[SHOT_SEP]`.
</ParamField>

<ParamField body="response.script.visual.visual_prompt_asset_id" type="string">
  Previously registered asset from `assets.visual_prompts`. Provide exactly one of `prompt` or `visual_prompt_asset_id`.
</ParamField>

<ParamField body="response.script.visual.duration_ms" type="integer">
  Target duration of each visual action segment in milliseconds. Defaults to `5000` when omitted. When `prompt` contains `[SHOT_SEP]`, this value applies to each segment, not the combined sequence. For example, three segments at `5000` milliseconds each plan about 15 seconds of motion. It does not change the duration of supplied audio or guarantee exact browser playback timing.
</ParamField>

**Script behavior.**

| | Behavior |
| - | - |
| **Model bypass** | `script` renders direct avatar output. No LLM reply generation or tool calling occurs, and exact vocal text is not rewritten by model or avatar instructions. |
| **video\_avatar\` mode** | Vocal output drives speech or supplied audio, and visual text guides avatar motion over RTC media using the current active avatar, voice, source image, and visual defaults. |
| **Conversation writing** | If `conversation` is `auto` or omitted, inline speech text or speech text asset content is written as an assistant message. If `conversation` is `none`, it is not. |
| **Lower-level orchestration** | Lower-level explicit media orchestration, such as explicit shot queues, track item ids, segment arrays, delete and clear operations, timed audio gaps, and asset-backed media composition, is not publicly available yet. |

**Response scheduling.**

The session is busy until the active response reaches terminal `response.done`. If rendering starts, `response.render.stopped` is emitted first and `response.done` follows.

<ParamField body="interrupt" type="When busy" required>
  Cancels active output, clears any pending `after_current_response`, and starts the new response. Interrupted responses still emit terminal lifecycle events. · When idle: Starts immediately. This is the default.
</ParamField>

<ParamField body="reject_if_busy" type="When busy" required>
  Rejects with an `error` if the session is active, rendering, or has a pending response. · When idle: Starts immediately.
</ParamField>

<ParamField body="after_current_response" type="When busy" required>
  Accepts at most one pending response. A second pending request is rejected with an `error`. If the session is busy, the response is queued until the current response finishes playing (in `video_avatar` mode, until it reaches terminal `response.done`). · When idle: Starts immediately.
</ParamField>

Pending `after_current_response` requests snapshot effective response inputs at submit time, including `mode`, `active_avatar_id`, `source_image_id`, selected avatar defaults and instructions, `conversation.instructions`, response fields, and referenced asset contents. Later `session.update`, `avatar.add`, or `asset.add` events do not silently change the pending response. Pending responses do not emit `response.created` until they start.

<RequestExample>
  ```json Default response.create theme={null}
  {
    "event_id": "evt_response_create_001",
    "type": "response.create"
  }
  ```

  ```json Reject if busy theme={null}
  {
    "event_id": "evt_response_create_reject_busy_001",
    "type": "response.create",
    "response": {
      "scheduling_policy": "reject_if_busy"
    }
  }
  ```

  ```json After current response theme={null}
  {
    "event_id": "evt_response_create_after_current_001",
    "type": "response.create",
    "response": {
      "scheduling_policy": "after_current_response",
      "script": {
        "vocal": {
          "type": "speech",
          "text": "I will cover that next."
        }
      }
    }
  }
  ```

  ```json Out-of-band response theme={null}
  {
    "event_id": "evt_response_create_002",
    "type": "response.create",
    "response": {
      "instructions": "Provide a concise answer.",
      "tools": [],
      "conversation": "none",
      "input": [
        {
          "type": "item_reference",
          "id": "item_12345"
        },
        {
          "type": "message",
          "role": "user",
          "content": [
            {
              "type": "input_text",
              "text": "Summarize the product in one sentence."
            }
          ]
        }
      ]
    }
  }
  ```

  ```json Scripted speech and visual asset theme={null}
  {
    "event_id": "evt_response_create_003",
    "type": "response.create",
    "response": {
      "script": {
        "vocal": {
          "type": "speech_asset",
          "speech_text_asset_id": "welcome_line"
        },
        "visual": {
          "visual_prompt_asset_id": "wave_small",
          "duration_ms": 2000
        }
      }
    }
  }
  ```

  ```json Scripted audio asset and visual motion theme={null}
  {
    "event_id": "evt_response_create_004",
    "type": "response.create",
    "response": {
      "script": {
        "vocal": {
          "type": "audio_asset",
          "audio_asset_id": "welcome_jingle"
        },
        "visual": {
          "prompt": "smile and look into the camera",
          "duration_ms": 3000
        }
      }
    }
  }
  ```

  ```json Scripted visual-only action theme={null}
  {
    "event_id": "evt_response_create_005",
    "type": "response.create",
    "response": {
      "script": {
        "visual": {
          "prompt": "look toward the product and nod",
          "duration_ms": 2000
        }
      }
    }
  }
  ```

  ```json Out-of-band scripted speech theme={null}
  {
    "event_id": "evt_response_create_006",
    "type": "response.create",
    "response": {
      "conversation": "none",
      "script": {
        "vocal": {
          "type": "speech",
          "text": "This line is rendered but not added to the conversation."
        }
      }
    }
  }
  ```

  ```json Invalid scripted response (script with instructions and input is rejected) theme={null}
  {
    "event_id": "evt_response_create_007",
    "type": "response.create",
    "response": {
      "instructions": "Make this more excited.",
      "input": [],
      "script": {
        "vocal": {
          "type": "speech",
          "text": "Welcome back."
        }
      }
    }
  }
  ```
</RequestExample>
