> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Avatar Video Generation

This page covers asynchronous video generation powered by Vivix-A1: you provide a first-frame image plus either an audio clip or a line of dialogue, and A1 generates a video of the character speaking and performing. From photorealistic portraits to anime characters, A1 is designed to keep expressions, lip sync, and motion working together while preserving the character identity and style throughout the video.

Unlike the realtime sessions of [Streaming Avatar](/streaming-avatar/overview), video generation is an asynchronous task: submit one request to get a `video_id`, poll the task status, and fetch the playable video URL once the task completes. Material is uploaded first, so the whole flow touches three endpoints:

1. `POST /v1/files/images` / `POST /v1/files/audios` — upload the first-frame image and the audio material, and keep the returned `url`;
2. `POST /v1/videos/generations` — submit a generation task;
3. `GET /v1/videos/generations/{video_id}` — query the task status and result.

For the capabilities of the A1 model itself and the product-line split, see [Models](/overview/models).

<h2 id="model-selection">
  Model Selection
</h2>

A1 video generation selects its model with the `model` field of the request body:

| model | Output resolution | Notes |
| - | - | - |
| `vivix-a1` | 720p | A1 avatar video generation |

For first-frame image sizes, see [Material Upload](#material-upload).

<h2 id="voice-source">
  Voice Source
</h2>

Provide one of the two: `audio` or `audio_data` for voice audio you already have, or `dialogue` + `voice_id` to synthesize speech on the server (set the provider in `tts_config`). For voice ids, see [Choose a voice](/streaming-avatar/character/voice) and [Clone Voice](/streaming-avatar/api-references/voices#2-clone-voice).

<h2 id="material-upload">
  Material Upload
</h2>

Upload the first-frame image and the audio material to your workspace first, then pass the `url` from the upload response in `first_frame_image` / `audio`.

* `POST /v1/files/images` — the first-frame image, up to 10MB per file;
* `POST /v1/files/audios` — the driving audio, up to 10MB per file.

Any other publicly reachable URL is accepted as well, but its availability is then yours to guarantee; uploading to the workspace is the supported path.

**First-frame image size.** Provide the first frame at the size of the target resolution; an image of any other size is automatically compressed to it:

| Output resolution | Landscape | Portrait |
| - | - | - |
| 720p (`vivix-a1`) | 1280×672 | 672×1280 |

<h2 id="calling-conventions">
  Calling Conventions
</h2>

The conventions below apply to every endpoint in this group.

### Base URL

`https://api.vivix.ai`

### Authentication

Every request must carry an API key: `Authorization: Bearer $VIVIX_API_KEY`.

### Response Envelope

Every REST response is wrapped in a `{code, message, data}` envelope, and all business fields live in `data`:

```json Response envelope theme={null}
{
  "code": 0,
  "message": "success",
  "data": { }
}
```

* `code = 0` means success; any non-zero value means failure (the HTTP status code is derived from the mapping of `code`, see [Error Codes](#error-codes)).
* JSON fields use `snake_case`; 64-bit integers may be returned as strings, for example `"cost_time_ms": "16562"`.

### Business Errors

Business errors from task-submission endpoints usually surface as HTTP 200 with `data.error_code != 0`. To decide whether a submission succeeded, check both the outer `code` and `data.error_code`. Authentication failures return HTTP 401/403 directly.

<h2 id="task-lifecycle">
  Task Lifecycle
</h2>

The complete lifecycle of a generation task:

1. Submit: `POST /v1/videos/generations`. On success the response carries a `video_id` and the task enters `queued`.
2. Poll: check the task status with `GET /v1/videos/generations/{video_id}`.
3. Terminal state: when the status is `succeeded`, fetch the video from `playback_url`; stop polling on `failed` / `timeout`.

Interpret each polling result with the rules below:

| Query result | Meaning | Next step |
| - | - | - |
| `items` is empty | The task has not been written yet, or the `video_id` does not exist | Retry later |
| `status = queued / processing / ready` | The task is in progress | Keep short polling |
| `status = succeeded` and `playback_url` is non-empty | The task is complete | Play or download from `playback_url` |
| `status = failed / timeout` | The task failed or timed out | Stop polling |

<h2 id="create-video-generation">
  Submit a Generation Task
</h2>

Submits a video generation task driven by an audio clip or by dialogue text.

**POST** `/v1/videos/generations`

### Body Parameters

<ParamField body={"model"} type={"string"}>
  Model name. For A1 generation use `vivix-a1` (720p).
</ParamField>

<ParamField body={"prompt"} type={"string"} required>
  Main video prompt describing the shot and the performance. Required.
</ParamField>

<ParamField body={"first_frame_image"} type={"string"} required>
  Character first-frame image URL. Required; for the expected size see [Material Upload](#material-upload).
</ParamField>

<ParamField body={"dialogue"} type={"string"} required>
  TTS dialogue text. Required when driving with dialogue.
</ParamField>

<ParamField body={"voice_id"} type={"string"} required>
  Voice id. Required when driving with dialogue; see [Voice Source](#voice-source).
</ParamField>

<ParamField body={"audio"} type={"string"} required>
  Audio URL. When driving with audio, send exactly one of `audio` or `audio_data`.
</ParamField>

<ParamField body={"audio_data"} type={"string"} required>
  Base64-encoded audio. When driving with audio, send exactly one of `audio` or `audio_data`.
</ParamField>

<ParamField body={"tts_config"} type={"object"}>
  Fine-grained TTS settings when driving with dialogue.

  <Expandable title="properties">
    <ParamField body={"vol"} type={"number"}>
      Volume.
    </ParamField>

    <ParamField body={"seed"} type={"integer"}>
      Random seed.
    </ParamField>

    <ParamField body={"speed"} type={"number"}>
      Speech rate.
    </ParamField>

    <ParamField body={"stability"} type={"number"}>
      Stability.
    </ParamField>

    <ParamField body={"tts_model_id"} type={"string"}>
      TTS model id matching tts\_provider.
    </ParamField>

    <ParamField body={"tts_provider"} type={"string"}>
      TTS provider, the same provider as voice\_id; see [Choose a voice](/streaming-avatar/character/voice). ElevenLabs is used when omitted; for a cloned voice (voice\_…) it is filled automatically.
    </ParamField>
  </Expandable>
</ParamField>

<ParamField body={"resolution"} type={"string"}>
  Resolution: `720p` / `480p`.
</ParamField>

<ParamField body={"video_ratio"} type={"string"}>
  Aspect ratio, for example `16:9` or `9:16`.
</ParamField>

### Response Fields

Fields of the response `data` object:

<ResponseField name={"video_id"} type={"string"}>
  Video task ID, used for subsequent queries.
</ResponseField>

<ResponseField name={"error_code"} type={"int32"}>
  Business error code; `0` means success.
</ResponseField>

<ResponseField name={"message"} type={"string"}>
  Human-readable status or error detail; empty on success.
</ResponseField>

<ResponseField name={"status"} type={"string"}>
  Usually `queued` on a successful submission.
</ResponseField>

<ResponseField name={"playback"} type={"object"}>
  Playback information.

  <Expandable title="properties">
    <ResponseField name={"type"} type={"string"}>
      Playback type.
    </ResponseField>

    <ResponseField name={"url"} type={"string"}>
      Playback URL.
    </ResponseField>
  </Expandable>
</ResponseField>

```bash Dialogue request (server-side TTS) theme={null}
curl -X POST https://api.vivix.ai/v1/videos/generations \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-a1",
  "prompt": "A confident host smiles and greets the audience.",
  "first_frame_image": "https://static.vivi-x.ai/images/2026/09/16/xxxxxxxxxxxxxxxx.png",
  "dialogue": "Welcome back. It is a beautiful day to show you something new.",
  "voice_id": "<elevenlabs_voice_id>",
  "tts_config": {
    "tts_provider": "elevenlabs",
    "tts_model_id": "eleven_v3",
    "speed": 1,
    "vol": 1
  },
  "resolution": "720p"
}'
```

```bash Audio request theme={null}
curl -X POST https://api.vivix.ai/v1/videos/generations \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-a1",
  "prompt": "A confident host speaks to the camera.",
  "first_frame_image": "https://static.vivi-x.ai/images/2026/09/16/xxxxxxxxxxxxxxxx.png",
  "audio": "https://static.vivi-x.ai/audios/2026/09/16/xxxxxxxxxxxxxxxx.mp3",
  "resolution": "720p"
}'
```

```json Response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "video_id": "vid_xxx",
    "error_code": 0,
    "message": "",
    "status": "queued"
  }
}
```

<h2 id="get-video-generation">
  Query Task Status
</h2>

Queries the status of a video generation task. See [Task Lifecycle](#task-lifecycle) for how to interpret each status.

**GET** `/v1/videos/generations/{video_id}`

### Path Parameters

<ParamField path={"video_id"} type={"string"} required>
  Video generation task ID.
</ParamField>

### Response Fields

The response `data.items` is a list of tasks (an empty list means the task has not been written yet or does not exist). Each item has the following fields:

<ResponseField name={"video_id"} type={"string"}>
  Video generation task ID.
</ResponseField>

<ResponseField name={"status"} type={"string"}>
  Current task status: `queued` / `processing` / `succeeded` / `failed` / `timeout`.
</ResponseField>

<ResponseField name={"event"} type={"string"}>
  Most recent task event: `status` / `completed` / `failed`.
</ResponseField>

<ResponseField name={"event_timestamp"} type={"timestamp"}>
  Time of the most recent event.
</ResponseField>

<ResponseField name={"playback_url"} type={"string"}>
  Playback URL of the generated video; may be empty while the task is still processing.
</ResponseField>

<ResponseField name={"watch_error"} type={"string"}>
  Error detail for the task when it failed; empty otherwise.
</ResponseField>

<ResponseField name={"cost_time_ms"} type={"string"}>
  Total processing time in ms (int64, returned as a string).
</ResponseField>

<ResponseField name={"task_created_at"} type={"timestamp"}>
  Task creation time.
</ResponseField>

<ResponseField name={"task_updated_at"} type={"timestamp"}>
  Task update time.
</ResponseField>

```bash Request theme={null}
curl https://api.vivix.ai/v1/videos/generations/vid_xxx \
  -H "Authorization: Bearer $VIVIX_API_KEY"
```

```json Response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "items": [
      {
        "video_id": "vid_xxx",
        "status": "succeeded",
        "event": "completed",
        "event_timestamp": "2026-06-18T08:19:50Z",
        "playback_url": "https://example.com/video.m3u8",
        "watch_error": "",
        "cost_time_ms": "16562",
        "task_created_at": "2026-06-18T08:19:34.385558Z",
        "task_updated_at": "2026-06-18T08:19:50.948557Z"
      }
    ]
  }
}
```

<h2 id="error-codes">
  Error Codes
</h2>

Error responses share a single format, with `data` set to `null`:

```json Error response theme={null}
{
  "code": 50001,
  "message": "internal error",
  "data": null
}
```

`code = 0` means success; any non-zero value means failure, and the HTTP status code is derived from the mapping of `code`. Common error codes for video generation:

| code | message | HTTP | Description |
| - | - | - | - |
| 0 | success | 200 | Success |
| 10001 | missing api key | 401 | API key is missing |
| 10003 | invalid api key | 403 | API key is invalid |
| 10004 | workspace billing is not active | 402 | Workspace billing is not active |
| 10005 | insufficient workspace balance | 402 | Workspace balance is insufficient |
| 10008 | API rate limit exceeded | 429 | API key rate limit hit (60 req/min by default; RPM is configurable per key). Responses include `Retry-After`, `X-RateLimit-Limit`, `X-RateLimit-Remaining`, and `X-RateLimit-Reset`. |
| 20001 | invalid request body | 400 | Request body is invalid |
| 20002 | invalid json body | 400 | Malformed JSON |
| 20003 | missing required parameter | 400 | A required parameter is missing |
| 20004 | invalid parameter | 400 | A parameter is invalid |
| 20005 | not found | 404 | Resource not found |
| 30001 | content rejected by moderation | 403 | Content failed moderation |
| 30004 | invalid argument | 422 | Parameter or business validation failed, for example neither `audio` / `audio_data` nor `dialogue` + `voice_id` was provided (`data.detail` is `generation mode invalid`) |
| 50001 | internal error | 500 | Internal error |
| 50003 | service unavailable | 503 | Service unavailable |

<Note>
  **Two error layers.** The outer `code` is a gateway-level error. Task-level business errors usually surface as HTTP 200 with an outer `code = 0`, but `data.error_code != 0` (submit endpoint) or `status = failed / timeout` (query endpoint).
</Note>

When reporting an issue, include the `X-Request-Id` and `X-Trace-Id` response headers so the request can be located in the pipeline logs.
