Skip to main content
This page covers asynchronous video generation powered by Vivix-A1: you provide a first-frame image plus either an audio clip or a line of dialogue, and A1 generates a video of the character speaking and performing. From photorealistic portraits to anime characters, A1 is designed to keep expressions, lip sync, and motion working together while preserving the character identity and style throughout the video. Unlike the realtime sessions of Streaming Avatar, video generation is an asynchronous task: submit one request to get a video_id, poll the task status, and fetch the playable video URL once the task completes. Material is uploaded first, so the whole flow touches three endpoints:
  1. POST /v1/files/images / POST /v1/files/audios — upload the first-frame image and the audio material, and keep the returned url;
  2. POST /v1/videos/generations — submit a generation task;
  3. GET /v1/videos/generations/{video_id} — query the task status and result.
For the capabilities of the A1 model itself and the product-line split, see Models.

Model Selection

A1 video generation selects its model with the model field of the request body: For first-frame image sizes, see Material Upload.

Voice Source

Provide one of the two: audio or audio_data for voice audio you already have, or dialogue + voice_id to synthesize speech on the server (set the provider in tts_config). For voice ids, see Choose a voice and Clone Voice.

Material Upload

Upload the first-frame image and the audio material to your workspace first, then pass the url from the upload response in first_frame_image / audio.
  • POST /v1/files/images — the first-frame image, up to 10MB per file;
  • POST /v1/files/audios — the driving audio, up to 10MB per file.
Any other publicly reachable URL is accepted as well, but its availability is then yours to guarantee; uploading to the workspace is the supported path. First-frame image size. Provide the first frame at the size of the target resolution; an image of any other size is automatically compressed to it:

Calling Conventions

The conventions below apply to every endpoint in this group.

Base URL

https://api.vivix.ai

Authentication

Every request must carry an API key: Authorization: Bearer $VIVIX_API_KEY.

Response Envelope

Every REST response is wrapped in a {code, message, data} envelope, and all business fields live in data:
Response envelope
  • code = 0 means success; any non-zero value means failure (the HTTP status code is derived from the mapping of code, see Error Codes).
  • JSON fields use snake_case; 64-bit integers may be returned as strings, for example "cost_time_ms": "16562".

Business Errors

Business errors from task-submission endpoints usually surface as HTTP 200 with data.error_code != 0. To decide whether a submission succeeded, check both the outer code and data.error_code. Authentication failures return HTTP 401/403 directly.

Task Lifecycle

The complete lifecycle of a generation task:
  1. Submit: POST /v1/videos/generations. On success the response carries a video_id and the task enters queued.
  2. Poll: check the task status with GET /v1/videos/generations/{video_id}.
  3. Terminal state: when the status is succeeded, fetch the video from playback_url; stop polling on failed / timeout.
Interpret each polling result with the rules below:

Submit a Generation Task

Submits a video generation task driven by an audio clip or by dialogue text. POST /v1/videos/generations

Body Parameters

string
Model name. For A1 generation use vivix-a1 (720p).
string
required
Main video prompt describing the shot and the performance. Required.
string
required
Character first-frame image URL. Required; for the expected size see Material Upload.
string
required
TTS dialogue text. Required when driving with dialogue.
string
required
Voice id. Required when driving with dialogue; see Voice Source.
string
required
Audio URL. When driving with audio, send exactly one of audio or audio_data.
string
required
Base64-encoded audio. When driving with audio, send exactly one of audio or audio_data.
object
Fine-grained TTS settings when driving with dialogue.
string
Resolution: 720p / 480p.
string
Aspect ratio, for example 16:9 or 9:16.

Response Fields

Fields of the response data object:
string
Video task ID, used for subsequent queries.
int32
Business error code; 0 means success.
string
Human-readable status or error detail; empty on success.
string
Usually queued on a successful submission.
object
Playback information.
Dialogue request (server-side TTS)
Audio request
Response

Query Task Status

Queries the status of a video generation task. See Task Lifecycle for how to interpret each status. GET /v1/videos/generations/{video_id}

Path Parameters

string
required
Video generation task ID.

Response Fields

The response data.items is a list of tasks (an empty list means the task has not been written yet or does not exist). Each item has the following fields:
string
Video generation task ID.
string
Current task status: queued / processing / succeeded / failed / timeout.
string
Most recent task event: status / completed / failed.
timestamp
Time of the most recent event.
string
Playback URL of the generated video; may be empty while the task is still processing.
string
Error detail for the task when it failed; empty otherwise.
string
Total processing time in ms (int64, returned as a string).
timestamp
Task creation time.
timestamp
Task update time.
Request
Response

Error Codes

Error responses share a single format, with data set to null:
Error response
code = 0 means success; any non-zero value means failure, and the HTTP status code is derived from the mapping of code. Common error codes for video generation:
Two error layers. The outer code is a gateway-level error. Task-level business errors usually surface as HTTP 200 with an outer code = 0, but data.error_code != 0 (submit endpoint) or status = failed / timeout (query endpoint).
When reporting an issue, include the X-Request-Id and X-Trace-Id response headers so the request can be located in the pipeline logs.