Skip to main content
Generate a finished world video asynchronously with Vivix-W1. One request carries a prompt, the requested duration and aspect ratio, and optionally reference images and reference audio. Submit one job to receive a job_id, poll the job until it reaches a terminal state, and play the result from video_url. Reference material is uploaded first, so a job that uses references touches four endpoints:
  1. POST /v1/files/images / POST /v1/files/audios — upload reference material and keep the returned url (skip this for text-to-video);
  2. POST /v1/videos/jobs — submit a generation job;
  3. GET /v1/videos/jobs/{job_id} — query one job;
  4. POST /v1/videos/jobs/batch-get — query several jobs at once.

Model Selection

Set the request model field to the W1 model: For the complete Vivix model lineup, see Models.

Reference Material

Upload every reference file to your workspace first, then pass the url from the upload response in ref_images / ref_audios. Reference URLs must be publicly reachable over HTTP(S); the file_id is used by other endpoints and is not accepted here. File type is detected from the file content, not from the file name. The server sniffs the leading bytes and compares the detected type with the declared Content-Type; it must be one of the accepted types above and must match the declared value, or the upload returns 20008. Declare the real media type in your multipart Content-Type.
Upload an image
Upload response
Upload audio

First and Last Frame

To pin exactly how the video opens or ends, pass first_frame_image / last_frame_image. Both take an uploaded url, exactly like reference material. Keyframes and reference material are two different modes, so first_frame_image / last_frame_image cannot be combined with ref_images / ref_audios; mixing them returns 20004.

Task Types

The task type is derived from the material you send — there is no mode field. prompt, duration and aspect_ratio apply to every type.
Text to video
How to write the prompt Describe who is in the scene, what happens and how it ends, then add the style, camera and sound details you care about. State appearance, blocking, action order and the final shot plainly when they matter, and use observable wording instead of vague mood words.
  • Write in natural language or in labelled fields — there is no fixed length or order.
  • Add “one continuous shot” when you do not want cuts, and advance with first / then / finally.
  • Speech, lyrics and on-screen text go in the finished language, in quotes, with the speaker named.
Image reference to video
How to write the prompt Say which aspect each reference image provides and which subject it applies to. Use the number or name the upload shows, and spell out anything you want changed or left out.
  • For a character: “Image 1 defines the character’s look, move the scene to a rainy street.”
  • With several images, name the pairing: “Image 1 for the character, image 2 only for the bold line work and palette.”
  • You do not have to list everything you are not referencing; only say so when it is easy to confuse.
Audio reference to video
How to write the prompt Say what the audio is for — voice timbre, music style or ambience — and which subject or segment it binds to. With several people, bind each track explicitly instead of relying on upload order.
  • Borrowing a voice: “Audio 1 is the girl’s timbre for image 2; deliver the lines below verbatim.”
  • Using the original track: name the purpose and entry point, for example “the first 8 seconds of audio 1 as the backing track”.
  • Style only: write “reference the style, do not use the original track”.
Image and audio reference to video
How to write the prompt Say which aspect each reference image provides and which subject it applies to, then describe the soundtrack the audio should guide. Bind every track to its subject explicitly when several are in play.
  • Name the pairing: “Image 1 and image 2 are the traveler and the vehicle; audio 1 sets the pacing and mood.”
  • State what changes and what stays: “keep the vehicle design from image 2 unchanged.”
  • Add the camera and sound you want, and call out anything that must be excluded.
First and last frame to video
How to write the prompt Write how the motion leaves the first frame and how it arrives at the last one. You do not need to restate either image — add only the key change and the details that are easy to misread.
  • Both frames: describe the path between them, for example “the character opens the umbrella and settles into the final pose and framing”.
  • First frame only: write the action that follows. Last frame only: write the process that leads up to it.
  • Say whether the camera slows, holds or stays in one continuous shot, and keep people and clothing consistent.

Generation Inputs

One job per request. The jobs array must contain exactly one item, and unknown fields are rejected with 20004.

Calling Conventions

These conventions apply to every endpoint on this page.

Base URL

https://api.vivix.ai

Authentication

Every request must carry an API key: Authorization: Bearer $VIVIX_API_KEY.

Response Envelope

REST responses use the shared {code, message, data} envelope. Business fields are returned inside data.
Response envelope
  • code = 0 means the request succeeded.
  • A non-zero code indicates an API-level error; the HTTP status is derived from it.
  • An accepted job can still fail later. Poll the job and inspect status and error.

Task Lifecycle

  1. Submit POST /v1/videos/jobs and store the returned job_id.
  2. Poll GET /v1/videos/jobs/{job_id} every few seconds while the job is queued or processing.
  3. When the job reaches succeeded, play or download the video from video_url. Stop polling on failed and read error.

Submit a Generation Job

Creates an asynchronous W1 video generation job from a prompt and optional reference material. POST /v1/videos/jobs

Body Parameters

"vivix-w1"
required
Required. The W1 model used for this job.
"standard" | "lite"
required
Optional. Defaults to standard. Select standard for the existing prompt enhancement workflow, or lite for a single multimodal enhancement call with image compression. This setting does not change video duration, resolution or inference mode.
integer | null
required
Optional. An integer from 0 to 2147483647. Omit or send null for a backend-generated random seed. When set, all tasks use the same seed, including zero. This controls video inference, not prompt enhancement randomness.
array
required
Required. Exactly one job.
string
required
Required. Natural-language description of the video.
integer
required
Required. Video duration in seconds, between 5 and 15.
string
required
Optional. One of "16:9", "9:16", "1:1", "4:3", "3:4". Defaults to "16:9".
{ url: string }[]
required
Optional. Up to 5 reference images.
{ url: string }[]
required
Optional. Up to 3 reference audio files.

Response Fields

Fields in the response data object:
string
Job identifier used by the query endpoints.
string
Initial job status, usually queued.
Request
Response

Query a Job

Returns the current state of one job. See Task Lifecycle for terminal-state handling. GET /v1/videos/jobs/{job_id}

Path Parameters

string
required
Job identifier returned on submission.

Response Fields

string
Job identifier.
string
queued / processing / succeeded / failed.
string
HLS (m3u8) playback URL. Returned only when the job succeeded.
string
Aspect ratio of the generated video.
number
Requested duration in seconds, echoed back. The encoded media can be marginally longer.
object | null
Non-null only when the job failed. error.code is generation_failed, generation_timeout or generation_cancelled.
timestamp
Job creation time.
timestamp
Most recent job update time.
An unknown job id returns 20005 (HTTP 404).
Request
Response
Failed job

Query Several Jobs

Returns the same job view for several ids in one call. Results follow the order of job_ids; unknown ids are omitted from items. POST /v1/videos/jobs/batch-get

Body Parameters

string[]
required
Job identifiers to query.

Response Fields

array
Job objects with the same fields as GET /v1/videos/jobs/{job_id}.
Request
Response

Error Codes

API errors return the common response envelope with data: null:
Error response
Two error layers. The envelope code covers submission and query errors. A job accepted with code = 0 can still end as status = failed, carrying error.code.
When contacting support, include the X-Request-Id and X-Trace-Id response headers.