> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vivix.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# World Video Generation

Generate a finished world video asynchronously with Vivix-W1. One request carries a prompt, the requested duration and aspect ratio, and optionally reference images and reference audio.

Submit one job to receive a `job_id`, poll the job until it reaches a terminal state, and play the result from `video_url`. Reference material is uploaded first, so a job that uses references touches four endpoints:

1. `POST /v1/files/images` / `POST /v1/files/audios` — upload reference material and keep the returned `url` (skip this for text-to-video);
2. `POST /v1/videos/jobs` — submit a generation job;
3. `GET /v1/videos/jobs/{job_id}` — query one job;
4. `POST /v1/videos/jobs/batch-get` — query several jobs at once.

<h2 id="world-model-selection">
  Model Selection
</h2>

Set the request `model` field to the W1 model:

| model | Product line |
| - | - |
| `vivix-w1` | Vivix-W1 world video generation |

For the complete Vivix model lineup, see [Models](/overview/models).

<h2 id="world-reference-material">
  Reference Material
</h2>

Upload every reference file to your workspace first, then pass the `url` from the upload response in `ref_images` / `ref_audios`. Reference URLs must be publicly reachable over HTTP(S); the `file_id` is used by other endpoints and is not accepted here.

| Endpoint | Accepted types | Limit |
| - | - | - |
| `POST /v1/files/images` | jpg, png, webp, heic | 10MB per file |
| `POST /v1/files/audios` | wav, mp3, m4a, ogg, webm (mp3 and wav recommended) | 10MB per file |

File type is detected from the file content, not from the file name. The server sniffs the leading bytes and compares the detected type with the declared `Content-Type`; it must be one of the accepted types above and must match the declared value, or the upload returns `20008`. Declare the real media type in your multipart `Content-Type`.

```bash Upload an image theme={null}
curl -X POST https://api.vivix.ai/v1/files/images \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -F "file=@traveler.png"
```

```json Upload response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "file_id": "file_img_xxxxxxxx-xxx",
    "file_type": "image",
    "mime_type": "image/png",
    "size_bytes": 412876,
    "url": "https://example.com/images/xxxxxxxx.png",
    "created_at": "2026-09-15T08:00:00.000000000Z"
  }
}
```

```bash Upload audio theme={null}
curl -X POST https://api.vivix.ai/v1/files/audios \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -F "file=@market.mp3"

{
  "code": 0,
  "message": "success",
  "data": {
    "file_id": "file_aud_xxxxxxxx-xxx",
    "file_type": "audio",
    "mime_type": "audio/mpeg",
    "size_bytes": 335822,
    "url": "https://example.com/audios/xxxxxxxx.mp3",
    "created_at": "2026-09-15T08:00:12.000000000Z"
  }
}
```

<h2 id="world-keyframes">
  First and Last Frame
</h2>

To pin exactly how the video opens or ends, pass `first_frame_image` / `last_frame_image`. Both take an uploaded `url`, exactly like reference material.

| Fields | Result |
| - | - |
| `first_frame_image` only | Generation starts from this image. |
| `last_frame_image` only | Generation converges onto this image. |
| Both | Starts at the first and lands on the last; the same image twice gives a seamless loop. |

Keyframes and reference material are two different modes, so `first_frame_image` / `last_frame_image` cannot be combined with `ref_images` / `ref_audios`; mixing them returns `20004`.

<h2 id="world-task-types">
  Task Types
</h2>

The task type is derived from the material you send — there is no `mode` field. `prompt`, `duration` and `aspect_ratio` apply to every type.

| Task type | Material | What it does |
| - | - | - |
| Text to video | `prompt` only | Builds the whole scene from your description. |
| Image reference to video | `prompt` + `ref_images` | Keeps characters, objects or style consistent with the reference images. |
| Audio reference to video | `prompt` + `ref_audios` | Generates the video together with a soundtrack guided by the reference audio. |
| Image and audio reference to video | `prompt` + `ref_images` + `ref_audios` | Combines both kinds of reference material in one job. |
| First frame to video | `prompt` + `first_frame_image` | Opens on the given image and moves forward from it. |
| Last frame to video | `prompt` + `last_frame_image` | Ends on the given image, converging toward it. |
| First and last frame to video | `prompt` + `first_frame_image` + `last_frame_image` | Travels from the first image to the last one. |

```bash Text to video theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-w1",
  "jobs": [
    {
      "prompt": "A traveler crosses a windswept salt flat at sunset as the camera slowly pulls back.",
      "aspect_ratio": "16:9",
      "duration": 8
    }
  ]
}'
```

**How to write the prompt**

Describe who is in the scene, what happens and how it ends, then add the style, camera and sound details you care about. State appearance, blocking, action order and the final shot plainly when they matter, and use observable wording instead of vague mood words.

* Write in natural language or in labelled fields — there is no fixed length or order.
* Add "one continuous shot" when you do not want cuts, and advance with first / then / finally.
* Speech, lyrics and on-screen text go in the finished language, in quotes, with the speaker named.

```bash Image reference to video theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-w1",
  "jobs": [
    {
      "prompt": "The traveler from the reference image walks across a windswept salt flat at sunset.",
      "aspect_ratio": "16:9",
      "duration": 8,
      "ref_images": [
        { "url": "https://example.com/images/xxxxxxxx.png" }
      ]
    }
  ]
}'
```

**How to write the prompt**

Say which aspect each reference image provides and which subject it applies to. Use the number or name the upload shows, and spell out anything you want changed or left out.

* For a character: "Image 1 defines the character’s look, move the scene to a rainy street."
* With several images, name the pairing: "Image 1 for the character, image 2 only for the bold line work and palette."
* You do not have to list everything you are not referencing; only say so when it is easy to confuse.

```bash Audio reference to video theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-w1",
  "jobs": [
    {
      "prompt": "A night market comes alive, matching the mood of the reference audio.",
      "aspect_ratio": "9:16",
      "duration": 10,
      "ref_audios": [
        { "url": "https://example.com/audios/xxxxxxxx.mp3" }
      ]
    }
  ]
}'
```

**How to write the prompt**

Say what the audio is for — voice timbre, music style or ambience — and which subject or segment it binds to. With several people, bind each track explicitly instead of relying on upload order.

* Borrowing a voice: "Audio 1 is the girl’s timbre for image 2; deliver the lines below verbatim."
* Using the original track: name the purpose and entry point, for example "the first 8 seconds of audio 1 as the backing track".
* Style only: write "reference the style, do not use the original track".

```bash Image and audio reference to video theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-w1",
  "jobs": [
    {
      "prompt": "The traveler and the vehicle from the reference images cross the salt flat while the reference music plays.",
      "aspect_ratio": "16:9",
      "duration": 12,
      "ref_images": [
        { "url": "https://example.com/images/xxxxxxxx.png" },
        { "url": "https://example.com/images/yyyyyyyy.png" }
      ],
      "ref_audios": [
        { "url": "https://example.com/audios/yyyyyyyy.mp3" }
      ]
    }
  ]
}'
```

**How to write the prompt**

Say which aspect each reference image provides and which subject it applies to, then describe the soundtrack the audio should guide. Bind every track to its subject explicitly when several are in play.

* Name the pairing: "Image 1 and image 2 are the traveler and the vehicle; audio 1 sets the pacing and mood."
* State what changes and what stays: "keep the vehicle design from image 2 unchanged."
* Add the camera and sound you want, and call out anything that must be excluded.

```bash First and last frame to video theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-w1",
  "jobs": [
    {
      "prompt": "The traveler walks from the dawn dunes toward the salt flat as the light turns golden.",
      "aspect_ratio": "16:9",
      "duration": 8,
      "first_frame_image": "https://static.vivi-x.ai/images/2026/09/15/3a9c7f14b2e85d60.png",
      "last_frame_image": "https://static.vivi-x.ai/images/2026/09/15/c05d8e37a9f16b42.png"
    }
  ]
}'
```

**How to write the prompt**

Write how the motion leaves the first frame and how it arrives at the last one. You do not need to restate either image — add only the key change and the details that are easy to misread.

* Both frames: describe the path between them, for example "the character opens the umbrella and settles into the final pose and framing".
* First frame only: write the action that follows. Last frame only: write the process that leads up to it.
* Say whether the camera slows, holds or stays in one continuous shot, and keep people and clothing consistent.

<h2 id="world-generation-inputs">
  Generation Inputs
</h2>

| Field | Accepted value | Purpose |
| - | - | - |
| `prompt` | non-empty string | Required. Natural-language description of the video. |
| `duration` | integer, 5–15 | Required. Requested video duration in seconds. |
| `aspect_ratio` | `"16:9"` / `"9:16"` / `"1:1"` / `"4:3"` / `"3:4"` | Optional. Defaults to `"16:9"`. |
| `ref_images` | `{ url }[]`, at most 5 | Optional. Reference images. Each item carries a url only. |
| `ref_audios` | `{ url }[]`, at most 3 | Optional. Reference audio. Each item carries a url only. |
| `first_frame_image` | string (URL) | Optional. Pins the opening frame. Cannot be combined with reference material. |
| `last_frame_image` | string (URL) | Optional. Pins the closing frame. Cannot be combined with reference material. |

<Note>
  **One job per request.** The `jobs` array must contain exactly one item, and unknown fields are rejected with `20004`.
</Note>

<h2 id="world-calling-conventions">
  Calling Conventions
</h2>

These conventions apply to every endpoint on this page.

### Base URL

`https://api.vivix.ai`

### Authentication

Every request must carry an API key: `Authorization: Bearer $VIVIX_API_KEY`.

### Response Envelope

REST responses use the shared `{code, message, data}` envelope. Business fields are returned inside `data`.

```json Response envelope theme={null}
{
  "code": 0,
  "message": "success",
  "data": {}
}
```

* `code = 0` means the request succeeded.
* A non-zero `code` indicates an API-level error; the HTTP status is derived from it.
* An accepted job can still fail later. Poll the job and inspect `status` and `error`.

<h2 id="world-task-lifecycle">
  Task Lifecycle
</h2>

1. Submit `POST /v1/videos/jobs` and store the returned `job_id`.
2. Poll `GET /v1/videos/jobs/{job_id}` every few seconds while the job is `queued` or `processing`.
3. When the job reaches `succeeded`, play or download the video from `video_url`. Stop polling on `failed` and read `error`.

| Status | Meaning | Client action |
| - | - | - |
| `queued` | The job is waiting to start. | Keep polling. |
| `processing` | The video is being generated. | Keep polling. |
| `succeeded` | The generated video is ready. | Read `video_url`. |
| `failed` | The job did not complete successfully. | Stop polling and read `error.code`. |

<h2 id="create-world-video-generation">
  Submit a Generation Job
</h2>

Creates an asynchronous W1 video generation job from a prompt and optional reference material.

**POST** `/v1/videos/jobs`

### Body Parameters

<ParamField body={"model"} type={"\"vivix-w1\""} required>
  Required. The W1 model used for this job.
</ParamField>

<ParamField body={"pe_mode"} type={"\"standard\" | \"lite\""} required>
  Optional. Defaults to standard. Select standard for the existing prompt enhancement workflow, or lite for a single multimodal enhancement call with image compression. This setting does not change video duration, resolution or inference mode.
</ParamField>

<ParamField body={"seed"} type={"integer | null"} required>
  Optional. An integer from 0 to 2147483647. Omit or send null for a backend-generated random seed. When set, all tasks use the same seed, including zero. This controls video inference, not prompt enhancement randomness.
</ParamField>

<ParamField body={"jobs"} type={"array"} required>
  Required. Exactly one job.
</ParamField>

<ParamField body={"jobs[].prompt"} type={"string"} required>
  Required. Natural-language description of the video.
</ParamField>

<ParamField body={"jobs[].duration"} type={"integer"} required>
  Required. Video duration in seconds, between 5 and 15.
</ParamField>

<ParamField body={"jobs[].aspect_ratio"} type={"string"} required>
  Optional. One of `"16:9"`, `"9:16"`, `"1:1"`, `"4:3"`, `"3:4"`. Defaults to `"16:9"`.
</ParamField>

<ParamField body={"jobs[].ref_images"} type={"{ url: string }[]"} required>
  Optional. Up to 5 reference images.
</ParamField>

<ParamField body={"jobs[].ref_audios"} type={"{ url: string }[]"} required>
  Optional. Up to 3 reference audio files.
</ParamField>

### Response Fields

Fields in the response `data` object:

<ResponseField name={"results[].job_id"} type={"string"}>
  Job identifier used by the query endpoints.
</ResponseField>

<ResponseField name={"results[].status"} type={"string"}>
  Initial job status, usually `queued`.
</ResponseField>

```bash Request theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "vivix-w1",
  "pe_mode": "standard",
  "seed": 12345,
  "jobs": [
    {
      "prompt": "A traveler crosses a windswept salt flat at sunset as the camera slowly pulls back.",
      "aspect_ratio": "16:9",
      "duration": 8
    }
  ]
}'
```

```json Response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "results": [
      {
        "job_id": "job_xxxxxxxx",
        "status": "queued"
      }
    ]
  }
}
```

<h2 id="get-world-video-generation">
  Query a Job
</h2>

Returns the current state of one job. See [Task Lifecycle](#world-task-lifecycle) for terminal-state handling.

**GET** `/v1/videos/jobs/{job_id}`

### Path Parameters

<ParamField path={"job_id"} type={"string"} required>
  Job identifier returned on submission.
</ParamField>

### Response Fields

<ResponseField name={"job_id"} type={"string"}>
  Job identifier.
</ResponseField>

<ResponseField name={"status"} type={"string"}>
  `queued` / `processing` / `succeeded` / `failed`.
</ResponseField>

<ResponseField name={"video_url"} type={"string"}>
  HLS (m3u8) playback URL. Returned only when the job succeeded.
</ResponseField>

<ResponseField name={"aspect_ratio"} type={"string"}>
  Aspect ratio of the generated video.
</ResponseField>

<ResponseField name={"duration"} type={"number"}>
  Requested duration in seconds, echoed back. The encoded media can be marginally longer.
</ResponseField>

<ResponseField name={"error"} type={"object | null"}>
  Non-null only when the job failed. `error.code` is `generation_failed`, `generation_timeout` or `generation_cancelled`.
</ResponseField>

<ResponseField name={"created_at"} type={"timestamp"}>
  Job creation time.
</ResponseField>

<ResponseField name={"updated_at"} type={"timestamp"}>
  Most recent job update time.
</ResponseField>

An unknown job id returns `20005` (HTTP 404).

```bash Request theme={null}
curl https://api.vivix.ai/v1/videos/jobs/job_xxxxxxxx \
  -H "Authorization: Bearer $VIVIX_API_KEY"
```

```json Response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "job_id": "job_xxxxxxxx",
    "status": "succeeded",
    "video_url": "https://example.com/clips/job_xxxxxxxx/job_xxxxxxxx.m3u8",
    "aspect_ratio": "16:9",
    "duration": 8,
    "error": null,
    "created_at": "2026-09-15T07:27:40.938208Z",
    "updated_at": "2026-09-15T07:28:45.207793Z"
  }
}
```

```json Failed job theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "job_id": "job_xxxxxxxx",
    "status": "failed",
    "aspect_ratio": "16:9",
    "duration": 8,
    "error": {
      "code": "generation_failed",
      "message": "Video generation failed"
    },
    "created_at": "2026-09-15T07:27:40.938208Z",
    "updated_at": "2026-09-15T07:31:02.771204Z"
  }
}
```

<h2 id="batch-get-world-video-generation">
  Query Several Jobs
</h2>

Returns the same job view for several ids in one call. Results follow the order of `job_ids`; unknown ids are omitted from `items`.

**POST** `/v1/videos/jobs/batch-get`

### Body Parameters

<ParamField body={"job_ids"} type={"string[]"} required>
  Job identifiers to query.
</ParamField>

### Response Fields

<ResponseField name={"items"} type={"array"}>
  Job objects with the same fields as `GET /v1/videos/jobs/{job_id}`.
</ResponseField>

```bash Request theme={null}
curl -X POST https://api.vivix.ai/v1/videos/jobs/batch-get \
  -H "Authorization: Bearer $VIVIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "job_ids": [
    "job_xxxxxxxx",
    "job_yyyyyyyy"
  ]
}'
```

```json Response theme={null}
{
  "code": 0,
  "message": "success",
  "data": {
    "items": [
      {
        "job_id": "job_xxxxxxxx",
        "status": "succeeded",
        "video_url": "https://example.com/clips/job_xxxxxxxx/job_xxxxxxxx.m3u8",
        "aspect_ratio": "16:9",
        "duration": 8,
        "error": null,
        "created_at": "2026-09-15T07:27:40.938208Z",
        "updated_at": "2026-09-15T07:28:45.207793Z"
      },
      {
        "job_id": "job_yyyyyyyy",
        "status": "processing",
        "aspect_ratio": "9:16",
        "duration": 10,
        "error": null,
        "created_at": "2026-09-15T07:30:11.102338Z",
        "updated_at": "2026-09-15T07:30:44.918227Z"
      }
    ]
  }
}
```

<h2 id="world-video-error-codes">
  Error Codes
</h2>

API errors return the common response envelope with `data: null`:

```json Error response theme={null}
{
  "code": 20004,
  "message": "invalid parameter: jobs[0].duration must be an integer between 5 and 15",
  "data": null
}
```

| code | HTTP | Meaning |
| - | - | - |
| `10001` | 401 | API key is missing. |
| `10003` | 403 | API key is invalid. |
| `10004` / `10005` | 402 | Workspace billing is inactive or the balance is insufficient. |
| `10008` | 429 | API rate limit exceeded. |
| `20004` | 400 | Invalid parameter: unsupported model, more than one job, missing prompt or duration, duration outside 5–15, unsupported aspect ratio, too many references, an unreachable reference URL, or an unknown field. |
| `20005` | 404 | Job not found. |
| `20008` | 415 | Unsupported media type on upload: the detected file content is not an accepted type, or it does not match the declared Content-Type. The filename extension is not checked. |
| `30001` | 403 | Content rejected by moderation. |
| `30008` | 429 | Model capacity is exhausted. Retry later. |
| `50001` | 500 | Internal error. |
| `50003` | 503 | Service unavailable. |

<Note>
  **Two error layers.** The envelope `code` covers submission and query errors. A job accepted with `code = 0` can still end as `status = failed`, carrying `error.code`.
</Note>

When contacting support, include the `X-Request-Id` and `X-Trace-Id` response headers.
