Skip to main content
Voice endpoints clone, list, and retrieve the voices available to the current workspace. After a clone succeeds, pass the returned voice_id as pipeline_config.tts_config.tts_voice_id when creating a Streaming Avatar session or saving a character. Upload the reference audio to the current workspace first: POST /v1/files/audios. Response envelope. REST responses are wrapped in {code, message, data}: code = 0 means success. Field descriptions below refer to fields inside data.

1. Upload Audio

Uploads an audio file and returns a file_id for cloning. Files are scoped to the workspace. POST /v1/files/audios multipart/form-data, field name file. Supported types: audio/wav, audio/mpeg, audio/mp3, audio/mp4, audio/ogg, audio/webm. Maximum 10MB. Clone sample requirements: 10–20 seconds recommended (60 seconds max), WAV (16bit) / MP3 / M4A, sample rate ≥16kHz, mono or stereo (first channel only), with at least 5 seconds of continuous clear speech.

Request Body Parameters

file: binary Audio file. Supported types: audio/wav, audio/mpeg, audio/mp3, audio/mp4, audio/ogg, audio/webm. Maximum 10MB; WAV (16bit) / MP3 are recommended.

Response Fields

file_id: string File id, shaped like file_aud_…. Pass this as file_id when cloning. file_type: string Always audio. mime_type: string Detected MIME type. size_bytes: integer File size. url: string CDN URL. created_at: timestamp Creation time.

Error Cases

Missing file field. · API error code: 20010 missing file field Empty file. · API error code: 20006 empty file Larger than 10MB. · API error code: 20007 file too large Unsupported audio type. · API error code: 20008 unsupported media type Invalid multipart form. · API error code: 20009 invalid multipart form Request
Response

2. Clone Voice

Clones a voice from an uploaded audio file_id and returns a voice_id. Pass that id as pipeline_config.tts_config.tts_voice_id when you start a session. POST /v1/voices

Body Parameters

name: string Display name, at most 128 characters. file_id: string Audio file id in the current workspace (call Upload Audio first). language: optional string Sample language: zh (Chinese), en (English), fr (French), de (German), ja (Japanese), ko (Korean), ru (Russian), pt (Portuguese), th (Thai), id (Indonesian), vi (Vietnamese), it (Italian), es (Spanish), ms (Malay), fil (Filipino), ar (Arabic).

Response Fields

voice_id: string Voice id (shaped like voice_…). Pass it unchanged as tts_voice_id. name: string Display name. language: string Language sent on create; empty string when omitted.

Error Cases

Missing name or file_id. · API error code: 20003 missing required parameter language / name is invalid. · API error code: 20004 invalid parameter file_id does not exist or is not in the current workspace. · API error code: 20005 not found file_id is not an audio file. · API error code: 20004 invalid parameter The clone provider call failed (HTTP 502). · API error code: 30007 voice clone provider failed The API key is missing or invalid. · API error code: 10001 missing api key or 10003 invalid api key Request
Response
Use the voice on Create Session
Only the tts_voice_id is required: the service fills tts_provider / tts_model_id for that voice.

3. List Voices

Lists voices visible to the current workspace: voices cloned in this workspace, plus the public voices Vivix provides. GET /v1/voices

Response Fields

voices: array Visible voices. voices[].voice_id: string Voice id (shaped like voice_…). voices[].name: string Display name. voices[].language: string One of zh (Chinese), en (English), fr (French), de (German), ja (Japanese), ko (Korean), ru (Russian), pt (Portuguese), th (Thai), id (Indonesian), vi (Vietnamese), it (Italian), es (Spanish), ms (Malay), fil (Filipino), ar (Arabic), or an empty string. Request
Response

4. Get Voice

Looks up one voice visible to the current workspace by voice_id. Same shape as Clone Voice. GET /v1/voices/{voice_id}

Path Parameters

voice_id: string The voice_id returned by clone or list.

Response Fields

voice_id: string Voice id (shaped like voice_…). Pass it unchanged as tts_voice_id. name: string Display name. language: string Language sent on create; empty string when omitted.

Error Cases

The voice does not exist or is not visible to the current workspace. · API error code: 20005 not found Request
Response
Choose a built-in voice

Continue reading

Character instructions · Speech recognition · Text model · Text to speech · Opening acknowledgements