voice_id as pipeline_config.tts_config.tts_voice_id when creating a Streaming Avatar session or saving a character.
Upload the reference audio to the current workspace first: POST /v1/files/audios.
Response envelope. REST responses are wrapped in {code, message, data}: code = 0 means success. Field descriptions below refer to fields inside data.
1. Upload Audio
Uploads an audio file and returns afile_id for cloning. Files are scoped to the workspace.
POST /v1/files/audios
multipart/form-data, field name file. Supported types: audio/wav, audio/mpeg, audio/mp3, audio/mp4, audio/ogg, audio/webm. Maximum 10MB.
Clone sample requirements: 10–20 seconds recommended (60 seconds max), WAV (16bit) / MP3 / M4A, sample rate ≥16kHz, mono or stereo (first channel only), with at least 5 seconds of continuous clear speech.
Request Body Parameters
file: binary Audio file. Supported types:audio/wav, audio/mpeg, audio/mp3, audio/mp4, audio/ogg, audio/webm. Maximum 10MB; WAV (16bit) / MP3 are recommended.
Response Fields
file_id: string File id, shaped likefile_aud_…. Pass this as file_id when cloning.
file_type: string
Always audio.
mime_type: string
Detected MIME type.
size_bytes: integer
File size.
url: string
CDN URL.
created_at: timestamp
Creation time.
Error Cases
Missingfile field. · API error code: 20010 missing file field
Empty file. · API error code: 20006 empty file
Larger than 10MB. · API error code: 20007 file too large
Unsupported audio type. · API error code: 20008 unsupported media type
Invalid multipart form. · API error code: 20009 invalid multipart form
Request
2. Clone Voice
Clones a voice from an uploaded audiofile_id and returns a voice_id. Pass that id as pipeline_config.tts_config.tts_voice_id when you start a session.
POST /v1/voices
Body Parameters
name: string Display name, at most 128 characters. file_id: string Audio file id in the current workspace (call Upload Audio first).language: optional string
Sample language: zh (Chinese), en (English), fr (French), de (German), ja (Japanese), ko (Korean), ru (Russian), pt (Portuguese), th (Thai), id (Indonesian), vi (Vietnamese), it (Italian), es (Spanish), ms (Malay), fil (Filipino), ar (Arabic).
Response Fields
voice_id: string Voice id (shaped likevoice_…). Pass it unchanged as tts_voice_id.
name: string
Display name.
language: string
Language sent on create; empty string when omitted.
Error Cases
Missingname or file_id. · API error code: 20003 missing required parameter
language / name is invalid. · API error code: 20004 invalid parameter
file_id does not exist or is not in the current workspace. · API error code: 20005 not found
file_id is not an audio file. · API error code: 20004 invalid parameter
The clone provider call failed (HTTP 502). · API error code: 30007 voice clone provider failed
The API key is missing or invalid. · API error code: 10001 missing api key or 10003 invalid api key
Request
tts_voice_id is required: the service fills tts_provider / tts_model_id for that voice.
3. List Voices
Lists voices visible to the current workspace: voices cloned in this workspace, plus the public voices Vivix provides. GET/v1/voices
Response Fields
voices: array Visible voices.voices[].voice_id: string
Voice id (shaped like voice_…).
voices[].name: string
Display name.
voices[].language: string
One of zh (Chinese), en (English), fr (French), de (German), ja (Japanese), ko (Korean), ru (Russian), pt (Portuguese), th (Thai), id (Indonesian), vi (Vietnamese), it (Italian), es (Spanish), ms (Malay), fil (Filipino), ar (Arabic), or an empty string.
Request
4. Get Voice
Looks up one voice visible to the current workspace byvoice_id. Same shape as Clone Voice.
GET /v1/voices/{voice_id}
Path Parameters
voice_id: string Thevoice_id returned by clone or list.
Response Fields
voice_id: string Voice id (shaped likevoice_…). Pass it unchanged as tts_voice_id.
name: string
Display name.
language: string
Language sent on create; empty string when omitted.
Error Cases
The voice does not exist or is not visible to the current workspace. · API error code:20005 not found
Request