Skip to main content

1. Choose a built-in voice or bring your own account

Text to speech turns the character’s text into audio. Use a platform voice without a separate provider account, or supply your own speech-provider key. Vivix TTS can work with your own dialogue model. If you already have audio, play it directly instead of synthesizing it again. Put these settings in pipeline_config.tts_config when creating a Session. Keep the existing image, instructions, and other configuration. For built-in voice choices, see Choose a voice.
This uses the default Qwen Audio service. You do not have to add tts_config to every session, but when you supply it, include the voice ID. For another service, select a matching provider, model, and voice rather than changing the voice ID alone.

2. Supply your own API key

Your Vivix API key authenticates the HTTP request to Vivix through the Authorization header. Your speech-provider key belongs in tts_config.tts_api_key. For an ElevenLabs account:
Replace both placeholders with your provider key and a voice available to that account. A nonempty tts_api_key takes precedence over the platform-configured credential. Omit it to use platform credentials. Access through Vivix does not imply that your provider account has access to the same voice. For your own DashScope account:
Read the key from an environment variable or secret store on your server, then insert it into the request. Keep real credentials out of browser code, public examples, and ordinary logs. If you save a key in a Character configuration, control access to that configuration; do not assume every character-detail response will hide it.

3. Configuration fields

For a cloned voice, put the voice_id returned by the Voices API in tts_voice_id. When switching from another service, replace the old tts_config rather than retaining its settings. See Voices API. Session-level tts_config applies to every avatar in the session and takes precedence over avatars[].voice. A single-character application usually needs only one configuration method.

4. Tune speech settings

Listen to a complete line with the defaults first, then change one setting at a time. This Qwen Audio example uses the normal speed multiplier:
Qwen Audio maps speed to its rate setting. If your DashScope setup requires a workspace, add tts_workspace_id inside payload; otherwise omit it. The ElevenLabs HTTP adapter accepts speed, stability, seed, and optimize_streaming_latency. Speed and stability map to the provider’s voice settings. Accepted values and effects depend on the selected model. Do not treat the latency option as a universal speed switch. Start with payload: {"speed": 1.0} in the matching ElevenLabs configuration. Add stability only with a value accepted by that model. One stability number is not appropriate for every model, and Qwen Audio parameters should not be assumed to apply to Qwen3-TTS. payload must be an object. Flat extension fields in tts_config are also accepted; when the same key appears in both places, the value inside payload wins. Prefer one form. The object is limited to 32 KiB and eight nesting levels. Keep credentials and routing fields outside it. Service-managed fields such as text, voice_id, output_format, sample_rate, and previous_text cannot be overridden through extensions.

5. When to set tts_endpoint

Neither built-in voices nor your own key requires an endpoint override. Use one only when you need a supported provider address. This field is not a universal proxy for arbitrary TTS protocols. Qwen Audio accepts only the corresponding official DashScope HTTPS or WSS endpoint, including the correct host and path. It does not accept an arbitrary compatible server URL. ElevenLabs HTTP expects a complete request URL and does not append its speech endpoint path; the WS adapter expects a complete WebSocket URL. Keep the default address unless you have a specific reason to change it.

6. Play existing audio without TTS

Send this through the session’s control WebSocket. Use audio/mpeg for MP3 or audio/wav for WAV, with a URL the service can fetch. Alternatively, supply Base64 file bytes in audio instead of url. The recording keeps its original voice; it is not converted to the selected TTS voice. See Speech and performance for the full flow, or External audio for continuous input.