Speech-to-speech
Understand when to use speech-to-speech mode for a voice Agent.
GetVoiceBot supports standard voice processing and speech-to-speech modes. In speech-to-speech mode, one real-time model listens and responds with audio, which can make conversations feel faster and more expressive.
Providers and models
GetVoiceBot supports speech-to-speech through OpenAI and Google Gemini. Available providers, models, and compatible voices can vary by account as new capabilities become available. Upcoming OpenAI support includes gpt-live-1.
When to use it
Speech-to-speech is a good fit when natural turn-taking, low latency, and expressive delivery matter more than selecting separate speech-recognition and voice providers.
Standard mode is a better fit when you need the broader set of speech-recognition controls, voice providers, or tuning options.
Configure and test
- Open the Agent editor and select Voice.
- Choose speech-to-speech mode and an available provider.
- Select a compatible voice.
- Test realistic calls, including interruptions, names, numbers, and background noise.
- Publish only after the staged Agent behaves as expected.
Changing speech mode can change which voices and controls are available. Review the complete Voice and speech mode guide before switching a live Agent.
Configuring speech mode through the API? See Language, voice, and speech.
