The Voice Relay WebSocket enables real-time speech-to-text (STT) and text-to-speech (TTS) interaction between your application and an active Sinch voice call.
Enables real-time voice-to-text and text-to-speech communication between the Sinch platform and customer endpoints. The Voice Relay service transcribes incoming audio to text, relays it to your AI chatbot or agent system for processing, and synthesizes text responses back into audio for playback on active voice channels.
It allows you to:
- Receive live transcriptions of caller audio
- Send text to be synthesized into audio
- Stream audio chunks
- Build conversational AI agents and IVRs
You can download the full AsyncAPI specification here: