Atlas Knowledge Base
Dashboard
Text to Speech

Text to Speech


Text to Speech

The Text to Speech service turns written text into spoken audio and plays it into a call - for announcements, prompts, and other spoken messages.

Overview

When SBN needs to say something on a call - an announcement, an instruction, a prompt - the Text to Speech service produces the audio for it. You give it the text, the language, and the voice; it generates the speech and plays it into the call.

The speech itself is produced by a cloud voice provider. Two are supported:

  • ElevenLabs
  • AWS Polly

Generated audio is cached, so a message that is used repeatedly is only produced once.

  • Runs: When configured. One per host.
  • Required: Only where SBN plays spoken messages into calls. Not needed otherwise.
  • Depends on: NATS; the device that plays the audio; and outbound internet access to the chosen voice provider, with that provider's credentials.

Configuration

Text to Speech is configured under the texttospeech namespace. View the defaults with ./sbn-media config eject and set overrides in sbn-media.local.yaml.

SettingDefaultDescription
texttospeech.defaultLocaleenThe language used when a request does not name one.
texttospeech.voices(per locale)A map of language to the voice used for it - the voice's ID, the provider, and the engine/model.
texttospeech.cacheLocation./tts-cacheWhere generated audio is cached so repeated messages are not re-generated.
texttospeech.elevenlabs.api(none)The ElevenLabs API key, required to use ElevenLabs voices.
texttospeech.awspolly.accesskey / awspolly.credentials(none)AWS credentials for Polly. If left unset, the standard AWS credential sources are used.

Each entry in texttospeech.voices ties a language to a voice and a provider, so different languages can use different voices. The provider you choose must have valid credentials and be reachable from the host.

After changing these in sbn-media.local.yaml, apply them with ./sbn-media config reload.

FAQ

Which voice is used for a message?

The one mapped to the message's language in texttospeech.voices, or the voice for defaultLocale if the request does not specify a language. A request can also name the provider, voice, and engine directly.

Speech is not being produced.

Check that the chosen provider's credentials are set (texttospeech.elevenlabs.api, or AWS credentials for Polly) and that the host can reach the provider over the internet. A provider error or a blocked outbound connection will stop the audio.

Does it need internet access?

Yes. The voice is generated by a cloud provider (ElevenLabs or AWS Polly), so the host needs outbound access to that provider. Cached messages are replayed without contacting the provider again.

How do I add another language?

Add an entry to texttospeech.voices mapping that language to a voice and provider.

Related pages

  • SBN Media Overview (SBN-Media/overview)
  • Installing and Configuring SBN Media (SBN-Media/installation)
  • SIP Media (SBN-Media/Telephony/sip-media)


Was this helpful?