Skip to content

feat(sarvam): add bulbul:v4-flash TTS model support - #5844

Open
dhruvladia-sarvam wants to merge 1 commit into
pipecat-ai:mainfrom
dhruvladia-sarvam:feat/sarvam-tts-bulbul-v4
Open

dhruvladia-sarvam wants to merge 1 commit into
pipecat-ai:mainfrom
dhruvladia-sarvam:feat/sarvam-tts-bulbul-v4

Conversation

@dhruvladia-sarvam

@dhruvladia-sarvam dhruvladia-sarvam commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds Sarvam's bulbul:v4-flash to SarvamTTSService and SarvamHttpTTSService.

Both services already select the model by name and drive parameters from TTS_MODEL_CONFIGS, so this is a new enum value plus the model's capability entry:

  • Supports pitch and loudness, which bulbul:v3 ignores.
  • Ignores temperature, which Sarvam pins to 0.6, so the service drops it rather than sending a value that has no effect.
  • Pace range 0.5–2.0, default sample rate 24000 Hz, preprocessing always on.

It streams over the same /text-to-speech/ws endpoint as the other models, pinned by the model query parameter, so no endpoint handling changed.

Speakers

bulbul:v4-flash has its own catalogue and does not accept the bulbul:v3 names. Its config carries an empty speakers tuple: callers pass a speaker from Sarvam's docs, and Sarvam's API rejects an unknown one with a 400 that lists what the model accepts. Omitting voice falls back to shubh_en_narration_gentle, which is the API's own default.

Testing

tests/test_sarvam_tts.py covers the model's defaults, that pitch and loudness reach the config message, that temperature is dropped, and that pace clamps to the model's range.

Verified against the live API with a key that had v4-flash access: WebSocket synthesis at 24 kHz and 8 kHz in English and Hindi with pitch/loudness/pace applied, HTTP synthesis via /text-to-speech, and bulbul:v3 unchanged.

Access to bulbul:v4-flash is gated per subscription; the key lost access partway through, so the latest live run returns the 422 beta-access notice rather than audio. A reviewer with access can re-confirm. bulbul:v3 needs no special access and is verified on the final code.

@codecov

codecov Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/pipecat/services/sarvam/tts.py 51.08% <100.00%> (+4.24%) ⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@dhruvladia-sarvam
dhruvladia-sarvam force-pushed the feat/sarvam-tts-bulbul-v4 branch 2 times, most recently from 38ee17a to 4f67cfd Compare September 18, 2026 16:05

@markbackman markbackman left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before getting too deep into the review, can you please submit changes required only for this model? It looks like you're adding helpers for error handling, which I've pushed back on before. Barring new API, this change should be adding new params supported and any model specific capabilties.

Comment thread src/pipecat/services/sarvam/tts.py Outdated
SOPHIA = "sophia"


class SarvamTTSSpeakerV4Flash(StrEnum):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't define voices in the Pipecat code. Developers should find voices in your docs and your server would ideally return an error if a voice ID is incorrect.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arguably, the same should be true of the v2 and v3 voices.

@dhruvladia-sarvam dhruvladia-sarvam Sep 21, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't define voices in the Pipecat code. Developers should find voices in your docs and your server would ideally return an error if a voice ID is incorrect.

Thanks, both fair. I cut it down to the model entry and its capabilities: dropped the speaker enum, the error-handling helpers, the sample-rate check, the Language additions, and the refactors. speakers is empty for v4-flash so get_speakers_for_model returns [] for it. Have left the default speaker value in the code for a minimal visibility via the plugin.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arguably, the same should be true of the v2 and v3 voices.

Left the v2/v3 speaker enums alone since removing them is breaking; happy to do it as a separate PR.

bulbul:v4-flash supports pitch and loudness, which bulbul:v3 ignores, and
ignores temperature, which Sarvam pins to 0.6. It has its own speaker
catalogue, so its config carries no speaker list: callers pass a speaker from
Sarvam's docs, and Sarvam's API validates it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants