feat(sarvam): add bulbul:v4-flash TTS model support - #5844
dhruvladia-sarvam wants to merge 1 commit into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests.
🚀 New features to boost your workflow:
|
38ee17a to
4f67cfd
Compare
markbackman
left a comment
There was a problem hiding this comment.
Before getting too deep into the review, can you please submit changes required only for this model? It looks like you're adding helpers for error handling, which I've pushed back on before. Barring new API, this change should be adding new params supported and any model specific capabilties.
| SOPHIA = "sophia" | ||
|
|
||
|
|
||
| class SarvamTTSSpeakerV4Flash(StrEnum): |
There was a problem hiding this comment.
We shouldn't define voices in the Pipecat code. Developers should find voices in your docs and your server would ideally return an error if a voice ID is incorrect.
There was a problem hiding this comment.
Arguably, the same should be true of the v2 and v3 voices.
There was a problem hiding this comment.
We shouldn't define voices in the Pipecat code. Developers should find voices in your docs and your server would ideally return an error if a voice ID is incorrect.
Thanks, both fair. I cut it down to the model entry and its capabilities: dropped the speaker enum, the error-handling helpers, the sample-rate check, the Language additions, and the refactors. speakers is empty for v4-flash so get_speakers_for_model returns [] for it. Have left the default speaker value in the code for a minimal visibility via the plugin.
There was a problem hiding this comment.
Arguably, the same should be true of the v2 and v3 voices.
Left the v2/v3 speaker enums alone since removing them is breaking; happy to do it as a separate PR.
bulbul:v4-flash supports pitch and loudness, which bulbul:v3 ignores, and ignores temperature, which Sarvam pins to 0.6. It has its own speaker catalogue, so its config carries no speaker list: callers pass a speaker from Sarvam's docs, and Sarvam's API validates it.
b0b853f to
15c2327
Compare
Summary
Adds Sarvam's
bulbul:v4-flashtoSarvamTTSServiceandSarvamHttpTTSService.Both services already select the model by name and drive parameters from
TTS_MODEL_CONFIGS, so this is a new enum value plus the model's capability entry:pitchandloudness, whichbulbul:v3ignores.temperature, which Sarvam pins to 0.6, so the service drops it rather than sending a value that has no effect.It streams over the same
/text-to-speech/wsendpoint as the other models, pinned by themodelquery parameter, so no endpoint handling changed.Speakers
bulbul:v4-flashhas its own catalogue and does not accept thebulbul:v3names. Its config carries an emptyspeakerstuple: callers pass a speaker from Sarvam's docs, and Sarvam's API rejects an unknown one with a 400 that lists what the model accepts. Omittingvoicefalls back toshubh_en_narration_gentle, which is the API's own default.Testing
tests/test_sarvam_tts.pycovers the model's defaults, thatpitchandloudnessreach the config message, thattemperatureis dropped, and thatpaceclamps to the model's range.Verified against the live API with a key that had v4-flash access: WebSocket synthesis at 24 kHz and 8 kHz in English and Hindi with pitch/loudness/pace applied, HTTP synthesis via
/text-to-speech, andbulbul:v3unchanged.Access to
bulbul:v4-flashis gated per subscription; the key lost access partway through, so the latest live run returns the 422 beta-access notice rather than audio. A reviewer with access can re-confirm.bulbul:v3needs no special access and is verified on the final code.