mirror of
https://github.com/jambonz/speech-utils.git
synced 2026-10-04 07:43:59 +00:00
Adds synthFishaudio with both arms: the say: streaming url consumed by the mediajam dialect, and a POST /v1/tts cache render. The render asks for raw pcm at 8k and returns extension r8 because fish's wav output carries a placeholder RIFF size, the same problem gradium has. Fish is a voice-cloning vendor, so the voice is a reference_id; the sentinel 'default' means send none and use fish's own default voice. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>