Compare commits

...
16 Commits
Author SHA1 Message Date
Dave Horton f4c3c7dc8b 1.0.18 2026-08-23 13:12:37 -04:00
Dave HortonandClaude Opus 5 a32f71ac5a feat: add fishaudio (Fish Audio) TTS support (#159)
Adds synthFishaudio with both arms: the say: streaming url consumed by the
mediajam dialect, and a POST /v1/tts cache render. The render asks for raw pcm
at 8k and returns extension r8 because fish's wav output carries a placeholder
RIFF size, the same problem gradium has.

Fish is a voice-cloning vendor, so the voice is a reference_id; the sentinel
'default' means send none and use fish's own default voice.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 13:11:59 -04:00
Dave HortonandClaude Opus 5 c8850998b5 fix(tts): use a current nineninesix voice id in the synth test (#158)
The vendor replaced its voice catalog and now rejects unknown ids, so
the env-gated test failed whenever NINENINESIX_API_KEY was set.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:19:59 -04:00
Dave Horton e8b2009d29 1.0.17 2026-08-10 16:45:32 -04:00
Dave HortonandClaude Opus 5 27e07bb551 fix(tts): forward gradium json_config to the streaming path (#157)
synthGradium destructured json_config from options but only forwarded
pronunciation_id into the say: param block, so the setting could never
reach mediajam's streaming dialect — only the cache-render POST. The
say: parser is brace-aware, so a nested json object survives intact.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 16:45:07 -04:00
Dave Horton 33456b93d9 1.0.16 2026-08-08 23:07:43 -04:00
Dave HortonandClaude Opus 5 70e91b5057 test(inworld): gate the say: param test on INWORLD_API_KEY (#156)
The test I added in #155 ran unconditionally, which broke `npm test` without
credentials — the path husky's pre-commit hook takes, so `npm version patch`
could not commit.

Two causes, both addressed:

- no credential gate, unlike every other vendor test in this file. Now skips
  without INWORLD_API_KEY, and closes its redis client on that path so the
  run can still exit.
- it assumed streaming was enabled. The Google non-streaming test sets
  JAMBONES_DISABLE_TTS_STREAMING and, on its no-credentials skip path,
  deletes the env var WITHOUT clearing the require cache (unlike its finally
  block, which clears both) — so lib/config still held 'true' further down
  the file and synthInworld took the non-streaming branch, attempting a real
  vendor call. The test now re-requires with streaming enabled so it does not
  depend on what ran before it.

Verified both ways: skips and exits 0 with no key; 11/11 with a key even
under the leaked state. Re-introducing the #155 bug still fails 3 assertions,
so the regression value is intact.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:06:52 -04:00
Dave HortonandClaude Opus 5 c066840f85 fix(inworld): read pitch and speakingRate from audioConfig in the say: params (#155)
The streaming say: path guarded on opts.audioConfig?.pitch and
opts.audioConfig?.speakingRate but interpolated opts.pitch and
opts.speakingRate, which are undefined — so anyone setting them under
audioConfig (what the docs and the portal defaults tell you to do) got
'pitch=undefined,speakingRate=undefined' on the wire and their setting
silently dropped.

Adds a test for the say: params that needs no credentials, since the
streaming branch builds the path without calling the vendor.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:49:50 -04:00
Dave Horton d89457f67d 1.0.15 2026-08-06 10:23:47 -04:00
Dave HortonandClaude Opus 5 3c15976669 fix(deps): remove junk "23" dependency (#154)
"23": "^0.0.0" is not a real dependency - it is an empty placeholder
package (0.0.0, no deps, unrelated third-party maintainer) that landed
here from a stray npm install. Nothing in the package references it.

Beyond the noise, it is a small supply-chain liability: a dependency on
a squatted single-number name owned by nobody we know, shipped to every
consumer of speech-utils.

Lint passes; full test suite passes 105/105 with live vendor credentials.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:23:31 -04:00
Dave Horton 70abb04941 1.0.14 2026-08-06 10:13:51 -04:00
Dave HortonandClaude Opus 5 724af7fcab fix(deps): drop unused undici dependency (#153)
undici was declared as a direct dependency but is never required
anywhere in this package - grep across index.js, lib/ and stubs/ finds
no reference to undici, ProxyAgent, setGlobalDispatcher or Dispatcher.
The one HTTP call in lib/synth-audio.js uses the global fetch, and
Azure proxy support goes through the SDK's own setProxy plus the
http_proxy_ip/http_proxy_port params.

Removing it clears all seven open undici advisories from npm audit for
this package and its consumers, and stops speech-utils pulling a
duplicate undici 7.x into trees where the consumer already depends on
undici 8 (feature-server does).

Lint passes; full test suite passes 105/105 with live vendor credentials.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:13:31 -04:00
Dave Horton c0a79090c2 1.0.13 2026-08-06 10:04:26 -04:00
Dave HortonandClaude Opus 5 7a9a3e2468 fix(deps): bump microsoft-cognitiveservices-speech-sdk to ^1.51.0 (#152)
The Azure speech SDK was pinned to exactly 1.38.0 since the initial
commit. It pulls in uuid <11.1.1, which carries GHSA-w5hq-g745-h8pq
(missing buffer bounds check in v3/v5/v6 when buf is provided). The
advisory covers Azure SDK 1.14.0-1.50.0; 1.51.0 clears it.

Loosened the exact pin to a caret range so future patches come in
without another PR. The SDK surface used here (SpeechConfig,
SpeechSynthesizer, ResultReason, CancellationDetails,
SpeechSynthesisOutputFormat) is unchanged in 1.51.0.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:03:59 -04:00
Dave Horton ab2b3f7fb2 1.0.12 2026-08-04 09:04:50 -04:00
Dave HortonandClaude Opus 5 e848d0276e feat(tts): add gradium as a TTS vendor (#151)
Streaming arm returns a say: url for the mediajam dialect; the cache-render
arm posts to /api/post/speech/tts with only_audio and pcm_8000, which is bare
r8 samples and avoids gradium's streaming wav header (0xffffffff RIFF size).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 09:02:07 -04:00
5 changed files with 494 additions and 84 deletions
+84
View File
@@ -0,0 +1,84 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Overview
`@jambonz/speech-utils` is a Node.js library providing TTS (Text-to-Speech) utilities for the jambonz CPaaS platform. It handles speech synthesis with caching through Redis and supports multiple TTS vendors.
## Commands
```bash
# Run tests (requires Docker for Redis)
npm test
# Run linter
npm run jslint
# Auto-fix lint issues
npm run jslint:fix
# Generate coverage report
npm run coverage
```
## Testing
Tests use `tape` and require Redis. The test harness automatically starts/stops Redis via Docker Compose (`test/docker-compose-testbed.yaml`).
Most tests are conditional based on environment variables for vendor credentials:
- `GCP_FILE` or `GCP_JSON_KEY` - Google TTS
- `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_REGION` - AWS Polly
- `MICROSOFT_API_KEY`, `MICROSOFT_REGION` - Azure TTS
- `ELEVENLABS_API_KEY`, `ELEVENLABS_VOICE_ID`, `ELEVENLABS_MODEL_ID` - ElevenLabs
- `OPENAI_API_KEY` - OpenAI Whisper TTS
- And others per vendor
Redis config is in `config/test.json` (port 3379).
## Architecture
### Entry Point
`index.js` exports a factory function that takes Redis options and a logger, returning an object with these methods:
- `synthAudio` - Main synthesis function
- `getTtsVoices` - List available voices for a vendor
- `purgeTtsCache` / `getTtsSize` / `addFileToCache` - Cache management
- `getAwsAuthToken` - Token management
### Core Module: `lib/synth-audio.js`
The `synthAudio` function handles synthesis for all vendors. Key behaviors:
1. **Cache check**: Generates SHA1 hash key from (vendor, language, voice, engine, model, text, instructions)
2. **Streaming vs non-streaming**: When `JAMBONES_DISABLE_TTS_STREAMING` is not set and `renderForCaching=false`, returns `say:{params}text` format for FreeSWITCH streaming playback instead of generating files
3. **Vendor dispatch**: Switch statement routes to vendor-specific synth functions (`synthGoogle`, `synthPolly`, `synthMicrosoft`, etc.)
4. **Caching**: Stores audio as base64 JSON in Redis with configurable TTL (default 4 hours)
### Supported Vendors
google, aws/polly, microsoft/azure, nvidia (Riva), wellsaid, elevenlabs, cartesia, inworld, rimelabs, whisper (OpenAI), deepgram, resemble, custom:*
### gRPC Stubs
`stubs/riva/` contains generated protobuf/gRPC code for NVIDIA Riva.
## Environment Variables
Key configuration via env vars (see `lib/config.js`):
- `JAMBONES_DISABLE_TTS_STREAMING` - Force non-streaming mode
- `JAMBONES_DISABLE_AZURE_TTS_STREAMING` - Azure-specific streaming disable
- `JAMBONES_TTS_CACHE_DURATION_MINS` - Cache TTL in minutes (default: 240)
- `JAMBONES_TTS_TRIM_SILENCE` - Trim trailing silence from audio
- `JAMBONES_TMP_FOLDER` - Temp folder for audio files (default: /tmp)
- `JAMBONES_HTTP_PROXY_IP`, `JAMBONES_HTTP_PROXY_PORT` - HTTP proxy for Azure
- `JAMBONES_AZURE_ENABLE_SSML` - Force SSML wrapper for Azure plain text
## Key Dependencies
- `@jambonz/realtimedb-helpers` - Redis client and hash utilities
- `@google-cloud/text-to-speech` - Google TTS
- `@aws-sdk/client-polly` - AWS Polly
- `microsoft-cognitiveservices-speech-sdk` - Azure TTS
- `@grpc/grpc-js` - gRPC for Riva
- `openai` - OpenAI Whisper TTS
- `bent` - HTTP client for REST-based vendors
+158 -4
View File
@@ -80,8 +80,8 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
logger = logger || noopLogger;
assert.ok(['google', 'aws', 'polly', 'microsoft', 'wellsaid', 'nvidia', 'elevenlabs',
'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'nineninesix', 'inworld', 'resemble',
'murf', 'xai']
'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'gradium', 'nineninesix', 'inworld', 'resemble',
'murf', 'xai', 'fishaudio']
.includes(vendor) ||
vendor.startsWith('custom'),
`synthAudio supported vendors are google, aws, microsoft, nvidia and wellsaid ..etc, not ${vendor}`);
@@ -136,6 +136,9 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
} else if ('cartesia' === vendor) {
assert.ok(credentials.api_key, 'synthAudio requires api_key when cartesia is used');
assert.ok(credentials.model_id, 'synthAudio requires model_id when cartesia is used');
} else if ('gradium' === vendor) {
assert.ok(voice, 'synthAudio requires voice when gradium is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when gradium is used');
} else if ('nineninesix' === vendor) {
assert.ok(voice, 'synthAudio requires voice when nineninesix is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when nineninesix is used');
@@ -146,6 +149,10 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
} else if (vendor === 'resemble') {
assert.ok(voice, 'synthAudio requires voice when resemble is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when resemble is used');
} else if ('fishaudio' === vendor) {
/* no voice assert: fish synthesizes with its own default voice when
reference_id is omitted, which is what the 'default' selection means */
assert.ok(credentials.api_key, 'synthAudio requires api_key when fishaudio is used');
}
const key = makeSynthKey({
@@ -217,6 +224,16 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'fishaudio':
audioData = await synthFishaudio(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'gradium':
audioData = await synthGradium(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'nineninesix':
audioData = await synthNineninesix(logger, {
credentials, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
@@ -933,8 +950,9 @@ const synthInworld = async(logger, {
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (opts.temperature) params += `,temperature=${opts.temperature}`;
if (opts.audioConfig?.pitch) params += `,pitch=${opts.pitch}`;
if (opts.audioConfig?.speakingRate) params += `,speakingRate=${opts.speakingRate}`;
/* pitch and speakingRate are nested under audioConfig, matching Inworld's API */
if (opts.audioConfig?.pitch) params += `,pitch=${opts.audioConfig.pitch}`;
if (opts.audioConfig?.speakingRate) params += `,speakingRate=${opts.audioConfig.speakingRate}`;
params += '}';
return {
@@ -1433,8 +1451,144 @@ const synthCartesia = async(logger, {
};
/* gradium.ai — json websocket for streaming, and a POST endpoint for the cache
render. only_audio:true + pcm_8000 returns bare little-endian 16-bit samples,
which is exactly the r8 container, so we avoid gradium's streaming wav header
(it carries 0xffffffff as the RIFF size, since length is unknown up front).
*/
const synthGradium = async(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
const {api_key, model_id} = credentials;
const {json_config, pronunciation_id} = options || {};
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=gradium';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (model_id) params += `,model_id=${model_id}`;
if (pronunciation_id) params += `,pronunciation_id=${pronunciation_id}`;
/* the say: param parser is brace-aware, so a nested json object survives intact */
if (json_config) {
params += `,json_config=${typeof json_config === 'string' ? json_config : JSON.stringify(json_config)}`;
}
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const post = bent('https://api.gradium.ai', 'POST', 'buffer', {
'x-api-key': api_key,
'Content-Type': 'application/json'
});
const audioContent = await post('/api/post/speech/tts', {
text,
voice_id: voice,
model_name: model_id || 'default',
output_format: `pcm_${sampleRate}`,
only_audio: true,
...(json_config && {json_config}),
...(pronunciation_id && {pronunciation_id})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth gradium returned error');
stats.increment('tts.count', ['vendor:gradium', 'accepted:no']);
throw err;
}
};
/* nineninesix.ai — a Cartesia-compatible API, but only raw/wav come back
(mp3 is rejected), so the cache render asks for wav rather than mp3. */
/* fish.audio — msgpack websocket for streaming, and a POST endpoint for the cache
render. format:pcm + sample_rate:8000 returns bare little-endian 16-bit samples,
which is exactly the r8 container. we avoid fish's wav output because its RIFF
header carries a placeholder size (0xffffff24) — length is unknown up front, as
with gradium.
fish is a voice-cloning vendor: the "voice" is a reference_id returned by
POST /model, and omitting it entirely synthesizes with fish's default voice.
the sentinel value 'default' (the bundled fallback entry in the portal) means
exactly that — send no reference_id.
*/
const synthFishaudio = async(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
const {api_key, model_id, fishaudio_tts_uri} = credentials;
const {reference_id, latency, chunk_length, speed, volume} = options || {};
/* free-text reference_id in the vendor options wins over the voice selector */
const refId = reference_id || (voice && voice !== 'default' ? voice : null);
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=fishaudio';
params += `,voice=${refId || 'default'}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (model_id) params += `,model_id=${model_id}`;
if (latency) params += `,latency=${latency}`;
if (chunk_length) params += `,chunk_length=${chunk_length}`;
if (speed) params += `,speed=${speed}`;
if (volume) params += `,volume=${volume}`;
if (fishaudio_tts_uri) params += `,endpoint=${fishaudio_tts_uri}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const post = bent(fishaudio_tts_uri || 'https://api.fish.audio', 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json',
/* the model is selected by header, not in the body */
'model': model_id || 's2.1-pro'
});
const audioContent = await post('/v1/tts', {
text,
format: 'pcm',
sample_rate: sampleRate,
...(refId && {reference_id: refId}),
...(latency && {latency}),
...(chunk_length && {chunk_length: parseInt(chunk_length, 10)}),
...((speed || volume) && {prosody: {
...(speed && {speed: parseFloat(speed)}),
...(volume && {volume: parseFloat(volume)})
}})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth fishaudio returned error');
stats.increment('tts.count', ['vendor:fishaudio', 'accepted:no']);
throw err;
}
};
const synthNineninesix = async(logger, {
credentials, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
+115 -74
View File
@@ -1,15 +1,14 @@
{
"name": "@jambonz/speech-utils",
"version": "1.0.11",
"version": "1.0.18",
"lockfileVersion": 2,
"requires": true,
"packages": {
"": {
"name": "@jambonz/speech-utils",
"version": "1.0.11",
"version": "1.0.18",
"license": "MIT",
"dependencies": {
"23": "^0.0.0",
"@aws-sdk/client-polly": "^3.496.0",
"@aws-sdk/client-sts": "^3.496.0",
"@cartesia/cartesia-js": "^2.2.7",
@@ -19,9 +18,8 @@
"bent": "^7.3.12",
"debug": "^4.3.4",
"google-protobuf": "^3.21.2",
"microsoft-cognitiveservices-speech-sdk": "1.38.0",
"openai": "^4.98.0",
"undici": "^7.5.0"
"microsoft-cognitiveservices-speech-sdk": "^1.51.0",
"openai": "^4.98.0"
},
"devDependencies": {
"config": "^4.2.0",
@@ -760,6 +758,46 @@
"node": ">=18.0.0"
}
},
"node_modules/@azure/abort-controller": {
"version": "2.2.0",
"resolved": "https://registry.npmjs.org/@azure/abort-controller/-/abort-controller-2.2.0.tgz",
"integrity": "sha512-fNAjWnA/nZ2jz31kxR/AqRaUT8ewHBw/WuBIosK0moMy1C9e5ValbDfFdIxJzVOOYaYkV/b2F1S4H/aHiqfVQg==",
"license": "MIT",
"dependencies": {
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/@azure/core-auth": {
"version": "1.11.0",
"resolved": "https://registry.npmjs.org/@azure/core-auth/-/core-auth-1.11.0.tgz",
"integrity": "sha512-IUZydyTUkDnYdstOW9pFOOUQlBjAepK5teihDE3x6yxsPJs/hsAaaYpeGxdxrgtOiJbBKSjKW7MDk7AEhb4LRg==",
"license": "MIT",
"dependencies": {
"@azure/abort-controller": "^2.1.2",
"@azure/core-util": "^1.13.0",
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/@azure/core-util": {
"version": "1.14.0",
"resolved": "https://registry.npmjs.org/@azure/core-util/-/core-util-1.14.0.tgz",
"integrity": "sha512-9n2pWK61veAuN0V20t9lOuoV4CFMdyAZ1ygZzvBGk/pBBJRib/PjL9PLXa/aI2CcPpyHfqVsxxqLCYl6uZlfDw==",
"license": "MIT",
"dependencies": {
"@azure/abort-controller": "^2.1.2",
"@typespec/ts-http-runtime": "^0.3.0",
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/@babel/code-frame": {
"version": "7.26.2",
"resolved": "https://registry.npmjs.org/@babel/code-frame/-/code-frame-7.26.2.tgz",
@@ -2391,11 +2429,19 @@
"resolved": "https://registry.npmjs.org/@types/webrtc/-/webrtc-0.0.37.tgz",
"integrity": "sha512-JGAJC/ZZDhcrrmepU4sPLQLIOIAgs5oIK+Ieq90K8fdaNMhfdfqmYatJdgif1NDQtvrSlTOGJDUYHIDunuufOg=="
},
"node_modules/23": {
"version": "0.0.0",
"resolved": "https://registry.npmjs.org/23/-/23-0.0.0.tgz",
"integrity": "sha512-uAETf9Okr72trtp1pNXYKhFCTTI1EKGcYMA8gw3jLGhlbaDX+grrNToEWrpt8luxRAvrZWBTvyB5wk3PpNjGQQ==",
"license": "ISC"
"node_modules/@typespec/ts-http-runtime": {
"version": "0.3.8",
"resolved": "https://registry.npmjs.org/@typespec/ts-http-runtime/-/ts-http-runtime-0.3.8.tgz",
"integrity": "sha512-bLMpVcWZNzq6lYOybwFwOAR1IXKcHnhUNqYeHjl1bET/qE3jFPFH+p8Wrh3rU4xwdnifPxmKNESBYnvnmc75aA==",
"license": "MIT",
"dependencies": {
"http-proxy-agent": "^7.0.0",
"https-proxy-agent": "^7.0.0",
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/abort-controller": {
"version": "3.0.0",
@@ -5278,17 +5324,18 @@
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk": {
"version": "1.38.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.38.0.tgz",
"integrity": "sha512-NA6J4eIDkeR9iN83rcn77Kn5AWQcizDEn1tLMjzRvSovUNB1FrZe0mWYO0fsGltUwMl3Ns5OZ3lGw42PU4fEYA==",
"version": "1.51.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.51.0.tgz",
"integrity": "sha512-BLLovv5PegOr5Lp52h4CSgL2c/ViFAh9LoU49ILNPyb/Z0giGI7nPV3xNLcmTtHLT+kfYNRGUIA62vORLzP8TA==",
"license": "MIT",
"dependencies": {
"@azure/core-auth": "^1.9.0",
"@types/webrtc": "^0.0.37",
"agent-base": "^6.0.1",
"bent": "^7.3.12",
"https-proxy-agent": "^4.0.0",
"uuid": "^9.0.0",
"ws": "^7.5.6"
"uuid": "^11.1.1",
"ws": "^8.21.0"
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk/node_modules/https-proxy-agent": {
@@ -5312,36 +5359,16 @@
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk/node_modules/uuid": {
"version": "9.0.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-9.0.1.tgz",
"integrity": "sha512-b+1eJOlsR9K8HJpow9Ok3fiWOWSIcIzXodvv0rQjVoOVNpWMpxf1wZNpt4y9h10odCNrqnYp1OBzRktckBe3sA==",
"version": "11.1.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.1.tgz",
"integrity": "sha512-vIYxrBCC/N/K+Js3qSN88go7kIfNPssr/hHCesKCQNAjmgvYS2oqr69kIufEG+O4+PfezOH4EbIeHCfFov8ZgQ==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"bin": {
"uuid": "dist/bin/uuid"
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk/node_modules/ws": {
"version": "7.5.10",
"resolved": "https://registry.npmjs.org/ws/-/ws-7.5.10.tgz",
"integrity": "sha512-+dbF1tHwZpXcbOJdVOkzLDxZP1ailvSxM6ZweXTegylPny803bFhA+vqBYw4s31NSAk4S2Qz+AKXK9a4wkdjcQ==",
"license": "MIT",
"engines": {
"node": ">=8.3.0"
},
"peerDependencies": {
"bufferutil": "^4.0.1",
"utf-8-validate": "^5.0.2"
},
"peerDependenciesMeta": {
"bufferutil": {
"optional": true
},
"utf-8-validate": {
"optional": true
}
"bin": {
"uuid": "dist/esm/bin/uuid"
}
},
"node_modules/mime-db": {
@@ -7031,15 +7058,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/undici": {
"version": "7.24.5",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.24.5.tgz",
"integrity": "sha512-3IWdCpjgxp15CbJnsi/Y9TCDE7HWVN19j1hmzVhoAkY/+CJx449tVxT5wZc1Gwg8J+P0LWvzlBzxYRnHJ+1i7Q==",
"license": "MIT",
"engines": {
"node": ">=20.18.1"
}
},
"node_modules/undici-types": {
"version": "5.26.5",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-5.26.5.tgz",
@@ -7342,11 +7360,6 @@
}
},
"dependencies": {
"23": {
"version": "0.0.0",
"resolved": "https://registry.npmjs.org/23/-/23-0.0.0.tgz",
"integrity": "sha512-uAETf9Okr72trtp1pNXYKhFCTTI1EKGcYMA8gw3jLGhlbaDX+grrNToEWrpt8luxRAvrZWBTvyB5wk3PpNjGQQ=="
},
"@aashutoshrathi/word-wrap": {
"version": "1.2.6",
"resolved": "https://registry.npmjs.org/@aashutoshrathi/word-wrap/-/word-wrap-1.2.6.tgz",
@@ -7924,6 +7937,34 @@
"resolved": "https://registry.npmjs.org/@aws/lambda-invoke-store/-/lambda-invoke-store-0.2.4.tgz",
"integrity": "sha512-iY8yvjE0y651BixKNPgmv1WrQc+GZ142sb0z4gYnChDDY2YqI4P/jsSopBWrKfAt7LOJAkOXt7rC/hms+WclQQ=="
},
"@azure/abort-controller": {
"version": "2.2.0",
"resolved": "https://registry.npmjs.org/@azure/abort-controller/-/abort-controller-2.2.0.tgz",
"integrity": "sha512-fNAjWnA/nZ2jz31kxR/AqRaUT8ewHBw/WuBIosK0moMy1C9e5ValbDfFdIxJzVOOYaYkV/b2F1S4H/aHiqfVQg==",
"requires": {
"tslib": "^2.6.2"
}
},
"@azure/core-auth": {
"version": "1.11.0",
"resolved": "https://registry.npmjs.org/@azure/core-auth/-/core-auth-1.11.0.tgz",
"integrity": "sha512-IUZydyTUkDnYdstOW9pFOOUQlBjAepK5teihDE3x6yxsPJs/hsAaaYpeGxdxrgtOiJbBKSjKW7MDk7AEhb4LRg==",
"requires": {
"@azure/abort-controller": "^2.1.2",
"@azure/core-util": "^1.13.0",
"tslib": "^2.6.2"
}
},
"@azure/core-util": {
"version": "1.14.0",
"resolved": "https://registry.npmjs.org/@azure/core-util/-/core-util-1.14.0.tgz",
"integrity": "sha512-9n2pWK61veAuN0V20t9lOuoV4CFMdyAZ1ygZzvBGk/pBBJRib/PjL9PLXa/aI2CcPpyHfqVsxxqLCYl6uZlfDw==",
"requires": {
"@azure/abort-controller": "^2.1.2",
"@typespec/ts-http-runtime": "^0.3.0",
"tslib": "^2.6.2"
}
},
"@babel/code-frame": {
"version": "7.26.2",
"resolved": "https://registry.npmjs.org/@babel/code-frame/-/code-frame-7.26.2.tgz",
@@ -9118,6 +9159,16 @@
"resolved": "https://registry.npmjs.org/@types/webrtc/-/webrtc-0.0.37.tgz",
"integrity": "sha512-JGAJC/ZZDhcrrmepU4sPLQLIOIAgs5oIK+Ieq90K8fdaNMhfdfqmYatJdgif1NDQtvrSlTOGJDUYHIDunuufOg=="
},
"@typespec/ts-http-runtime": {
"version": "0.3.8",
"resolved": "https://registry.npmjs.org/@typespec/ts-http-runtime/-/ts-http-runtime-0.3.8.tgz",
"integrity": "sha512-bLMpVcWZNzq6lYOybwFwOAR1IXKcHnhUNqYeHjl1bET/qE3jFPFH+p8Wrh3rU4xwdnifPxmKNESBYnvnmc75aA==",
"requires": {
"http-proxy-agent": "^7.0.0",
"https-proxy-agent": "^7.0.0",
"tslib": "^2.6.2"
}
},
"abort-controller": {
"version": "3.0.0",
"resolved": "https://registry.npmjs.org/abort-controller/-/abort-controller-3.0.0.tgz",
@@ -11106,16 +11157,17 @@
"integrity": "sha512-/IXtbwEk5HTPyEwyKX6hGkYXxM9nbj64B+ilVJnC/R6B0pH5G4V3b0pVbL7DBj4tkhBAppbQUlf6F6Xl9LHu1g=="
},
"microsoft-cognitiveservices-speech-sdk": {
"version": "1.38.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.38.0.tgz",
"integrity": "sha512-NA6J4eIDkeR9iN83rcn77Kn5AWQcizDEn1tLMjzRvSovUNB1FrZe0mWYO0fsGltUwMl3Ns5OZ3lGw42PU4fEYA==",
"version": "1.51.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.51.0.tgz",
"integrity": "sha512-BLLovv5PegOr5Lp52h4CSgL2c/ViFAh9LoU49ILNPyb/Z0giGI7nPV3xNLcmTtHLT+kfYNRGUIA62vORLzP8TA==",
"requires": {
"@azure/core-auth": "^1.9.0",
"@types/webrtc": "^0.0.37",
"agent-base": "^6.0.1",
"bent": "^7.3.12",
"https-proxy-agent": "^4.0.0",
"uuid": "^9.0.0",
"ws": "^7.5.6"
"uuid": "^11.1.1",
"ws": "^8.21.0"
},
"dependencies": {
"https-proxy-agent": {
@@ -11135,15 +11187,9 @@
}
},
"uuid": {
"version": "9.0.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-9.0.1.tgz",
"integrity": "sha512-b+1eJOlsR9K8HJpow9Ok3fiWOWSIcIzXodvv0rQjVoOVNpWMpxf1wZNpt4y9h10odCNrqnYp1OBzRktckBe3sA=="
},
"ws": {
"version": "7.5.10",
"resolved": "https://registry.npmjs.org/ws/-/ws-7.5.10.tgz",
"integrity": "sha512-+dbF1tHwZpXcbOJdVOkzLDxZP1ailvSxM6ZweXTegylPny803bFhA+vqBYw4s31NSAk4S2Qz+AKXK9a4wkdjcQ==",
"requires": {}
"version": "11.1.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.1.tgz",
"integrity": "sha512-vIYxrBCC/N/K+Js3qSN88go7kIfNPssr/hHCesKCQNAjmgvYS2oqr69kIufEG+O4+PfezOH4EbIeHCfFov8ZgQ=="
}
}
},
@@ -12347,11 +12393,6 @@
"which-boxed-primitive": "^1.0.2"
}
},
"undici": {
"version": "7.24.5",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.24.5.tgz",
"integrity": "sha512-3IWdCpjgxp15CbJnsi/Y9TCDE7HWVN19j1hmzVhoAkY/+CJx449tVxT5wZc1Gwg8J+P0LWvzlBzxYRnHJ+1i7Q=="
},
"undici-types": {
"version": "5.26.5",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-5.26.5.tgz",
+3 -5
View File
@@ -1,6 +1,6 @@
{
"name": "@jambonz/speech-utils",
"version": "1.0.11",
"version": "1.0.18",
"description": "TTS-related speech utilities for jambonz",
"main": "index.js",
"author": "Dave Horton",
@@ -25,7 +25,6 @@
},
"homepage": "https://github.com/jambonz/speech-utils#readme",
"dependencies": {
"23": "^0.0.0",
"@aws-sdk/client-polly": "^3.496.0",
"@aws-sdk/client-sts": "^3.496.0",
"@cartesia/cartesia-js": "^2.2.7",
@@ -35,9 +34,8 @@
"bent": "^7.3.12",
"debug": "^4.3.4",
"google-protobuf": "^3.21.2",
"microsoft-cognitiveservices-speech-sdk": "1.38.0",
"openai": "^4.98.0",
"undici": "^7.5.0"
"microsoft-cognitiveservices-speech-sdk": "^1.51.0",
"openai": "^4.98.0"
},
"devDependencies": {
"config": "^4.2.0",
+134 -1
View File
@@ -1058,6 +1058,76 @@ test('Cartesia speech synth tests', async(t) => {
client.quit();
});
test('gradium speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GRADIUM_API_KEY) {
t.pass('skipping gradium speech synth tests since GRADIUM_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'gradium',
credentials: {
api_key: process.env.GRADIUM_API_KEY,
model_id: 'default'
},
voice: 'YTpq7expH9539ERJ',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed gradium audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('fishaudio speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.FISHAUDIO_API_KEY) {
t.pass('skipping fishaudio speech synth tests since FISHAUDIO_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
/* voice 'default' means "send no reference_id" — fish's own default voice */
const o = await synthAudio(stats, {
vendor: 'fishaudio',
credentials: {
api_key: process.env.FISHAUDIO_API_KEY,
model_id: 's2.1-pro'
},
voice: 'default',
text,
renderForCaching: true
});
t.ok(!o.servedFromCache, `successfully synthed fishaudio audio to ${o.filePath}`);
/* the cache render must be raw 8k pcm (r8): fish's wav header carries a
placeholder RIFF size, so we never ask for wav */
const o2 = await synthAudio(stats, {
vendor: 'fishaudio',
credentials: {api_key: process.env.FISHAUDIO_API_KEY},
voice: 'default',
text: text + ' two',
renderForCaching: true,
disableTtsCache: true
});
t.ok(!o2.servedFromCache, 'fishaudio synthed a second uncached render');
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('nineninesix speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -1075,7 +1145,7 @@ test('nineninesix speech synth tests', async(t) => {
model_id: 'gepard-1.0'
},
language: 'en',
voice: '3ad7a827-7fd1-4954-bf35-47d4cc33d9ed',
voice: '775e4dfd-819c-4325-92ec-250f487ef7e3',
text,
renderForCaching: true
});
@@ -1118,6 +1188,69 @@ test('inworld speech synth', async(t) => {
client.quit();
});
test('inworld streaming say: params', async(t) => {
/* This test asserts the streaming say: path, so it must run with streaming
enabled. The Google non-streaming test above sets
JAMBONES_DISABLE_TTS_STREAMING and, on its no-credentials skip path,
deletes the env var WITHOUT clearing the require cache — so lib/config can
still be holding 'true' by the time we get here. Re-require to be
independent of what ran before us.
*/
delete process.env.JAMBONES_DISABLE_TTS_STREAMING;
delete require.cache[require.resolve('../lib/config')];
delete require.cache[require.resolve('../lib/synth-audio')];
delete require.cache[require.resolve('..')];
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.INWORLD_API_KEY) {
t.pass('skipping inworld streaming say: param tests since INWORLD_API_KEY is not provided');
client.quit();
return t.end();
}
try {
let result = await synthAudio(stats, {
vendor: 'inworld',
credentials: {api_key: process.env.INWORLD_API_KEY, model_id: 'inworld-tts-1.5-mini'},
language: 'en',
voice: 'Ashley',
text: 'This is a test of inworld streaming.',
options: {temperature: 0.9, audioConfig: {pitch: 2.5, speakingRate: 1.2}},
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'inworld returns streaming say: path');
t.ok(result.filePath.includes('vendor=inworld'), 'streaming path contains vendor=inworld');
t.ok(result.filePath.includes('voice=Ashley'), 'streaming path contains voice');
t.ok(result.filePath.includes('model_id=inworld-tts-1.5-mini'), 'streaming path contains model_id');
t.ok(result.filePath.includes('temperature=0.9'), 'streaming path contains temperature');
/* pitch and speakingRate are nested under audioConfig; they used to be read
from the top level and emitted as "undefined"
*/
t.ok(result.filePath.includes('speakingRate=1.2'), 'audioConfig.speakingRate reaches the say: params');
t.ok(result.filePath.includes('pitch=2.5'), 'audioConfig.pitch reaches the say: params');
t.ok(!result.filePath.includes('undefined'), 'no undefined values in the say: params');
/* options omitted entirely: no stray keys */
result = await synthAudio(stats, {
vendor: 'inworld',
credentials: {api_key: process.env.INWORLD_API_KEY, model_id: 'inworld-tts-1.5-mini'},
language: 'en',
voice: 'Ashley',
text: 'This is a test of inworld streaming.',
disableTtsCache: true
});
t.ok(!result.filePath.includes('speakingRate='), 'speakingRate omitted when unset');
t.ok(!result.filePath.includes('pitch='), 'pitch omitted when unset');
t.ok(!result.filePath.includes('undefined'), 'no undefined values when options are omitted');
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('resemble speech synth', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);