Compare commits

...
27 Commits
Author SHA1 Message Date
Dave Horton 5dc26c916d 1.0.21 2026-09-30 23:04:20 -04:00
Dave Horton 12e3364461 speechify 2026-09-30 23:03:54 -04:00
Hoan Luu HuuandClaude Opus 5.5 6c67f6faa0 feat(tts): add speechify as a TTS vendor (#165)
- streaming returns a say: url for the mediajam starter
- cache render POSTs /v1/audio/stream for raw 24 kHz PCM (r24)
- credential options merge under the verb's options
- no model: simba-3.2 for English, simba-3.0 otherwise

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:01:18 -04:00
Dave Horton 18917f49ff 1.0.19 2026-09-29 14:22:33 -04:00
Dave HortonandClaude Opus 5.5 fa705694e1 feat(tts): add kugelaudio as a TTS vendor (#163)
Streaming returns a say: url for the mediajam dialect; the cache render POSTs
/v1/tts/generate for raw 8 kHz PCM (r8). Language is reduced to its primary
subtag and dropped when KugelAudio does not support it.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:21:41 -04:00
Dave HortonandClaude Opus 5 89fe0b87e9 fix: stop the RoleArn synth test printing AWS credentials into CI logs (#161)
With renderForCaching unset, synthPolly returns the mediajam streaming params
string, and lib/synth-audio.js:337-340 embeds accessKeyId, secretAccessKey and
sessionToken in it. The test interpolated that value into its assertion message,
so run 32995180996 printed live (1 hour) AWS credentials into a public build log.

Every other synth test in this file passes renderForCaching: true; this one was
the only exception. It now does the same, and the assertion no longer interpolates
opts.filePath at all, so the streaming params string cannot leak this way again.

Latent since the test was written -- it only surfaced once AWS_ROLE_ARN was wired
up and the test actually ran.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 13:44:05 -04:00
Dave HortonandClaude Opus 5 b928820906 ci: authenticate to AWS via GitHub OIDC instead of stored access keys (#160)
* ci: authenticate to AWS via GitHub OIDC instead of stored access keys

CI held a long-lived AWS access key as repository secrets. This replaces it with
a short-lived credential obtained through the GitHub OIDC provider, so no AWS key
is stored in the repo at all.

The tests could not simply inherit the OIDC credentials. They passed
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY explicitly, which routes
lib/get-aws-sts-token.js down its accessKeyId branch and calls GetSessionToken --
and AWS rejects GetSessionToken when it is called with session credentials.

test/aws-credentials.js centralises the decision. When AWS_SESSION_TOKEN is present
the credentials are temporary and only the region is passed, so the SDK's default
credential provider chain is used. Static keys still work unchanged, which keeps
local run-tests.sh working and also lets it run off an 'aws sso login' session with
no credentials in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: supply a cache key when AWS credentials come from the default chain

getAwsAuthToken derives its cache key as roleArn || accessKeyId || speech_credential_sid.
With temporary credentials the test passes none of those, so makeAwsKey received
undefined and hash.update() threw ERR_INVALID_ARG_TYPE.

Production never hits this: the instance-profile path always carries a
speech_credential_sid from the database, which is exactly what the comment above
that line describes. The test now supplies one the same way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: reference the CI role via secret rather than inlining the account id

speech-utils is public, so the role ARN (and with it the AWS account id) should not
be committed. It now comes from the AWS_ROLE_ARN secret, which also feeds the
'AWS speech synth tests by RoleArn' test -- previously always skipped, so the
AssumeRole credential path had no coverage here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: give the AWS RoleArn synth test its own cache key

The RoleArn test synthesized the same vendor/voice/language/text as the plain AWS
synth test that runs before it, so it hit that test's cache entry, servedFromCache
came back true and the !servedFromCache assertion failed.

Latent since the test was written -- it never ran, because AWS_ROLE_ARN was never
supplied. Wiring the secret in activated it and exposed the collision. Distinct text
makes the test independent of execution order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 13:38:32 -04:00
Dave Horton f4c3c7dc8b 1.0.18 2026-08-23 13:12:37 -04:00
Dave HortonandClaude Opus 5 a32f71ac5a feat: add fishaudio (Fish Audio) TTS support (#159)
Adds synthFishaudio with both arms: the say: streaming url consumed by the
mediajam dialect, and a POST /v1/tts cache render. The render asks for raw pcm
at 8k and returns extension r8 because fish's wav output carries a placeholder
RIFF size, the same problem gradium has.

Fish is a voice-cloning vendor, so the voice is a reference_id; the sentinel
'default' means send none and use fish's own default voice.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 13:11:59 -04:00
Dave HortonandClaude Opus 5 c8850998b5 fix(tts): use a current nineninesix voice id in the synth test (#158)
The vendor replaced its voice catalog and now rejects unknown ids, so
the env-gated test failed whenever NINENINESIX_API_KEY was set.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:19:59 -04:00
Dave Horton e8b2009d29 1.0.17 2026-08-10 16:45:32 -04:00
Dave HortonandClaude Opus 5 27e07bb551 fix(tts): forward gradium json_config to the streaming path (#157)
synthGradium destructured json_config from options but only forwarded
pronunciation_id into the say: param block, so the setting could never
reach mediajam's streaming dialect — only the cache-render POST. The
say: parser is brace-aware, so a nested json object survives intact.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 16:45:07 -04:00
Dave Horton 33456b93d9 1.0.16 2026-08-08 23:07:43 -04:00
Dave HortonandClaude Opus 5 70e91b5057 test(inworld): gate the say: param test on INWORLD_API_KEY (#156)
The test I added in #155 ran unconditionally, which broke `npm test` without
credentials — the path husky's pre-commit hook takes, so `npm version patch`
could not commit.

Two causes, both addressed:

- no credential gate, unlike every other vendor test in this file. Now skips
  without INWORLD_API_KEY, and closes its redis client on that path so the
  run can still exit.
- it assumed streaming was enabled. The Google non-streaming test sets
  JAMBONES_DISABLE_TTS_STREAMING and, on its no-credentials skip path,
  deletes the env var WITHOUT clearing the require cache (unlike its finally
  block, which clears both) — so lib/config still held 'true' further down
  the file and synthInworld took the non-streaming branch, attempting a real
  vendor call. The test now re-requires with streaming enabled so it does not
  depend on what ran before it.

Verified both ways: skips and exits 0 with no key; 11/11 with a key even
under the leaked state. Re-introducing the #155 bug still fails 3 assertions,
so the regression value is intact.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:06:52 -04:00
Dave HortonandClaude Opus 5 c066840f85 fix(inworld): read pitch and speakingRate from audioConfig in the say: params (#155)
The streaming say: path guarded on opts.audioConfig?.pitch and
opts.audioConfig?.speakingRate but interpolated opts.pitch and
opts.speakingRate, which are undefined — so anyone setting them under
audioConfig (what the docs and the portal defaults tell you to do) got
'pitch=undefined,speakingRate=undefined' on the wire and their setting
silently dropped.

Adds a test for the say: params that needs no credentials, since the
streaming branch builds the path without calling the vendor.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:49:50 -04:00
Dave Horton d89457f67d 1.0.15 2026-08-06 10:23:47 -04:00
Dave HortonandClaude Opus 5 3c15976669 fix(deps): remove junk "23" dependency (#154)
"23": "^0.0.0" is not a real dependency - it is an empty placeholder
package (0.0.0, no deps, unrelated third-party maintainer) that landed
here from a stray npm install. Nothing in the package references it.

Beyond the noise, it is a small supply-chain liability: a dependency on
a squatted single-number name owned by nobody we know, shipped to every
consumer of speech-utils.

Lint passes; full test suite passes 105/105 with live vendor credentials.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:23:31 -04:00
Dave Horton 70abb04941 1.0.14 2026-08-06 10:13:51 -04:00
Dave HortonandClaude Opus 5 724af7fcab fix(deps): drop unused undici dependency (#153)
undici was declared as a direct dependency but is never required
anywhere in this package - grep across index.js, lib/ and stubs/ finds
no reference to undici, ProxyAgent, setGlobalDispatcher or Dispatcher.
The one HTTP call in lib/synth-audio.js uses the global fetch, and
Azure proxy support goes through the SDK's own setProxy plus the
http_proxy_ip/http_proxy_port params.

Removing it clears all seven open undici advisories from npm audit for
this package and its consumers, and stops speech-utils pulling a
duplicate undici 7.x into trees where the consumer already depends on
undici 8 (feature-server does).

Lint passes; full test suite passes 105/105 with live vendor credentials.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:13:31 -04:00
Dave Horton c0a79090c2 1.0.13 2026-08-06 10:04:26 -04:00
Dave HortonandClaude Opus 5 7a9a3e2468 fix(deps): bump microsoft-cognitiveservices-speech-sdk to ^1.51.0 (#152)
The Azure speech SDK was pinned to exactly 1.38.0 since the initial
commit. It pulls in uuid <11.1.1, which carries GHSA-w5hq-g745-h8pq
(missing buffer bounds check in v3/v5/v6 when buf is provided). The
advisory covers Azure SDK 1.14.0-1.50.0; 1.51.0 clears it.

Loosened the exact pin to a caret range so future patches come in
without another PR. The SDK surface used here (SpeechConfig,
SpeechSynthesizer, ResultReason, CancellationDetails,
SpeechSynthesisOutputFormat) is unchanged in 1.51.0.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:03:59 -04:00
Dave Horton ab2b3f7fb2 1.0.12 2026-08-04 09:04:50 -04:00
Dave HortonandClaude Opus 5 e848d0276e feat(tts): add gradium as a TTS vendor (#151)
Streaming arm returns a say: url for the mediajam dialect; the cache-render
arm posts to /api/post/speech/tts with only_audio and pcm_8000, which is bare
r8 samples and avoids gradium's streaming wav header (0xffffffff RIFF size).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 09:02:07 -04:00
Dave Horton d2447d477d 1.0.11 2026-08-03 12:04:47 -04:00
Dave HortonandClaude Opus 5 ec8bbacc2c feat(tts): add nineninesix.ai synthesis (#150)
Streaming goes through mediajam's say: url; the cache render posts to
/tts/bytes for wav, since the service rejects mp3.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 12:03:52 -04:00
Dave Horton 9694714a7f 1.0.10 2026-07-28 20:59:05 -04:00
Hoan Luu Huu 6a6de9b7a1 support google tts configuraiton (#149) 2026-07-28 20:58:26 -04:00
9 changed files with 1053 additions and 117 deletions
+12 -3
View File
@@ -7,11 +7,18 @@ on:
jobs:
build:
runs-on: ubuntu-latest
permissions:
id-token: write # required to request the GitHub OIDC token for AWS
contents: read
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '20'
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
aws-region: us-east-1
- run: npm install
- run: npm run jslint
- run: sudo apt update && sudo apt install -y squid
@@ -19,10 +26,12 @@ jobs:
- run: sudo systemctl start squid
- run: npm test
env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_REGION: ${{ secrets.AWS_REGION }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
# AWS_* are exported by configure-aws-credentials above as short-lived
# OIDC credentials; no AWS keys are stored as repository secrets.
GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }}
# Enables the "AWS speech synth tests by RoleArn" test, which exercises the
# AssumeRole credential path used in production.
AWS_ROLE_ARN: ${{ secrets.AWS_ROLE_ARN }}
MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }}
MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }}
+84
View File
@@ -0,0 +1,84 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Overview
`@jambonz/speech-utils` is a Node.js library providing TTS (Text-to-Speech) utilities for the jambonz CPaaS platform. It handles speech synthesis with caching through Redis and supports multiple TTS vendors.
## Commands
```bash
# Run tests (requires Docker for Redis)
npm test
# Run linter
npm run jslint
# Auto-fix lint issues
npm run jslint:fix
# Generate coverage report
npm run coverage
```
## Testing
Tests use `tape` and require Redis. The test harness automatically starts/stops Redis via Docker Compose (`test/docker-compose-testbed.yaml`).
Most tests are conditional based on environment variables for vendor credentials:
- `GCP_FILE` or `GCP_JSON_KEY` - Google TTS
- `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_REGION` - AWS Polly
- `MICROSOFT_API_KEY`, `MICROSOFT_REGION` - Azure TTS
- `ELEVENLABS_API_KEY`, `ELEVENLABS_VOICE_ID`, `ELEVENLABS_MODEL_ID` - ElevenLabs
- `OPENAI_API_KEY` - OpenAI Whisper TTS
- And others per vendor
Redis config is in `config/test.json` (port 3379).
## Architecture
### Entry Point
`index.js` exports a factory function that takes Redis options and a logger, returning an object with these methods:
- `synthAudio` - Main synthesis function
- `getTtsVoices` - List available voices for a vendor
- `purgeTtsCache` / `getTtsSize` / `addFileToCache` - Cache management
- `getAwsAuthToken` - Token management
### Core Module: `lib/synth-audio.js`
The `synthAudio` function handles synthesis for all vendors. Key behaviors:
1. **Cache check**: Generates SHA1 hash key from (vendor, language, voice, engine, model, text, instructions)
2. **Streaming vs non-streaming**: When `JAMBONES_DISABLE_TTS_STREAMING` is not set and `renderForCaching=false`, returns `say:{params}text` format for FreeSWITCH streaming playback instead of generating files
3. **Vendor dispatch**: Switch statement routes to vendor-specific synth functions (`synthGoogle`, `synthPolly`, `synthMicrosoft`, etc.)
4. **Caching**: Stores audio as base64 JSON in Redis with configurable TTL (default 4 hours)
### Supported Vendors
google, aws/polly, microsoft/azure, nvidia (Riva), wellsaid, elevenlabs, cartesia, inworld, rimelabs, whisper (OpenAI), deepgram, resemble, custom:*
### gRPC Stubs
`stubs/riva/` contains generated protobuf/gRPC code for NVIDIA Riva.
## Environment Variables
Key configuration via env vars (see `lib/config.js`):
- `JAMBONES_DISABLE_TTS_STREAMING` - Force non-streaming mode
- `JAMBONES_DISABLE_AZURE_TTS_STREAMING` - Azure-specific streaming disable
- `JAMBONES_TTS_CACHE_DURATION_MINS` - Cache TTL in minutes (default: 240)
- `JAMBONES_TTS_TRIM_SILENCE` - Trim trailing silence from audio
- `JAMBONES_TMP_FOLDER` - Temp folder for audio files (default: /tmp)
- `JAMBONES_HTTP_PROXY_IP`, `JAMBONES_HTTP_PROXY_PORT` - HTTP proxy for Azure
- `JAMBONES_AZURE_ENABLE_SSML` - Force SSML wrapper for Azure plain text
## Key Dependencies
- `@jambonz/realtimedb-helpers` - Redis client and hash utilities
- `@google-cloud/text-to-speech` - Google TTS
- `@aws-sdk/client-polly` - AWS Polly
- `microsoft-cognitiveservices-speech-sdk` - Azure TTS
- `@grpc/grpc-js` - gRPC for Riva
- `openai` - OpenAI Whisper TTS
- `bent` - HTTP client for REST-based vendors
+436 -3
View File
@@ -80,7 +80,8 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
logger = logger || noopLogger;
assert.ok(['google', 'aws', 'polly', 'microsoft', 'wellsaid', 'nvidia', 'elevenlabs',
'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'inworld', 'resemble', 'murf', 'xai']
'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'gradium', 'kugelaudio', 'nineninesix', 'inworld',
'resemble', 'murf', 'xai', 'fishaudio', 'speechify']
.includes(vendor) ||
vendor.startsWith('custom'),
`synthAudio supported vendors are google, aws, microsoft, nvidia and wellsaid ..etc, not ${vendor}`);
@@ -135,12 +136,29 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
} else if ('cartesia' === vendor) {
assert.ok(credentials.api_key, 'synthAudio requires api_key when cartesia is used');
assert.ok(credentials.model_id, 'synthAudio requires model_id when cartesia is used');
} else if ('gradium' === vendor) {
assert.ok(voice, 'synthAudio requires voice when gradium is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when gradium is used');
} else if ('speechify' === vendor) {
assert.ok(voice, 'synthAudio requires voice when speechify is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when speechify is used');
} else if ('kugelaudio' === vendor) {
assert.ok(voice, 'synthAudio requires voice when kugelaudio is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when kugelaudio is used');
} else if ('nineninesix' === vendor) {
assert.ok(voice, 'synthAudio requires voice when nineninesix is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when nineninesix is used');
assert.ok(credentials.model_id, 'synthAudio requires model_id when nineninesix is used');
} else if ('murf' === vendor) {
assert.ok(voice, 'synthAudio requires voice when murf is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when murf is used');
} else if (vendor === 'resemble') {
assert.ok(voice, 'synthAudio requires voice when resemble is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when resemble is used');
} else if ('fishaudio' === vendor) {
/* no voice assert: fish synthesizes with its own default voice when
reference_id is omitted, which is what the 'default' selection means */
assert.ok(credentials.api_key, 'synthAudio requires api_key when fishaudio is used');
}
const key = makeSynthKey({
@@ -212,6 +230,31 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'fishaudio':
audioData = await synthFishaudio(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'gradium':
audioData = await synthGradium(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'kugelaudio':
audioData = await synthKugelaudio(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'speechify':
audioData = await synthSpeechify(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'nineninesix':
audioData = await synthNineninesix(logger, {
credentials, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'inworld':
audioData = await synthInworld(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
@@ -383,6 +426,32 @@ const synthPolly = async(createHash, retrieveHash, logger,
}
};
/* google AudioConfig settings we support, as [google camelCase name, freeswitch param name] */
const GOOGLE_AUDIO_SETTINGS = [
['speakingRate', 'speaking_rate'],
['pitch', 'pitch'],
['volumeGainDb', 'volume_gain_db']
];
/**
* Extract google AudioConfig settings from the synthesizer options. They may be supplied
* nested under an audioConfig property (mirroring google's AudioConfig object) or flat at
* the top level, and either google's camelCase or snake_case names are accepted.
* @see https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig
* @returns object keyed by google's camelCase names, holding only valid numeric settings
*/
const googleAudioConfig = (options) => {
const provided = {...options, ...(options?.audioConfig || {})};
const audioConfig = {};
for (const [name, snakeName] of GOOGLE_AUDIO_SETTINGS) {
const value = provided[name] ?? provided[snakeName];
/* note: 0 is meaningful for pitch and volumeGainDb, so check for absence explicitly */
if (value === undefined || value === null || value === '') continue;
const num = Number(value);
if (Number.isFinite(num)) audioConfig[name] = num;
}
return audioConfig;
};
const synthGoogle = async(logger, {
credentials, stats, language, voice, gender, key, text, model, options, instructions,
@@ -418,6 +487,15 @@ const synthGoogle = async(logger, {
// comma is used to separate parameters in freeswitch tts module
const prompt = options?.prompt || instructions;
if (prompt) params += `,prompt=${prompt.replace(/\n/g, ' ').replace(/,/g, ';')}`;
/**
* AudioConfig settings. Note google only honors these in some api modes:
* tts applies all of them, live (HD voices) applies speakingRate only,
* and gemini ignores them entirely (use prompt instead for style control).
*/
const audioSettings = googleAudioConfig(options);
for (const [name, snakeName] of GOOGLE_AUDIO_SETTINGS) {
if (name in audioSettings) params += `,${snakeName}=${audioSettings[name]}`;
}
params += '}';
return {
@@ -486,6 +564,9 @@ const synthGoogle = async(logger, {
sampleRate = 8000;
}
/* gemini voices do not support the AudioConfig settings; they use prompt for style control */
if (!isGemini) Object.assign(audioConfig, googleAudioConfig(options));
const opts = { input, voice: voiceParams, audioConfig };
try {
@@ -885,8 +966,9 @@ const synthInworld = async(logger, {
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (opts.temperature) params += `,temperature=${opts.temperature}`;
if (opts.audioConfig?.pitch) params += `,pitch=${opts.pitch}`;
if (opts.audioConfig?.speakingRate) params += `,speakingRate=${opts.speakingRate}`;
/* pitch and speakingRate are nested under audioConfig, matching Inworld's API */
if (opts.audioConfig?.pitch) params += `,pitch=${opts.audioConfig.pitch}`;
if (opts.audioConfig?.speakingRate) params += `,speakingRate=${opts.audioConfig.speakingRate}`;
params += '}';
return {
@@ -1385,6 +1467,357 @@ const synthCartesia = async(logger, {
};
/* gradium.ai — json websocket for streaming, and a POST endpoint for the cache
render. only_audio:true + pcm_8000 returns bare little-endian 16-bit samples,
which is exactly the r8 container, so we avoid gradium's streaming wav header
(it carries 0xffffffff as the RIFF size, since length is unknown up front).
*/
const synthGradium = async(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
const {api_key, model_id} = credentials;
const {json_config, pronunciation_id} = options || {};
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=gradium';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (model_id) params += `,model_id=${model_id}`;
if (pronunciation_id) params += `,pronunciation_id=${pronunciation_id}`;
/* the say: param parser is brace-aware, so a nested json object survives intact */
if (json_config) {
params += `,json_config=${typeof json_config === 'string' ? json_config : JSON.stringify(json_config)}`;
}
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const post = bent('https://api.gradium.ai', 'POST', 'buffer', {
'x-api-key': api_key,
'Content-Type': 'application/json'
});
const audioContent = await post('/api/post/speech/tts', {
text,
voice_id: voice,
model_name: model_id || 'default',
output_format: `pcm_${sampleRate}`,
only_audio: true,
...(json_config && {json_config}),
...(pronunciation_id && {pronunciation_id})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth gradium returned error');
stats.increment('tts.count', ['vendor:gradium', 'accepted:no']);
throw err;
}
};
/* kugelaudio — json websocket (/ws/tts/stream) for streaming, and POST /v1/tts/generate
for the cache render. the POST streams back bare little-endian 16-bit samples at the
requested sample_rate, which is exactly the r8 container at 8000.
voices are numeric ids (or public handles). language is an ISO 639-1 code that drives
text normalization; jambonz carries BCP-47, so only the primary subtag is sent. the
api rejects codes outside its list, so an unsupported or unset language is omitted
and the voice's own language applies. options.api_uri pins a region
(e.g. api.eu.kugelaudio.com).
*/
const KUGELAUDIO_LANGUAGES = ['ar', 'bg', 'bn', 'cs', 'da', 'de', 'el', 'en', 'es', 'fa', 'fi', 'fr', 'he', 'hi',
'hr', 'hu', 'id', 'it', 'ja', 'ko', 'ms', 'nl', 'no', 'pl', 'pt', 'ro', 'ru', 'sk', 'sl', 'sr', 'sv', 'ta', 'th',
'tr', 'uk', 'ur', 'vi', 'yue', 'zh'];
const synthKugelaudio = async(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache
}) => {
const {api_key, model_id} = credentials;
const {api_uri, speed, cfg_scale, temperature, normalize, project_id, dictionary_ids} = options || {};
const isSet = (v) => v !== null && v !== undefined;
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=kugelaudio';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
params += `,model_id=${model_id || 'kugel-3'}`;
if (language) params += `,language=${language}`;
if (api_uri) params += `,api_uri=${api_uri}`;
if (isSet(speed)) params += `,speed=${speed}`;
if (isSet(cfg_scale)) params += `,cfg_scale=${cfg_scale}`;
if (isSet(temperature)) params += `,temperature=${temperature}`;
if (isSet(normalize)) params += `,normalize=${normalize}`;
if (isSet(project_id)) params += `,project_id=${project_id}`;
/* the say: param parser is bracket-aware, so the json array survives intact */
if (Array.isArray(dictionary_ids)) params += `,dictionary_ids=${JSON.stringify(dictionary_ids)}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const host = (api_uri || 'api.kugelaudio.com').replace(/^[a-z]+:\/\//, '').replace(/\/$/, '');
const post = bent(`https://${host}`, 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json; charset=utf-8'
});
const voiceId = /^\d+$/.test(`${voice}`) ? Number(voice) : voice;
const lang = language && language.split('-')[0].toLowerCase();
const audioContent = await post('/v1/tts/generate', {
text,
voice_id: voiceId,
model_id: model_id || 'kugel-3',
sample_rate: sampleRate,
...(KUGELAUDIO_LANGUAGES.includes(lang) && {language: lang}),
...(isSet(speed) && {speed: Number(speed)}),
...(isSet(cfg_scale) && {cfg_scale: Number(cfg_scale)}),
...(isSet(temperature) && {temperature: Number(temperature)}),
...(isSet(normalize) && {normalize: normalize === true || normalize === 'true'}),
...(isSet(project_id) && {project_id: Number(project_id)}),
...(Array.isArray(dictionary_ids) && {dictionary_ids})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth kugelaudio returned error');
stats.increment('tts.count', ['vendor:kugelaudio', 'accepted:no']);
throw err;
}
};
/* simba-3.2 is English only; with no model set, other languages get simba-3.0 */
const speechifyModel = (model_id, language) => {
if (model_id) return model_id;
return !language || /^en/i.test(language) ? 'simba-3.2' : 'simba-3.0';
};
const synthSpeechify = async(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache
}) => {
const {api_key, model_id} = credentials;
let credOptions = credentials.options || {};
if (typeof credOptions === 'string') {
try {
credOptions = JSON.parse(credOptions);
} catch {
credOptions = {};
}
}
const {api_uri, loudness_normalization, text_normalization} = {...credOptions, ...options};
const isSet = (v) => v !== null && v !== undefined;
const model = speechifyModel(options?.model_id || model_id, language);
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=speechify';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
params += `,model_id=${model}`;
if (language) params += `,language=${language}`;
if (api_uri) params += `,api_uri=${api_uri}`;
if (isSet(loudness_normalization)) params += `,loudness_normalization=${loudness_normalization}`;
if (isSet(text_normalization)) params += `,text_normalization=${text_normalization}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
/* other pcm rates are mislabelled 24 kHz on workspaces pinned before 2026-09-30 */
const sampleRate = 24000;
const host = (api_uri || 'api.speechify.ai').replace(/^[a-z]+:\/\//, '').replace(/\/$/, '');
const post = bent(`https://${host}`, 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json',
'Speechify-Caller': 'jambonz'
});
const toBool = (v) => v === true || v === 'true';
const speechOptions = {
...(isSet(loudness_normalization) && {loudness_normalization: toBool(loudness_normalization)}),
...(isSet(text_normalization) && {text_normalization: toBool(text_normalization)})
};
const audioContent = await post('/v1/audio/stream', {
input: text,
voice_id: voice,
model,
output_format: `pcm_${sampleRate}`,
...(language && {language}),
...(Object.keys(speechOptions).length && {options: speechOptions})
});
return {
audioContent,
extension: 'r24',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth speechify returned error');
stats.increment('tts.count', ['vendor:speechify', 'accepted:no']);
throw err;
}
};
/* nineninesix.ai — a Cartesia-compatible API, but only raw/wav come back
(mp3 is rejected), so the cache render asks for wav rather than mp3. */
/* fish.audio — msgpack websocket for streaming, and a POST endpoint for the cache
render. format:pcm + sample_rate:8000 returns bare little-endian 16-bit samples,
which is exactly the r8 container. we avoid fish's wav output because its RIFF
header carries a placeholder size (0xffffff24) — length is unknown up front, as
with gradium.
fish is a voice-cloning vendor: the "voice" is a reference_id returned by
POST /model, and omitting it entirely synthesizes with fish's default voice.
the sentinel value 'default' (the bundled fallback entry in the portal) means
exactly that — send no reference_id.
*/
const synthFishaudio = async(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
const {api_key, model_id, fishaudio_tts_uri} = credentials;
const {reference_id, latency, chunk_length, speed, volume} = options || {};
/* free-text reference_id in the vendor options wins over the voice selector */
const refId = reference_id || (voice && voice !== 'default' ? voice : null);
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=fishaudio';
params += `,voice=${refId || 'default'}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (model_id) params += `,model_id=${model_id}`;
if (latency) params += `,latency=${latency}`;
if (chunk_length) params += `,chunk_length=${chunk_length}`;
if (speed) params += `,speed=${speed}`;
if (volume) params += `,volume=${volume}`;
if (fishaudio_tts_uri) params += `,endpoint=${fishaudio_tts_uri}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const post = bent(fishaudio_tts_uri || 'https://api.fish.audio', 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json',
/* the model is selected by header, not in the body */
'model': model_id || 's2.1-pro'
});
const audioContent = await post('/v1/tts', {
text,
format: 'pcm',
sample_rate: sampleRate,
...(refId && {reference_id: refId}),
...(latency && {latency}),
...(chunk_length && {chunk_length: parseInt(chunk_length, 10)}),
...((speed || volume) && {prosody: {
...(speed && {speed: parseFloat(speed)}),
...(volume && {volume: parseFloat(volume)})
}})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth fishaudio returned error');
stats.increment('tts.count', ['vendor:fishaudio', 'accepted:no']);
throw err;
}
};
const synthNineninesix = async(logger, {
credentials, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
const {api_key, model_id} = credentials;
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += `,model_id=${model_id}`;
params += ',vendor=nineninesix';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (language) params += `,language=${language}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const post = bent('https://api.nineninesix.ai', 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json'
});
const audioContent = await post('/tts/bytes', {
model_id,
transcript: text,
voice: {mode: 'id', id: voice},
...(language && {language}),
output_format: {
container: 'wav',
encoding: 'pcm_s16le',
sample_rate: sampleRate
}
});
return {
audioContent,
extension: 'wav',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth nineninesix returned error');
stats.increment('tts.count', ['vendor:nineninesix', 'accepted:no']);
throw err;
}
};
const synthResemble = async(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
+115 -74
View File
@@ -1,15 +1,14 @@
{
"name": "@jambonz/speech-utils",
"version": "1.0.9",
"version": "1.0.21",
"lockfileVersion": 2,
"requires": true,
"packages": {
"": {
"name": "@jambonz/speech-utils",
"version": "1.0.9",
"version": "1.0.21",
"license": "MIT",
"dependencies": {
"23": "^0.0.0",
"@aws-sdk/client-polly": "^3.496.0",
"@aws-sdk/client-sts": "^3.496.0",
"@cartesia/cartesia-js": "^2.2.7",
@@ -19,9 +18,8 @@
"bent": "^7.3.12",
"debug": "^4.3.4",
"google-protobuf": "^3.21.2",
"microsoft-cognitiveservices-speech-sdk": "1.38.0",
"openai": "^4.98.0",
"undici": "^7.5.0"
"microsoft-cognitiveservices-speech-sdk": "^1.51.0",
"openai": "^4.98.0"
},
"devDependencies": {
"config": "^4.2.0",
@@ -760,6 +758,46 @@
"node": ">=18.0.0"
}
},
"node_modules/@azure/abort-controller": {
"version": "2.2.0",
"resolved": "https://registry.npmjs.org/@azure/abort-controller/-/abort-controller-2.2.0.tgz",
"integrity": "sha512-fNAjWnA/nZ2jz31kxR/AqRaUT8ewHBw/WuBIosK0moMy1C9e5ValbDfFdIxJzVOOYaYkV/b2F1S4H/aHiqfVQg==",
"license": "MIT",
"dependencies": {
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/@azure/core-auth": {
"version": "1.11.0",
"resolved": "https://registry.npmjs.org/@azure/core-auth/-/core-auth-1.11.0.tgz",
"integrity": "sha512-IUZydyTUkDnYdstOW9pFOOUQlBjAepK5teihDE3x6yxsPJs/hsAaaYpeGxdxrgtOiJbBKSjKW7MDk7AEhb4LRg==",
"license": "MIT",
"dependencies": {
"@azure/abort-controller": "^2.1.2",
"@azure/core-util": "^1.13.0",
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/@azure/core-util": {
"version": "1.14.0",
"resolved": "https://registry.npmjs.org/@azure/core-util/-/core-util-1.14.0.tgz",
"integrity": "sha512-9n2pWK61veAuN0V20t9lOuoV4CFMdyAZ1ygZzvBGk/pBBJRib/PjL9PLXa/aI2CcPpyHfqVsxxqLCYl6uZlfDw==",
"license": "MIT",
"dependencies": {
"@azure/abort-controller": "^2.1.2",
"@typespec/ts-http-runtime": "^0.3.0",
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/@babel/code-frame": {
"version": "7.26.2",
"resolved": "https://registry.npmjs.org/@babel/code-frame/-/code-frame-7.26.2.tgz",
@@ -2391,11 +2429,19 @@
"resolved": "https://registry.npmjs.org/@types/webrtc/-/webrtc-0.0.37.tgz",
"integrity": "sha512-JGAJC/ZZDhcrrmepU4sPLQLIOIAgs5oIK+Ieq90K8fdaNMhfdfqmYatJdgif1NDQtvrSlTOGJDUYHIDunuufOg=="
},
"node_modules/23": {
"version": "0.0.0",
"resolved": "https://registry.npmjs.org/23/-/23-0.0.0.tgz",
"integrity": "sha512-uAETf9Okr72trtp1pNXYKhFCTTI1EKGcYMA8gw3jLGhlbaDX+grrNToEWrpt8luxRAvrZWBTvyB5wk3PpNjGQQ==",
"license": "ISC"
"node_modules/@typespec/ts-http-runtime": {
"version": "0.3.8",
"resolved": "https://registry.npmjs.org/@typespec/ts-http-runtime/-/ts-http-runtime-0.3.8.tgz",
"integrity": "sha512-bLMpVcWZNzq6lYOybwFwOAR1IXKcHnhUNqYeHjl1bET/qE3jFPFH+p8Wrh3rU4xwdnifPxmKNESBYnvnmc75aA==",
"license": "MIT",
"dependencies": {
"http-proxy-agent": "^7.0.0",
"https-proxy-agent": "^7.0.0",
"tslib": "^2.6.2"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/abort-controller": {
"version": "3.0.0",
@@ -5278,17 +5324,18 @@
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk": {
"version": "1.38.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.38.0.tgz",
"integrity": "sha512-NA6J4eIDkeR9iN83rcn77Kn5AWQcizDEn1tLMjzRvSovUNB1FrZe0mWYO0fsGltUwMl3Ns5OZ3lGw42PU4fEYA==",
"version": "1.51.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.51.0.tgz",
"integrity": "sha512-BLLovv5PegOr5Lp52h4CSgL2c/ViFAh9LoU49ILNPyb/Z0giGI7nPV3xNLcmTtHLT+kfYNRGUIA62vORLzP8TA==",
"license": "MIT",
"dependencies": {
"@azure/core-auth": "^1.9.0",
"@types/webrtc": "^0.0.37",
"agent-base": "^6.0.1",
"bent": "^7.3.12",
"https-proxy-agent": "^4.0.0",
"uuid": "^9.0.0",
"ws": "^7.5.6"
"uuid": "^11.1.1",
"ws": "^8.21.0"
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk/node_modules/https-proxy-agent": {
@@ -5312,36 +5359,16 @@
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk/node_modules/uuid": {
"version": "9.0.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-9.0.1.tgz",
"integrity": "sha512-b+1eJOlsR9K8HJpow9Ok3fiWOWSIcIzXodvv0rQjVoOVNpWMpxf1wZNpt4y9h10odCNrqnYp1OBzRktckBe3sA==",
"version": "11.1.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.1.tgz",
"integrity": "sha512-vIYxrBCC/N/K+Js3qSN88go7kIfNPssr/hHCesKCQNAjmgvYS2oqr69kIufEG+O4+PfezOH4EbIeHCfFov8ZgQ==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"bin": {
"uuid": "dist/bin/uuid"
}
},
"node_modules/microsoft-cognitiveservices-speech-sdk/node_modules/ws": {
"version": "7.5.10",
"resolved": "https://registry.npmjs.org/ws/-/ws-7.5.10.tgz",
"integrity": "sha512-+dbF1tHwZpXcbOJdVOkzLDxZP1ailvSxM6ZweXTegylPny803bFhA+vqBYw4s31NSAk4S2Qz+AKXK9a4wkdjcQ==",
"license": "MIT",
"engines": {
"node": ">=8.3.0"
},
"peerDependencies": {
"bufferutil": "^4.0.1",
"utf-8-validate": "^5.0.2"
},
"peerDependenciesMeta": {
"bufferutil": {
"optional": true
},
"utf-8-validate": {
"optional": true
}
"bin": {
"uuid": "dist/esm/bin/uuid"
}
},
"node_modules/mime-db": {
@@ -7031,15 +7058,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/undici": {
"version": "7.24.5",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.24.5.tgz",
"integrity": "sha512-3IWdCpjgxp15CbJnsi/Y9TCDE7HWVN19j1hmzVhoAkY/+CJx449tVxT5wZc1Gwg8J+P0LWvzlBzxYRnHJ+1i7Q==",
"license": "MIT",
"engines": {
"node": ">=20.18.1"
}
},
"node_modules/undici-types": {
"version": "5.26.5",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-5.26.5.tgz",
@@ -7342,11 +7360,6 @@
}
},
"dependencies": {
"23": {
"version": "0.0.0",
"resolved": "https://registry.npmjs.org/23/-/23-0.0.0.tgz",
"integrity": "sha512-uAETf9Okr72trtp1pNXYKhFCTTI1EKGcYMA8gw3jLGhlbaDX+grrNToEWrpt8luxRAvrZWBTvyB5wk3PpNjGQQ=="
},
"@aashutoshrathi/word-wrap": {
"version": "1.2.6",
"resolved": "https://registry.npmjs.org/@aashutoshrathi/word-wrap/-/word-wrap-1.2.6.tgz",
@@ -7924,6 +7937,34 @@
"resolved": "https://registry.npmjs.org/@aws/lambda-invoke-store/-/lambda-invoke-store-0.2.4.tgz",
"integrity": "sha512-iY8yvjE0y651BixKNPgmv1WrQc+GZ142sb0z4gYnChDDY2YqI4P/jsSopBWrKfAt7LOJAkOXt7rC/hms+WclQQ=="
},
"@azure/abort-controller": {
"version": "2.2.0",
"resolved": "https://registry.npmjs.org/@azure/abort-controller/-/abort-controller-2.2.0.tgz",
"integrity": "sha512-fNAjWnA/nZ2jz31kxR/AqRaUT8ewHBw/WuBIosK0moMy1C9e5ValbDfFdIxJzVOOYaYkV/b2F1S4H/aHiqfVQg==",
"requires": {
"tslib": "^2.6.2"
}
},
"@azure/core-auth": {
"version": "1.11.0",
"resolved": "https://registry.npmjs.org/@azure/core-auth/-/core-auth-1.11.0.tgz",
"integrity": "sha512-IUZydyTUkDnYdstOW9pFOOUQlBjAepK5teihDE3x6yxsPJs/hsAaaYpeGxdxrgtOiJbBKSjKW7MDk7AEhb4LRg==",
"requires": {
"@azure/abort-controller": "^2.1.2",
"@azure/core-util": "^1.13.0",
"tslib": "^2.6.2"
}
},
"@azure/core-util": {
"version": "1.14.0",
"resolved": "https://registry.npmjs.org/@azure/core-util/-/core-util-1.14.0.tgz",
"integrity": "sha512-9n2pWK61veAuN0V20t9lOuoV4CFMdyAZ1ygZzvBGk/pBBJRib/PjL9PLXa/aI2CcPpyHfqVsxxqLCYl6uZlfDw==",
"requires": {
"@azure/abort-controller": "^2.1.2",
"@typespec/ts-http-runtime": "^0.3.0",
"tslib": "^2.6.2"
}
},
"@babel/code-frame": {
"version": "7.26.2",
"resolved": "https://registry.npmjs.org/@babel/code-frame/-/code-frame-7.26.2.tgz",
@@ -9118,6 +9159,16 @@
"resolved": "https://registry.npmjs.org/@types/webrtc/-/webrtc-0.0.37.tgz",
"integrity": "sha512-JGAJC/ZZDhcrrmepU4sPLQLIOIAgs5oIK+Ieq90K8fdaNMhfdfqmYatJdgif1NDQtvrSlTOGJDUYHIDunuufOg=="
},
"@typespec/ts-http-runtime": {
"version": "0.3.8",
"resolved": "https://registry.npmjs.org/@typespec/ts-http-runtime/-/ts-http-runtime-0.3.8.tgz",
"integrity": "sha512-bLMpVcWZNzq6lYOybwFwOAR1IXKcHnhUNqYeHjl1bET/qE3jFPFH+p8Wrh3rU4xwdnifPxmKNESBYnvnmc75aA==",
"requires": {
"http-proxy-agent": "^7.0.0",
"https-proxy-agent": "^7.0.0",
"tslib": "^2.6.2"
}
},
"abort-controller": {
"version": "3.0.0",
"resolved": "https://registry.npmjs.org/abort-controller/-/abort-controller-3.0.0.tgz",
@@ -11106,16 +11157,17 @@
"integrity": "sha512-/IXtbwEk5HTPyEwyKX6hGkYXxM9nbj64B+ilVJnC/R6B0pH5G4V3b0pVbL7DBj4tkhBAppbQUlf6F6Xl9LHu1g=="
},
"microsoft-cognitiveservices-speech-sdk": {
"version": "1.38.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.38.0.tgz",
"integrity": "sha512-NA6J4eIDkeR9iN83rcn77Kn5AWQcizDEn1tLMjzRvSovUNB1FrZe0mWYO0fsGltUwMl3Ns5OZ3lGw42PU4fEYA==",
"version": "1.51.0",
"resolved": "https://registry.npmjs.org/microsoft-cognitiveservices-speech-sdk/-/microsoft-cognitiveservices-speech-sdk-1.51.0.tgz",
"integrity": "sha512-BLLovv5PegOr5Lp52h4CSgL2c/ViFAh9LoU49ILNPyb/Z0giGI7nPV3xNLcmTtHLT+kfYNRGUIA62vORLzP8TA==",
"requires": {
"@azure/core-auth": "^1.9.0",
"@types/webrtc": "^0.0.37",
"agent-base": "^6.0.1",
"bent": "^7.3.12",
"https-proxy-agent": "^4.0.0",
"uuid": "^9.0.0",
"ws": "^7.5.6"
"uuid": "^11.1.1",
"ws": "^8.21.0"
},
"dependencies": {
"https-proxy-agent": {
@@ -11135,15 +11187,9 @@
}
},
"uuid": {
"version": "9.0.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-9.0.1.tgz",
"integrity": "sha512-b+1eJOlsR9K8HJpow9Ok3fiWOWSIcIzXodvv0rQjVoOVNpWMpxf1wZNpt4y9h10odCNrqnYp1OBzRktckBe3sA=="
},
"ws": {
"version": "7.5.10",
"resolved": "https://registry.npmjs.org/ws/-/ws-7.5.10.tgz",
"integrity": "sha512-+dbF1tHwZpXcbOJdVOkzLDxZP1ailvSxM6ZweXTegylPny803bFhA+vqBYw4s31NSAk4S2Qz+AKXK9a4wkdjcQ==",
"requires": {}
"version": "11.1.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.1.tgz",
"integrity": "sha512-vIYxrBCC/N/K+Js3qSN88go7kIfNPssr/hHCesKCQNAjmgvYS2oqr69kIufEG+O4+PfezOH4EbIeHCfFov8ZgQ=="
}
}
},
@@ -12347,11 +12393,6 @@
"which-boxed-primitive": "^1.0.2"
}
},
"undici": {
"version": "7.24.5",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.24.5.tgz",
"integrity": "sha512-3IWdCpjgxp15CbJnsi/Y9TCDE7HWVN19j1hmzVhoAkY/+CJx449tVxT5wZc1Gwg8J+P0LWvzlBzxYRnHJ+1i7Q=="
},
"undici-types": {
"version": "5.26.5",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-5.26.5.tgz",
+3 -5
View File
@@ -1,6 +1,6 @@
{
"name": "@jambonz/speech-utils",
"version": "1.0.9",
"version": "1.0.21",
"description": "TTS-related speech utilities for jambonz",
"main": "index.js",
"author": "Dave Horton",
@@ -25,7 +25,6 @@
},
"homepage": "https://github.com/jambonz/speech-utils#readme",
"dependencies": {
"23": "^0.0.0",
"@aws-sdk/client-polly": "^3.496.0",
"@aws-sdk/client-sts": "^3.496.0",
"@cartesia/cartesia-js": "^2.2.7",
@@ -35,9 +34,8 @@
"bent": "^7.3.12",
"debug": "^4.3.4",
"google-protobuf": "^3.21.2",
"microsoft-cognitiveservices-speech-sdk": "1.38.0",
"openai": "^4.98.0",
"undici": "^7.5.0"
"microsoft-cognitiveservices-speech-sdk": "^1.51.0",
"openai": "^4.98.0"
},
"devDependencies": {
"config": "^4.2.0",
+23
View File
@@ -0,0 +1,23 @@
/**
* Resolve AWS credentials for the test suite.
*
* Returns null when AWS is not configured, so the caller skips.
*
* When AWS_SESSION_TOKEN is set the credentials are temporary -- GitHub OIDC in CI, or
* `aws sso login` locally -- and the key pair must not be passed through. Doing so sends
* lib/get-aws-sts-token.js down its accessKeyId branch, which calls GetSessionToken, and
* AWS rejects GetSessionToken when it is called with session credentials. Returning the
* region alone routes to the SDK's default credential provider chain instead, which
* handles temporary credentials correctly.
*/
module.exports = () => {
const region = process.env.AWS_REGION;
if (!region) return null;
const accessKeyId = process.env.AWS_ACCESS_KEY_ID;
const secretAccessKey = process.env.AWS_SECRET_ACCESS_KEY;
if (process.env.AWS_SESSION_TOKEN) return {region};
if (accessKeyId && secretAccessKey) return {accessKeyId, secretAccessKey, region};
return {region};
};
+11 -11
View File
@@ -6,33 +6,33 @@ process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
});
const awsCredentials = require('./aws-credentials');
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
test('AWS - create and cache auth token', async(t) => {
const fn = require('..');
const {client, getAwsAuthToken} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) {
const credentials = awsCredentials();
if (!credentials) {
t.pass('skipping AWS auth token tests since no AWS credentials provided');
t.end();
client.quit();
return;
}
// getAwsAuthToken derives its cache key from roleArn || accessKeyId || speech_credential_sid.
// With temporary credentials none of the first two are passed, so supply a stable sid --
// which is what production does for instance-profile credentials.
const args = {...credentials, speech_credential_sid: 'test-aws-speech-credential'};
try {
let obj = await getAwsAuthToken({
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
let obj = await getAwsAuthToken(args);
//console.log({obj}, 'received auth token from AWS');
t.ok(obj.securityToken && !obj.servedFromCache, 'successfullY generated auth token from AWS');
await sleep(250);
obj = await getAwsAuthToken({
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
obj = await getAwsAuthToken(args);
//console.log({obj}, 'received auth token from AWS - second request');
t.ok(obj.securityToken && obj.servedFromCache, 'successfully received access token from cache');
+5 -7
View File
@@ -1,4 +1,5 @@
const test = require('tape').test ;
const awsCredentials = require('./aws-credentials');
const config = require('config');
const opts = config.get('redis');
const fs = require('fs');
@@ -45,18 +46,15 @@ test('AWS tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) {
t.pass('skipping AWS speech synth tests since AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, or AWS_REGION not provided');
const credentials = awsCredentials();
if (!credentials) {
t.pass('skipping AWS speech synth tests since AWS_REGION not provided');
return t.end();
}
try {
const opts = {
vendor: 'aws',
credentials: {
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
}
credentials
};
let result = await getTtsVoices(opts);
t.ok(result?.Voices?.length > 0, `GetVoices: successfully retrieved ${result.Voices.length} voices from AWS`);
+364 -14
View File
@@ -1,4 +1,5 @@
const test = require('tape').test;
const awsCredentials = require('./aws-credentials');
const config = require('config');
const opts = config.get('redis');
const fs = require('fs');
@@ -442,6 +443,95 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
});
t.ok(result.filePath.includes('api_mode=tts'), 'options.apiMode=tts overrides HD voice default');
/* AudioConfig settings (speakingRate, pitch, volumeGainDb) */
const googleCreds = {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
};
// Test 9: AudioConfig nested under options.audioConfig, as in google's API docs
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'Testing nested audioConfig settings.',
options: { audioConfig: { speakingRate: 1.4, pitch: -2.5, volumeGainDb: 6 } },
disableTtsCache: true
});
t.ok(result.filePath.includes(',speaking_rate=1.4'), 'nested audioConfig sets speaking_rate');
t.ok(result.filePath.includes(',pitch=-2.5'), 'nested audioConfig sets pitch');
t.ok(result.filePath.includes(',volume_gain_db=6'), 'nested audioConfig sets volume_gain_db');
// Test 10: AudioConfig flat at the top level of options, in camelCase or snake_case
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'Testing flat audioConfig settings.',
options: { speaking_rate: '0.8', volumeGainDb: 3 },
disableTtsCache: true
});
t.ok(result.filePath.includes(',speaking_rate=0.8'), 'flat snake_case speaking_rate is honored');
t.ok(result.filePath.includes(',volume_gain_db=3'), 'flat camelCase volumeGainDb is honored');
t.ok(!result.filePath.includes(',pitch='), 'unspecified audioConfig setting is omitted');
// Test 11: zero is a meaningful value for pitch and volumeGainDb, not an absent one
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'Testing zero audioConfig settings.',
options: { audioConfig: { pitch: 0, volumeGainDb: 0 } },
disableTtsCache: true
});
t.ok(result.filePath.includes(',pitch=0'), 'pitch=0 is passed through rather than dropped');
t.ok(result.filePath.includes(',volume_gain_db=0'), 'volume_gain_db=0 is passed through rather than dropped');
// Test 12: non-numeric values are ignored, so they cannot corrupt the freeswitch param string
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'Testing invalid audioConfig settings.',
options: { audioConfig: { speakingRate: 'fast,evil=1', pitch: null, volumeGainDb: '' } },
disableTtsCache: true
});
t.ok(!result.filePath.includes('speaking_rate'), 'non-numeric speakingRate is ignored');
t.ok(!result.filePath.includes('evil=1'), 'non-numeric value cannot inject extra params');
t.ok(!result.filePath.includes('pitch='), 'null pitch is ignored');
t.ok(!result.filePath.includes('volume_gain_db'), 'empty volumeGainDb is ignored');
// Test 13: no audioConfig supplied leaves the param string untouched
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'Testing absent audioConfig settings.',
disableTtsCache: true
});
t.ok(!/speaking_rate|pitch=|volume_gain_db/.test(result.filePath),
'no audioConfig params are added when none are supplied');
// Test 14: HD voice (api_mode=live) also carries speaking_rate, the one setting google streams support
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Chirp3-HD-Charon',
text: 'Testing audioConfig on an HD voice.',
options: { audioConfig: { speakingRate: 1.25 } },
disableTtsCache: true
});
t.ok(result.filePath.includes('api_mode=live'), 'HD voice with audioConfig still uses api_mode=live');
t.ok(result.filePath.includes(',speaking_rate=1.25'), 'HD voice streaming path carries speaking_rate');
} catch (err) {
console.error(err);
t.end(err);
@@ -526,6 +616,42 @@ test('Google TTS non-streaming tests (JAMBONES_DISABLE_TTS_STREAMING=true)', asy
t.ok(!result.filePath.startsWith('say:'), 'Gemini TTS does NOT return streaming say: path when disabled');
t.ok(result.filePath.endsWith('.mp3'), 'Gemini TTS returns mp3 file path');
const googleCreds = {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
};
/**
* Test 4: AudioConfig settings are accepted by the synthesize API.
* Google rejects out-of-range values with a 400, so a successful render also confirms
* the settings reached the request rather than being silently dropped.
*/
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'This is a test of audioConfig with streaming disabled.',
options: { audioConfig: { speakingRate: 1.4, pitch: -2.5, volumeGainDb: 6 } },
disableTtsCache: true
});
t.ok(result.filePath.endsWith('.mp3'), 'standard voice renders mp3 with audioConfig settings applied');
/* Test 5: gemini voices ignore the AudioConfig settings rather than failing on them */
result = await synthAudio(stats, {
vendor: 'google',
credentials: googleCreds,
language: 'en-US',
voice: 'Kore',
model: geminiModel,
text: 'This is a test of audioConfig on Gemini TTS.',
options: { audioConfig: { speakingRate: 1.4, pitch: -2.5, volumeGainDb: 6 } },
disableTtsCache: true
});
t.ok(result.filePath.endsWith('.mp3'), 'gemini voice renders mp3 with audioConfig settings skipped');
} catch (err) {
console.error(err);
t.end(err);
@@ -543,18 +669,15 @@ test('AWS speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) {
t.pass('skipping AWS speech synth tests since AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, or AWS_REGION not provided');
const credentials = awsCredentials();
if (!credentials) {
t.pass('skipping AWS speech synth tests since AWS_REGION not provided');
return t.end();
}
try {
let opts = await synthAudio(stats, {
vendor: 'aws',
credentials: {
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
},
credentials,
language: 'en-US',
voice: 'Joey',
text: 'This is a test. This is only a test',
@@ -564,11 +687,7 @@ test('AWS speech synth tests', async(t) => {
opts = await synthAudio(stats, {
vendor: 'aws',
credentials: {
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
},
credentials,
language: 'en-US',
voice: 'Joey',
text: 'This is a test. This is only a test',
@@ -599,9 +718,18 @@ test('AWS speech synth tests by RoleArn', async(t) => {
},
language: 'en-US',
voice: 'Joey',
text: 'This is a test. This is only a test',
// Distinct text on purpose: the 'AWS speech synth tests' above cache audio for
// the same vendor/voice/language, and identical text would hit that cache entry,
// making servedFromCache true and this assertion fail.
text: 'This is a roleArn test. This is only a roleArn test',
// Without this, synthPolly returns the mediajam streaming params string, which
// embeds accessKeyId/secretAccessKey/sessionToken (see lib/synth-audio.js).
// Every other synth test in this file renders for caching; this one did not,
// so it printed live credentials into a public CI log.
renderForCaching: true,
});
t.ok(!opts.servedFromCache, `successfully synthesized aws by roleArn audio to ${opts.filePath}`);
// Deliberately does not interpolate opts.filePath -- it can carry credentials.
t.ok(!opts.servedFromCache, 'successfully synthesized aws by roleArn audio');
} catch (err) {
console.error(err);
t.end(err);
@@ -933,6 +1061,165 @@ test('Cartesia speech synth tests', async(t) => {
client.quit();
});
test('gradium speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GRADIUM_API_KEY) {
t.pass('skipping gradium speech synth tests since GRADIUM_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'gradium',
credentials: {
api_key: process.env.GRADIUM_API_KEY,
model_id: 'default'
},
voice: 'YTpq7expH9539ERJ',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed gradium audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('kugelaudio speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.KUGELAUDIO_API_KEY) {
t.pass('skipping kugelaudio speech synth tests since KUGELAUDIO_API_KEY is not provided');
return t.end();
}
const text = 'Guten Tag und willkommen bei jambonz! Ihre Bestellung kostet 12,99 Euro. ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'kugelaudio',
credentials: {
api_key: process.env.KUGELAUDIO_API_KEY,
model_id: 'kugel-3'
},
language: 'de-DE',
voice: '1930',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed kugelaudio audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('speechify speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.SPEECHIFY_API_KEY) {
t.pass('skipping speechify speech synth tests since SPEECHIFY_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambonz! ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'speechify',
credentials: {
api_key: process.env.SPEECHIFY_API_KEY
},
language: 'en-US',
voice: 'geffen_32',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed speechify audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('fishaudio speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.FISHAUDIO_API_KEY) {
t.pass('skipping fishaudio speech synth tests since FISHAUDIO_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
/* voice 'default' means "send no reference_id" — fish's own default voice */
const o = await synthAudio(stats, {
vendor: 'fishaudio',
credentials: {
api_key: process.env.FISHAUDIO_API_KEY,
model_id: 's2.1-pro'
},
voice: 'default',
text,
renderForCaching: true
});
t.ok(!o.servedFromCache, `successfully synthed fishaudio audio to ${o.filePath}`);
/* the cache render must be raw 8k pcm (r8): fish's wav header carries a
placeholder RIFF size, so we never ask for wav */
const o2 = await synthAudio(stats, {
vendor: 'fishaudio',
credentials: {api_key: process.env.FISHAUDIO_API_KEY},
voice: 'default',
text: text + ' two',
renderForCaching: true,
disableTtsCache: true
});
t.ok(!o2.servedFromCache, 'fishaudio synthed a second uncached render');
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('nineninesix speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.NINENINESIX_API_KEY) {
t.pass('skipping nineninesix speech synth tests since NINENINESIX_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'nineninesix',
credentials: {
api_key: process.env.NINENINESIX_API_KEY,
model_id: 'gepard-1.0'
},
language: 'en',
voice: '775e4dfd-819c-4325-92ec-250f487ef7e3',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed nineninesix audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('inworld speech synth', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -963,6 +1250,69 @@ test('inworld speech synth', async(t) => {
client.quit();
});
test('inworld streaming say: params', async(t) => {
/* This test asserts the streaming say: path, so it must run with streaming
enabled. The Google non-streaming test above sets
JAMBONES_DISABLE_TTS_STREAMING and, on its no-credentials skip path,
deletes the env var WITHOUT clearing the require cache — so lib/config can
still be holding 'true' by the time we get here. Re-require to be
independent of what ran before us.
*/
delete process.env.JAMBONES_DISABLE_TTS_STREAMING;
delete require.cache[require.resolve('../lib/config')];
delete require.cache[require.resolve('../lib/synth-audio')];
delete require.cache[require.resolve('..')];
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.INWORLD_API_KEY) {
t.pass('skipping inworld streaming say: param tests since INWORLD_API_KEY is not provided');
client.quit();
return t.end();
}
try {
let result = await synthAudio(stats, {
vendor: 'inworld',
credentials: {api_key: process.env.INWORLD_API_KEY, model_id: 'inworld-tts-1.5-mini'},
language: 'en',
voice: 'Ashley',
text: 'This is a test of inworld streaming.',
options: {temperature: 0.9, audioConfig: {pitch: 2.5, speakingRate: 1.2}},
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'inworld returns streaming say: path');
t.ok(result.filePath.includes('vendor=inworld'), 'streaming path contains vendor=inworld');
t.ok(result.filePath.includes('voice=Ashley'), 'streaming path contains voice');
t.ok(result.filePath.includes('model_id=inworld-tts-1.5-mini'), 'streaming path contains model_id');
t.ok(result.filePath.includes('temperature=0.9'), 'streaming path contains temperature');
/* pitch and speakingRate are nested under audioConfig; they used to be read
from the top level and emitted as "undefined"
*/
t.ok(result.filePath.includes('speakingRate=1.2'), 'audioConfig.speakingRate reaches the say: params');
t.ok(result.filePath.includes('pitch=2.5'), 'audioConfig.pitch reaches the say: params');
t.ok(!result.filePath.includes('undefined'), 'no undefined values in the say: params');
/* options omitted entirely: no stray keys */
result = await synthAudio(stats, {
vendor: 'inworld',
credentials: {api_key: process.env.INWORLD_API_KEY, model_id: 'inworld-tts-1.5-mini'},
language: 'en',
voice: 'Ashley',
text: 'This is a test of inworld streaming.',
disableTtsCache: true
});
t.ok(!result.filePath.includes('speakingRate='), 'speakingRate omitted when unset');
t.ok(!result.filePath.includes('pitch='), 'pitch omitted when unset');
t.ok(!result.filePath.includes('undefined'), 'no undefined values when options are omitted');
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('resemble speech synth', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);