Compare commits

...
10 Commits
Author SHA1 Message Date
Dave Horton 5dc26c916d 1.0.21 2026-09-30 23:04:20 -04:00
Dave Horton 12e3364461 speechify 2026-09-30 23:03:54 -04:00
Hoan Luu HuuandClaude Opus 5.5 6c67f6faa0 feat(tts): add speechify as a TTS vendor (#165)
- streaming returns a say: url for the mediajam starter
- cache render POSTs /v1/audio/stream for raw 24 kHz PCM (r24)
- credential options merge under the verb's options
- no model: simba-3.2 for English, simba-3.0 otherwise

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:01:18 -04:00
Dave Horton 18917f49ff 1.0.19 2026-09-29 14:22:33 -04:00
Dave HortonandClaude Opus 5.5 fa705694e1 feat(tts): add kugelaudio as a TTS vendor (#163)
Streaming returns a say: url for the mediajam dialect; the cache render POSTs
/v1/tts/generate for raw 8 kHz PCM (r8). Language is reduced to its primary
subtag and dropped when KugelAudio does not support it.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:21:41 -04:00
Dave HortonandClaude Opus 5 89fe0b87e9 fix: stop the RoleArn synth test printing AWS credentials into CI logs (#161)
With renderForCaching unset, synthPolly returns the mediajam streaming params
string, and lib/synth-audio.js:337-340 embeds accessKeyId, secretAccessKey and
sessionToken in it. The test interpolated that value into its assertion message,
so run 32995180996 printed live (1 hour) AWS credentials into a public build log.

Every other synth test in this file passes renderForCaching: true; this one was
the only exception. It now does the same, and the assertion no longer interpolates
opts.filePath at all, so the streaming params string cannot leak this way again.

Latent since the test was written -- it only surfaced once AWS_ROLE_ARN was wired
up and the test actually ran.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 13:44:05 -04:00
Dave HortonandClaude Opus 5 b928820906 ci: authenticate to AWS via GitHub OIDC instead of stored access keys (#160)
* ci: authenticate to AWS via GitHub OIDC instead of stored access keys

CI held a long-lived AWS access key as repository secrets. This replaces it with
a short-lived credential obtained through the GitHub OIDC provider, so no AWS key
is stored in the repo at all.

The tests could not simply inherit the OIDC credentials. They passed
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY explicitly, which routes
lib/get-aws-sts-token.js down its accessKeyId branch and calls GetSessionToken --
and AWS rejects GetSessionToken when it is called with session credentials.

test/aws-credentials.js centralises the decision. When AWS_SESSION_TOKEN is present
the credentials are temporary and only the region is passed, so the SDK's default
credential provider chain is used. Static keys still work unchanged, which keeps
local run-tests.sh working and also lets it run off an 'aws sso login' session with
no credentials in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: supply a cache key when AWS credentials come from the default chain

getAwsAuthToken derives its cache key as roleArn || accessKeyId || speech_credential_sid.
With temporary credentials the test passes none of those, so makeAwsKey received
undefined and hash.update() threw ERR_INVALID_ARG_TYPE.

Production never hits this: the instance-profile path always carries a
speech_credential_sid from the database, which is exactly what the comment above
that line describes. The test now supplies one the same way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: reference the CI role via secret rather than inlining the account id

speech-utils is public, so the role ARN (and with it the AWS account id) should not
be committed. It now comes from the AWS_ROLE_ARN secret, which also feeds the
'AWS speech synth tests by RoleArn' test -- previously always skipped, so the
AssumeRole credential path had no coverage here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: give the AWS RoleArn synth test its own cache key

The RoleArn test synthesized the same vendor/voice/language/text as the plain AWS
synth test that runs before it, so it hit that test's cache entry, servedFromCache
came back true and the !servedFromCache assertion failed.

Latent since the test was written -- it never ran, because AWS_ROLE_ARN was never
supplied. Wiring the secret in activated it and exposed the collision. Distinct text
makes the test independent of execution order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 13:38:32 -04:00
Dave Horton f4c3c7dc8b 1.0.18 2026-08-23 13:12:37 -04:00
Dave HortonandClaude Opus 5 a32f71ac5a feat: add fishaudio (Fish Audio) TTS support (#159)
Adds synthFishaudio with both arms: the say: streaming url consumed by the
mediajam dialect, and a POST /v1/tts cache render. The render asks for raw pcm
at 8k and returns extension r8 because fish's wav output carries a placeholder
RIFF size, the same problem gradium has.

Fish is a voice-cloning vendor, so the voice is a reference_id; the sentinel
'default' means send none and use fish's own default voice.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 13:11:59 -04:00
Dave HortonandClaude Opus 5 c8850998b5 fix(tts): use a current nineninesix voice id in the synth test (#158)
The vendor replaced its voice catalog and now rejects unknown ids, so
the env-gated test failed whenever NINENINESIX_API_KEY was set.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:19:59 -04:00
8 changed files with 434 additions and 41 deletions
+12 -3
View File
@@ -7,11 +7,18 @@ on:
jobs: jobs:
build: build:
runs-on: ubuntu-latest runs-on: ubuntu-latest
permissions:
id-token: write # required to request the GitHub OIDC token for AWS
contents: read
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
- uses: actions/setup-node@v4 - uses: actions/setup-node@v4
with: with:
node-version: '20' node-version: '20'
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
aws-region: us-east-1
- run: npm install - run: npm install
- run: npm run jslint - run: npm run jslint
- run: sudo apt update && sudo apt install -y squid - run: sudo apt update && sudo apt install -y squid
@@ -19,10 +26,12 @@ jobs:
- run: sudo systemctl start squid - run: sudo systemctl start squid
- run: npm test - run: npm test
env: env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }} # AWS_* are exported by configure-aws-credentials above as short-lived
AWS_REGION: ${{ secrets.AWS_REGION }} # OIDC credentials; no AWS keys are stored as repository secrets.
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }} GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }}
# Enables the "AWS speech synth tests by RoleArn" test, which exercises the
# AssumeRole credential path used in production.
AWS_ROLE_ARN: ${{ secrets.AWS_ROLE_ARN }}
MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }} MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }}
MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }} MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }}
+262 -2
View File
@@ -80,8 +80,8 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
logger = logger || noopLogger; logger = logger || noopLogger;
assert.ok(['google', 'aws', 'polly', 'microsoft', 'wellsaid', 'nvidia', 'elevenlabs', assert.ok(['google', 'aws', 'polly', 'microsoft', 'wellsaid', 'nvidia', 'elevenlabs',
'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'gradium', 'nineninesix', 'inworld', 'resemble', 'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'gradium', 'kugelaudio', 'nineninesix', 'inworld',
'murf', 'xai'] 'resemble', 'murf', 'xai', 'fishaudio', 'speechify']
.includes(vendor) || .includes(vendor) ||
vendor.startsWith('custom'), vendor.startsWith('custom'),
`synthAudio supported vendors are google, aws, microsoft, nvidia and wellsaid ..etc, not ${vendor}`); `synthAudio supported vendors are google, aws, microsoft, nvidia and wellsaid ..etc, not ${vendor}`);
@@ -139,6 +139,12 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
} else if ('gradium' === vendor) { } else if ('gradium' === vendor) {
assert.ok(voice, 'synthAudio requires voice when gradium is used'); assert.ok(voice, 'synthAudio requires voice when gradium is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when gradium is used'); assert.ok(credentials.api_key, 'synthAudio requires api_key when gradium is used');
} else if ('speechify' === vendor) {
assert.ok(voice, 'synthAudio requires voice when speechify is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when speechify is used');
} else if ('kugelaudio' === vendor) {
assert.ok(voice, 'synthAudio requires voice when kugelaudio is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when kugelaudio is used');
} else if ('nineninesix' === vendor) { } else if ('nineninesix' === vendor) {
assert.ok(voice, 'synthAudio requires voice when nineninesix is used'); assert.ok(voice, 'synthAudio requires voice when nineninesix is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when nineninesix is used'); assert.ok(credentials.api_key, 'synthAudio requires api_key when nineninesix is used');
@@ -149,6 +155,10 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
} else if (vendor === 'resemble') { } else if (vendor === 'resemble') {
assert.ok(voice, 'synthAudio requires voice when resemble is used'); assert.ok(voice, 'synthAudio requires voice when resemble is used');
assert.ok(credentials.api_key, 'synthAudio requires api_key when resemble is used'); assert.ok(credentials.api_key, 'synthAudio requires api_key when resemble is used');
} else if ('fishaudio' === vendor) {
/* no voice assert: fish synthesizes with its own default voice when
reference_id is omitted, which is what the 'default' selection means */
assert.ok(credentials.api_key, 'synthAudio requires api_key when fishaudio is used');
} }
const key = makeSynthKey({ const key = makeSynthKey({
@@ -220,11 +230,26 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming, credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache}); disableTtsCache});
break; break;
case 'fishaudio':
audioData = await synthFishaudio(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'gradium': case 'gradium':
audioData = await synthGradium(logger, { audioData = await synthGradium(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache}); disableTtsCache});
break; break;
case 'kugelaudio':
audioData = await synthKugelaudio(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'speechify':
audioData = await synthSpeechify(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache});
break;
case 'nineninesix': case 'nineninesix':
audioData = await synthNineninesix(logger, { audioData = await synthNineninesix(logger, {
credentials, stats, language, voice, key, text, renderForCaching, disableTtsStreaming, credentials, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
@@ -1503,8 +1528,243 @@ const synthGradium = async(logger, {
} }
}; };
/* kugelaudio — json websocket (/ws/tts/stream) for streaming, and POST /v1/tts/generate
for the cache render. the POST streams back bare little-endian 16-bit samples at the
requested sample_rate, which is exactly the r8 container at 8000.
voices are numeric ids (or public handles). language is an ISO 639-1 code that drives
text normalization; jambonz carries BCP-47, so only the primary subtag is sent. the
api rejects codes outside its list, so an unsupported or unset language is omitted
and the voice's own language applies. options.api_uri pins a region
(e.g. api.eu.kugelaudio.com).
*/
const KUGELAUDIO_LANGUAGES = ['ar', 'bg', 'bn', 'cs', 'da', 'de', 'el', 'en', 'es', 'fa', 'fi', 'fr', 'he', 'hi',
'hr', 'hu', 'id', 'it', 'ja', 'ko', 'ms', 'nl', 'no', 'pl', 'pt', 'ro', 'ru', 'sk', 'sl', 'sr', 'sv', 'ta', 'th',
'tr', 'uk', 'ur', 'vi', 'yue', 'zh'];
const synthKugelaudio = async(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache
}) => {
const {api_key, model_id} = credentials;
const {api_uri, speed, cfg_scale, temperature, normalize, project_id, dictionary_ids} = options || {};
const isSet = (v) => v !== null && v !== undefined;
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=kugelaudio';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
params += `,model_id=${model_id || 'kugel-3'}`;
if (language) params += `,language=${language}`;
if (api_uri) params += `,api_uri=${api_uri}`;
if (isSet(speed)) params += `,speed=${speed}`;
if (isSet(cfg_scale)) params += `,cfg_scale=${cfg_scale}`;
if (isSet(temperature)) params += `,temperature=${temperature}`;
if (isSet(normalize)) params += `,normalize=${normalize}`;
if (isSet(project_id)) params += `,project_id=${project_id}`;
/* the say: param parser is bracket-aware, so the json array survives intact */
if (Array.isArray(dictionary_ids)) params += `,dictionary_ids=${JSON.stringify(dictionary_ids)}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const host = (api_uri || 'api.kugelaudio.com').replace(/^[a-z]+:\/\//, '').replace(/\/$/, '');
const post = bent(`https://${host}`, 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json; charset=utf-8'
});
const voiceId = /^\d+$/.test(`${voice}`) ? Number(voice) : voice;
const lang = language && language.split('-')[0].toLowerCase();
const audioContent = await post('/v1/tts/generate', {
text,
voice_id: voiceId,
model_id: model_id || 'kugel-3',
sample_rate: sampleRate,
...(KUGELAUDIO_LANGUAGES.includes(lang) && {language: lang}),
...(isSet(speed) && {speed: Number(speed)}),
...(isSet(cfg_scale) && {cfg_scale: Number(cfg_scale)}),
...(isSet(temperature) && {temperature: Number(temperature)}),
...(isSet(normalize) && {normalize: normalize === true || normalize === 'true'}),
...(isSet(project_id) && {project_id: Number(project_id)}),
...(Array.isArray(dictionary_ids) && {dictionary_ids})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth kugelaudio returned error');
stats.increment('tts.count', ['vendor:kugelaudio', 'accepted:no']);
throw err;
}
};
/* simba-3.2 is English only; with no model set, other languages get simba-3.0 */
const speechifyModel = (model_id, language) => {
if (model_id) return model_id;
return !language || /^en/i.test(language) ? 'simba-3.2' : 'simba-3.0';
};
const synthSpeechify = async(logger, {
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
disableTtsCache
}) => {
const {api_key, model_id} = credentials;
let credOptions = credentials.options || {};
if (typeof credOptions === 'string') {
try {
credOptions = JSON.parse(credOptions);
} catch {
credOptions = {};
}
}
const {api_uri, loudness_normalization, text_normalization} = {...credOptions, ...options};
const isSet = (v) => v !== null && v !== undefined;
const model = speechifyModel(options?.model_id || model_id, language);
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=speechify';
params += `,voice=${voice}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
params += `,model_id=${model}`;
if (language) params += `,language=${language}`;
if (api_uri) params += `,api_uri=${api_uri}`;
if (isSet(loudness_normalization)) params += `,loudness_normalization=${loudness_normalization}`;
if (isSet(text_normalization)) params += `,text_normalization=${text_normalization}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
/* other pcm rates are mislabelled 24 kHz on workspaces pinned before 2026-09-30 */
const sampleRate = 24000;
const host = (api_uri || 'api.speechify.ai').replace(/^[a-z]+:\/\//, '').replace(/\/$/, '');
const post = bent(`https://${host}`, 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json',
'Speechify-Caller': 'jambonz'
});
const toBool = (v) => v === true || v === 'true';
const speechOptions = {
...(isSet(loudness_normalization) && {loudness_normalization: toBool(loudness_normalization)}),
...(isSet(text_normalization) && {text_normalization: toBool(text_normalization)})
};
const audioContent = await post('/v1/audio/stream', {
input: text,
voice_id: voice,
model,
output_format: `pcm_${sampleRate}`,
...(language && {language}),
...(Object.keys(speechOptions).length && {options: speechOptions})
});
return {
audioContent,
extension: 'r24',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth speechify returned error');
stats.increment('tts.count', ['vendor:speechify', 'accepted:no']);
throw err;
}
};
/* nineninesix.ai — a Cartesia-compatible API, but only raw/wav come back /* nineninesix.ai — a Cartesia-compatible API, but only raw/wav come back
(mp3 is rejected), so the cache render asks for wav rather than mp3. */ (mp3 is rejected), so the cache render asks for wav rather than mp3. */
/* fish.audio — msgpack websocket for streaming, and a POST endpoint for the cache
render. format:pcm + sample_rate:8000 returns bare little-endian 16-bit samples,
which is exactly the r8 container. we avoid fish's wav output because its RIFF
header carries a placeholder size (0xffffff24) — length is unknown up front, as
with gradium.
fish is a voice-cloning vendor: the "voice" is a reference_id returned by
POST /model, and omitting it entirely synthesizes with fish's default voice.
the sentinel value 'default' (the bundled fallback entry in the portal) means
exactly that — send no reference_id.
*/
const synthFishaudio = async(logger, {
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => {
const {api_key, model_id, fishaudio_tts_uri} = credentials;
const {reference_id, latency, chunk_length, speed, volume} = options || {};
/* free-text reference_id in the vendor options wins over the voice selector */
const refId = reference_id || (voice && voice !== 'default' ? voice : null);
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
let params = '{';
params += `api_key=${api_key}`;
params += `,playback_id=${key}`;
params += ',vendor=fishaudio';
params += `,voice=${refId || 'default'}`;
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
if (model_id) params += `,model_id=${model_id}`;
if (latency) params += `,latency=${latency}`;
if (chunk_length) params += `,chunk_length=${chunk_length}`;
if (speed) params += `,speed=${speed}`;
if (volume) params += `,volume=${volume}`;
if (fishaudio_tts_uri) params += `,endpoint=${fishaudio_tts_uri}`;
params += '}';
return {
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
servedFromCache: false,
rtt: 0
};
}
try {
const sampleRate = 8000;
const post = bent(fishaudio_tts_uri || 'https://api.fish.audio', 'POST', 'buffer', {
'Authorization': `Bearer ${api_key}`,
'Content-Type': 'application/json',
/* the model is selected by header, not in the body */
'model': model_id || 's2.1-pro'
});
const audioContent = await post('/v1/tts', {
text,
format: 'pcm',
sample_rate: sampleRate,
...(refId && {reference_id: refId}),
...(latency && {latency}),
...(chunk_length && {chunk_length: parseInt(chunk_length, 10)}),
...((speed || volume) && {prosody: {
...(speed && {speed: parseFloat(speed)}),
...(volume && {volume: parseFloat(volume)})
}})
});
return {
audioContent,
extension: 'r8',
sampleRate
};
} catch (err) {
logger.info({err}, 'synth fishaudio returned error');
stats.increment('tts.count', ['vendor:fishaudio', 'accepted:no']);
throw err;
}
};
const synthNineninesix = async(logger, { const synthNineninesix = async(logger, {
credentials, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache credentials, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
}) => { }) => {
+2 -2
View File
@@ -1,12 +1,12 @@
{ {
"name": "@jambonz/speech-utils", "name": "@jambonz/speech-utils",
"version": "1.0.17", "version": "1.0.21",
"lockfileVersion": 2, "lockfileVersion": 2,
"requires": true, "requires": true,
"packages": { "packages": {
"": { "": {
"name": "@jambonz/speech-utils", "name": "@jambonz/speech-utils",
"version": "1.0.17", "version": "1.0.21",
"license": "MIT", "license": "MIT",
"dependencies": { "dependencies": {
"@aws-sdk/client-polly": "^3.496.0", "@aws-sdk/client-polly": "^3.496.0",
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "@jambonz/speech-utils", "name": "@jambonz/speech-utils",
"version": "1.0.17", "version": "1.0.21",
"description": "TTS-related speech utilities for jambonz", "description": "TTS-related speech utilities for jambonz",
"main": "index.js", "main": "index.js",
"author": "Dave Horton", "author": "Dave Horton",
+23
View File
@@ -0,0 +1,23 @@
/**
* Resolve AWS credentials for the test suite.
*
* Returns null when AWS is not configured, so the caller skips.
*
* When AWS_SESSION_TOKEN is set the credentials are temporary -- GitHub OIDC in CI, or
* `aws sso login` locally -- and the key pair must not be passed through. Doing so sends
* lib/get-aws-sts-token.js down its accessKeyId branch, which calls GetSessionToken, and
* AWS rejects GetSessionToken when it is called with session credentials. Returning the
* region alone routes to the SDK's default credential provider chain instead, which
* handles temporary credentials correctly.
*/
module.exports = () => {
const region = process.env.AWS_REGION;
if (!region) return null;
const accessKeyId = process.env.AWS_ACCESS_KEY_ID;
const secretAccessKey = process.env.AWS_SECRET_ACCESS_KEY;
if (process.env.AWS_SESSION_TOKEN) return {region};
if (accessKeyId && secretAccessKey) return {accessKeyId, secretAccessKey, region};
return {region};
};
+11 -11
View File
@@ -6,33 +6,33 @@ process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason); console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
}); });
const awsCredentials = require('./aws-credentials');
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
test('AWS - create and cache auth token', async(t) => { test('AWS - create and cache auth token', async(t) => {
const fn = require('..'); const fn = require('..');
const {client, getAwsAuthToken} = fn(opts, logger); const {client, getAwsAuthToken} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) { const credentials = awsCredentials();
if (!credentials) {
t.pass('skipping AWS auth token tests since no AWS credentials provided'); t.pass('skipping AWS auth token tests since no AWS credentials provided');
t.end(); t.end();
client.quit(); client.quit();
return; return;
} }
// getAwsAuthToken derives its cache key from roleArn || accessKeyId || speech_credential_sid.
// With temporary credentials none of the first two are passed, so supply a stable sid --
// which is what production does for instance-profile credentials.
const args = {...credentials, speech_credential_sid: 'test-aws-speech-credential'};
try { try {
let obj = await getAwsAuthToken({ let obj = await getAwsAuthToken(args);
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
//console.log({obj}, 'received auth token from AWS'); //console.log({obj}, 'received auth token from AWS');
t.ok(obj.securityToken && !obj.servedFromCache, 'successfullY generated auth token from AWS'); t.ok(obj.securityToken && !obj.servedFromCache, 'successfullY generated auth token from AWS');
await sleep(250); await sleep(250);
obj = await getAwsAuthToken({ obj = await getAwsAuthToken(args);
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
//console.log({obj}, 'received auth token from AWS - second request'); //console.log({obj}, 'received auth token from AWS - second request');
t.ok(obj.securityToken && obj.servedFromCache, 'successfully received access token from cache'); t.ok(obj.securityToken && obj.servedFromCache, 'successfully received access token from cache');
+5 -7
View File
@@ -1,4 +1,5 @@
const test = require('tape').test ; const test = require('tape').test ;
const awsCredentials = require('./aws-credentials');
const config = require('config'); const config = require('config');
const opts = config.get('redis'); const opts = config.get('redis');
const fs = require('fs'); const fs = require('fs');
@@ -45,18 +46,15 @@ test('AWS tests', async(t) => {
const fn = require('..'); const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger); const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) { const credentials = awsCredentials();
t.pass('skipping AWS speech synth tests since AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, or AWS_REGION not provided'); if (!credentials) {
t.pass('skipping AWS speech synth tests since AWS_REGION not provided');
return t.end(); return t.end();
} }
try { try {
const opts = { const opts = {
vendor: 'aws', vendor: 'aws',
credentials: { credentials
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
}
}; };
let result = await getTtsVoices(opts); let result = await getTtsVoices(opts);
t.ok(result?.Voices?.length > 0, `GetVoices: successfully retrieved ${result.Voices.length} voices from AWS`); t.ok(result?.Voices?.length > 0, `GetVoices: successfully retrieved ${result.Voices.length} voices from AWS`);
+118 -15
View File
@@ -1,4 +1,5 @@
const test = require('tape').test; const test = require('tape').test;
const awsCredentials = require('./aws-credentials');
const config = require('config'); const config = require('config');
const opts = config.get('redis'); const opts = config.get('redis');
const fs = require('fs'); const fs = require('fs');
@@ -668,18 +669,15 @@ test('AWS speech synth tests', async(t) => {
const fn = require('..'); const fn = require('..');
const {synthAudio, client} = fn(opts, logger); const {synthAudio, client} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) { const credentials = awsCredentials();
t.pass('skipping AWS speech synth tests since AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, or AWS_REGION not provided'); if (!credentials) {
t.pass('skipping AWS speech synth tests since AWS_REGION not provided');
return t.end(); return t.end();
} }
try { try {
let opts = await synthAudio(stats, { let opts = await synthAudio(stats, {
vendor: 'aws', vendor: 'aws',
credentials: { credentials,
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
},
language: 'en-US', language: 'en-US',
voice: 'Joey', voice: 'Joey',
text: 'This is a test. This is only a test', text: 'This is a test. This is only a test',
@@ -689,11 +687,7 @@ test('AWS speech synth tests', async(t) => {
opts = await synthAudio(stats, { opts = await synthAudio(stats, {
vendor: 'aws', vendor: 'aws',
credentials: { credentials,
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
},
language: 'en-US', language: 'en-US',
voice: 'Joey', voice: 'Joey',
text: 'This is a test. This is only a test', text: 'This is a test. This is only a test',
@@ -724,9 +718,18 @@ test('AWS speech synth tests by RoleArn', async(t) => {
}, },
language: 'en-US', language: 'en-US',
voice: 'Joey', voice: 'Joey',
text: 'This is a test. This is only a test', // Distinct text on purpose: the 'AWS speech synth tests' above cache audio for
// the same vendor/voice/language, and identical text would hit that cache entry,
// making servedFromCache true and this assertion fail.
text: 'This is a roleArn test. This is only a roleArn test',
// Without this, synthPolly returns the mediajam streaming params string, which
// embeds accessKeyId/secretAccessKey/sessionToken (see lib/synth-audio.js).
// Every other synth test in this file renders for caching; this one did not,
// so it printed live credentials into a public CI log.
renderForCaching: true,
}); });
t.ok(!opts.servedFromCache, `successfully synthesized aws by roleArn audio to ${opts.filePath}`); // Deliberately does not interpolate opts.filePath -- it can carry credentials.
t.ok(!opts.servedFromCache, 'successfully synthesized aws by roleArn audio');
} catch (err) { } catch (err) {
console.error(err); console.error(err);
t.end(err); t.end(err);
@@ -1087,6 +1090,106 @@ test('gradium speech synth tests', async(t) => {
client.quit(); client.quit();
}); });
test('kugelaudio speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.KUGELAUDIO_API_KEY) {
t.pass('skipping kugelaudio speech synth tests since KUGELAUDIO_API_KEY is not provided');
return t.end();
}
const text = 'Guten Tag und willkommen bei jambonz! Ihre Bestellung kostet 12,99 Euro. ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'kugelaudio',
credentials: {
api_key: process.env.KUGELAUDIO_API_KEY,
model_id: 'kugel-3'
},
language: 'de-DE',
voice: '1930',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed kugelaudio audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('speechify speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.SPEECHIFY_API_KEY) {
t.pass('skipping speechify speech synth tests since SPEECHIFY_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambonz! ' + Date.now();
try {
const opts = await synthAudio(stats, {
vendor: 'speechify',
credentials: {
api_key: process.env.SPEECHIFY_API_KEY
},
language: 'en-US',
voice: 'geffen_32',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthed speechify audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('fishaudio speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.FISHAUDIO_API_KEY) {
t.pass('skipping fishaudio speech synth tests since FISHAUDIO_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
/* voice 'default' means "send no reference_id" — fish's own default voice */
const o = await synthAudio(stats, {
vendor: 'fishaudio',
credentials: {
api_key: process.env.FISHAUDIO_API_KEY,
model_id: 's2.1-pro'
},
voice: 'default',
text,
renderForCaching: true
});
t.ok(!o.servedFromCache, `successfully synthed fishaudio audio to ${o.filePath}`);
/* the cache render must be raw 8k pcm (r8): fish's wav header carries a
placeholder RIFF size, so we never ask for wav */
const o2 = await synthAudio(stats, {
vendor: 'fishaudio',
credentials: {api_key: process.env.FISHAUDIO_API_KEY},
voice: 'default',
text: text + ' two',
renderForCaching: true,
disableTtsCache: true
});
t.ok(!o2.servedFromCache, 'fishaudio synthed a second uncached render');
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('nineninesix speech synth tests', async(t) => { test('nineninesix speech synth tests', async(t) => {
const fn = require('..'); const fn = require('..');
const {synthAudio, client} = fn(opts, logger); const {synthAudio, client} = fn(opts, logger);
@@ -1104,7 +1207,7 @@ test('nineninesix speech synth tests', async(t) => {
model_id: 'gepard-1.0' model_id: 'gepard-1.0'
}, },
language: 'en', language: 'en',
voice: '3ad7a827-7fd1-4954-bf35-47d4cc33d9ed', voice: '775e4dfd-819c-4325-92ec-250f487ef7e3',
text, text,
renderForCaching: true renderForCaching: true
}); });