Compare commits

..
92 Commits
Author SHA1 Message Date
Dave Horton 5dc26c916d 1.0.21 2026-09-30 23:04:20 -04:00
Dave Horton 12e3364461 speechify 2026-09-30 23:03:54 -04:00
Hoan Luu HuuandClaude Opus 5.5 6c67f6faa0 feat(tts): add speechify as a TTS vendor (#165)
- streaming returns a say: url for the mediajam starter
- cache render POSTs /v1/audio/stream for raw 24 kHz PCM (r24)
- credential options merge under the verb's options
- no model: simba-3.2 for English, simba-3.0 otherwise

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:01:18 -04:00
Dave Horton 18917f49ff 1.0.19 2026-09-29 14:22:33 -04:00
Dave HortonandClaude Opus 5.5 fa705694e1 feat(tts): add kugelaudio as a TTS vendor (#163)
Streaming returns a say: url for the mediajam dialect; the cache render POSTs
/v1/tts/generate for raw 8 kHz PCM (r8). Language is reduced to its primary
subtag and dropped when KugelAudio does not support it.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:21:41 -04:00
Dave HortonandClaude Opus 5 89fe0b87e9 fix: stop the RoleArn synth test printing AWS credentials into CI logs (#161)
With renderForCaching unset, synthPolly returns the mediajam streaming params
string, and lib/synth-audio.js:337-340 embeds accessKeyId, secretAccessKey and
sessionToken in it. The test interpolated that value into its assertion message,
so run 32995180996 printed live (1 hour) AWS credentials into a public build log.

Every other synth test in this file passes renderForCaching: true; this one was
the only exception. It now does the same, and the assertion no longer interpolates
opts.filePath at all, so the streaming params string cannot leak this way again.

Latent since the test was written -- it only surfaced once AWS_ROLE_ARN was wired
up and the test actually ran.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 13:44:05 -04:00
Dave HortonandClaude Opus 5 b928820906 ci: authenticate to AWS via GitHub OIDC instead of stored access keys (#160)
* ci: authenticate to AWS via GitHub OIDC instead of stored access keys

CI held a long-lived AWS access key as repository secrets. This replaces it with
a short-lived credential obtained through the GitHub OIDC provider, so no AWS key
is stored in the repo at all.

The tests could not simply inherit the OIDC credentials. They passed
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY explicitly, which routes
lib/get-aws-sts-token.js down its accessKeyId branch and calls GetSessionToken --
and AWS rejects GetSessionToken when it is called with session credentials.

test/aws-credentials.js centralises the decision. When AWS_SESSION_TOKEN is present
the credentials are temporary and only the region is passed, so the SDK's default
credential provider chain is used. Static keys still work unchanged, which keeps
local run-tests.sh working and also lets it run off an 'aws sso login' session with
no credentials in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: supply a cache key when AWS credentials come from the default chain

getAwsAuthToken derives its cache key as roleArn || accessKeyId || speech_credential_sid.
With temporary credentials the test passes none of those, so makeAwsKey received
undefined and hash.update() threw ERR_INVALID_ARG_TYPE.

Production never hits this: the instance-profile path always carries a
speech_credential_sid from the database, which is exactly what the comment above
that line describes. The test now supplies one the same way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: reference the CI role via secret rather than inlining the account id

speech-utils is public, so the role ARN (and with it the AWS account id) should not
be committed. It now comes from the AWS_ROLE_ARN secret, which also feeds the
'AWS speech synth tests by RoleArn' test -- previously always skipped, so the
AssumeRole credential path had no coverage here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: give the AWS RoleArn synth test its own cache key

The RoleArn test synthesized the same vendor/voice/language/text as the plain AWS
synth test that runs before it, so it hit that test's cache entry, servedFromCache
came back true and the !servedFromCache assertion failed.

Latent since the test was written -- it never ran, because AWS_ROLE_ARN was never
supplied. Wiring the secret in activated it and exposed the collision. Distinct text
makes the test independent of execution order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 13:38:32 -04:00
Dave Horton f4c3c7dc8b 1.0.18 2026-08-23 13:12:37 -04:00
Dave HortonandClaude Opus 5 a32f71ac5a feat: add fishaudio (Fish Audio) TTS support (#159)
Adds synthFishaudio with both arms: the say: streaming url consumed by the
mediajam dialect, and a POST /v1/tts cache render. The render asks for raw pcm
at 8k and returns extension r8 because fish's wav output carries a placeholder
RIFF size, the same problem gradium has.

Fish is a voice-cloning vendor, so the voice is a reference_id; the sentinel
'default' means send none and use fish's own default voice.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 13:11:59 -04:00
Dave HortonandClaude Opus 5 c8850998b5 fix(tts): use a current nineninesix voice id in the synth test (#158)
The vendor replaced its voice catalog and now rejects unknown ids, so
the env-gated test failed whenever NINENINESIX_API_KEY was set.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:19:59 -04:00
Dave Horton e8b2009d29 1.0.17 2026-08-10 16:45:32 -04:00
Dave HortonandClaude Opus 5 27e07bb551 fix(tts): forward gradium json_config to the streaming path (#157)
synthGradium destructured json_config from options but only forwarded
pronunciation_id into the say: param block, so the setting could never
reach mediajam's streaming dialect — only the cache-render POST. The
say: parser is brace-aware, so a nested json object survives intact.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 16:45:07 -04:00
Dave Horton 33456b93d9 1.0.16 2026-08-08 23:07:43 -04:00
Dave HortonandClaude Opus 5 70e91b5057 test(inworld): gate the say: param test on INWORLD_API_KEY (#156)
The test I added in #155 ran unconditionally, which broke `npm test` without
credentials — the path husky's pre-commit hook takes, so `npm version patch`
could not commit.

Two causes, both addressed:

- no credential gate, unlike every other vendor test in this file. Now skips
  without INWORLD_API_KEY, and closes its redis client on that path so the
  run can still exit.
- it assumed streaming was enabled. The Google non-streaming test sets
  JAMBONES_DISABLE_TTS_STREAMING and, on its no-credentials skip path,
  deletes the env var WITHOUT clearing the require cache (unlike its finally
  block, which clears both) — so lib/config still held 'true' further down
  the file and synthInworld took the non-streaming branch, attempting a real
  vendor call. The test now re-requires with streaming enabled so it does not
  depend on what ran before it.

Verified both ways: skips and exits 0 with no key; 11/11 with a key even
under the leaked state. Re-introducing the #155 bug still fails 3 assertions,
so the regression value is intact.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:06:52 -04:00
Dave HortonandClaude Opus 5 c066840f85 fix(inworld): read pitch and speakingRate from audioConfig in the say: params (#155)
The streaming say: path guarded on opts.audioConfig?.pitch and
opts.audioConfig?.speakingRate but interpolated opts.pitch and
opts.speakingRate, which are undefined — so anyone setting them under
audioConfig (what the docs and the portal defaults tell you to do) got
'pitch=undefined,speakingRate=undefined' on the wire and their setting
silently dropped.

Adds a test for the say: params that needs no credentials, since the
streaming branch builds the path without calling the vendor.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:49:50 -04:00
Dave Horton d89457f67d 1.0.15 2026-08-06 10:23:47 -04:00
Dave HortonandClaude Opus 5 3c15976669 fix(deps): remove junk "23" dependency (#154)
"23": "^0.0.0" is not a real dependency - it is an empty placeholder
package (0.0.0, no deps, unrelated third-party maintainer) that landed
here from a stray npm install. Nothing in the package references it.

Beyond the noise, it is a small supply-chain liability: a dependency on
a squatted single-number name owned by nobody we know, shipped to every
consumer of speech-utils.

Lint passes; full test suite passes 105/105 with live vendor credentials.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:23:31 -04:00
Dave Horton 70abb04941 1.0.14 2026-08-06 10:13:51 -04:00
Dave HortonandClaude Opus 5 724af7fcab fix(deps): drop unused undici dependency (#153)
undici was declared as a direct dependency but is never required
anywhere in this package - grep across index.js, lib/ and stubs/ finds
no reference to undici, ProxyAgent, setGlobalDispatcher or Dispatcher.
The one HTTP call in lib/synth-audio.js uses the global fetch, and
Azure proxy support goes through the SDK's own setProxy plus the
http_proxy_ip/http_proxy_port params.

Removing it clears all seven open undici advisories from npm audit for
this package and its consumers, and stops speech-utils pulling a
duplicate undici 7.x into trees where the consumer already depends on
undici 8 (feature-server does).

Lint passes; full test suite passes 105/105 with live vendor credentials.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:13:31 -04:00
Dave Horton c0a79090c2 1.0.13 2026-08-06 10:04:26 -04:00
Dave HortonandClaude Opus 5 7a9a3e2468 fix(deps): bump microsoft-cognitiveservices-speech-sdk to ^1.51.0 (#152)
The Azure speech SDK was pinned to exactly 1.38.0 since the initial
commit. It pulls in uuid <11.1.1, which carries GHSA-w5hq-g745-h8pq
(missing buffer bounds check in v3/v5/v6 when buf is provided). The
advisory covers Azure SDK 1.14.0-1.50.0; 1.51.0 clears it.

Loosened the exact pin to a caret range so future patches come in
without another PR. The SDK surface used here (SpeechConfig,
SpeechSynthesizer, ResultReason, CancellationDetails,
SpeechSynthesisOutputFormat) is unchanged in 1.51.0.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 10:03:59 -04:00
Dave Horton ab2b3f7fb2 1.0.12 2026-08-04 09:04:50 -04:00
Dave HortonandClaude Opus 5 e848d0276e feat(tts): add gradium as a TTS vendor (#151)
Streaming arm returns a say: url for the mediajam dialect; the cache-render
arm posts to /api/post/speech/tts with only_audio and pcm_8000, which is bare
r8 samples and avoids gradium's streaming wav header (0xffffffff RIFF size).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 09:02:07 -04:00
Dave Horton d2447d477d 1.0.11 2026-08-03 12:04:47 -04:00
Dave HortonandClaude Opus 5 ec8bbacc2c feat(tts): add nineninesix.ai synthesis (#150)
Streaming goes through mediajam's say: url; the cache render posts to
/tts/bytes for wav, since the service rejects mp3.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 12:03:52 -04:00
Dave Horton 9694714a7f 1.0.10 2026-07-28 20:59:05 -04:00
Hoan Luu Huu 6a6de9b7a1 support google tts configuraiton (#149) 2026-07-28 20:58:26 -04:00
Dave Horton 04d1fb548b 1.0.9 2026-07-21 06:52:05 -04:00
Hoan Luu Huu bbf0167b40 support deepgramflux tts (#148) 2026-07-21 06:49:38 -04:00
Dave Horton 4cfa286242 1.0.8 2026-07-05 17:42:09 -04:00
Hoan Luu Huu 42ac63adc6 support xaiTTS (#147) 2026-07-05 17:41:36 -04:00
Dave Horton 12a121672d 1.0.7 2026-07-03 07:12:20 -04:00
Hoan Luu HuuandClaude Opus 4.8 7b94a5a969 feat(murf): add Murf.ai TTS support to synthAudio (#146)
* feat(murf): add Murf.ai TTS support to synthAudio

Add synthMurf() following the rimelabs/cartesia pattern:
- streaming path returns a say:{vendor=murf,...} filePath consumed by the
  FreeSWITCH mod_murf_tts module
- non-streaming path calls POST /v1/speech/stream (api-key header) and returns
  WAV audio for cache rendering

Register murf in the supported-vendor assert list and the synth switch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(murf): drop Accept: audio/basic header (caused 406 Not Acceptable)

Murf's /v1/speech/stream rejects an unmatched Accept header with 406; the
response container is chosen by the `format` body field instead. Verified a
WAV request now returns 200 (valid RIFF/WAVE).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 07:11:39 -04:00
Dave Horton c47b4883c7 1.0.6 2026-06-17 16:20:30 -04:00
Dave HortonandClaude Opus 4.8 7d076bb8b4 chore: deprecate + remove verbio, nuance, playht speech vendor support (#144)
* chore: deprecate and remove verbio, nuance speech vendor support

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: also deprecate and remove PlayHT speech vendor

PlayHT was acquired and no longer provides the service.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:20:00 -04:00
Dave HortonandClaude Opus 4.8 142323a151 fix(synth-audio): include the Polly VoiceId in the aws say: url
synthPolly received voice but never emitted it into the say:{...} url
(unlike synthMicrosoft, which includes voice=). Polly's SynthesizeSpeech
requires a VoiceId, so the media-server-native (say:-url) Polly path had no
way to pick a voice. Add voice= to the params.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 22:31:57 -04:00
Dave Horton d176a644fe 1.0.4 2026-06-06 10:44:11 +02:00
Hoan Luu Huu 644f2918dc support cartesia sonic3.5 (#143) 2026-06-06 10:42:40 +02:00
Dave HortonandClaude Opus 4.5 4f430b9785 update publish workflow to use actions v4
Fixes npm warning about deprecated always-auth config by updating
setup-node from v3 to v4.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-06-03 08:51:56 -04:00
Dave Horton 18c658d20c 1.0.3 2026-06-03 07:27:44 -04:00
Dave Horton 062608cf13 update lock file 2026-06-03 07:27:28 -04:00
Hoan Luu Huu a2ea94c14a support rimelabs coda (#142) 2026-06-03 07:24:13 -04:00
Dave Horton 04bb85ef81 1.0.1 2026-03-25 10:08:57 -04:00
Dave Horton c123f19898 remove ibm speech since it is not used (to my knowledge) and has dependencies with vulnerabilities (#141) 2026-03-25 10:08:30 -04:00
Dave Horton 305695d068 0.2.30 2026-01-22 08:03:44 -05:00
Dave Horton 1477752d40 update dep 2026-01-22 08:03:24 -05:00
Dave Horton fbedfe947f vuln 2026-01-22 07:58:06 -05:00
Dave Horton ab88facd52 Merge pull request #139 from jambonz/feat/google_gemini_tts
google tts support api_mode
2026-01-22 07:52:04 -05:00
Hoan HL a2be64da89 google tts support api_mode 2026-01-22 16:36:57 +07:00
Dave Horton c5e0f256e6 0.2.28 2026-01-17 21:39:59 -05:00
Dave Horton d0378751ba Merge pull request #135 from jambonz/feat/gemini_tts
support gemini tts
2026-01-17 21:39:28 -05:00
Hoan HL a7e391fcb2 wip 2026-01-18 08:45:14 +07:00
Hoan HL 10095de25d wip 2026-01-17 16:16:11 +07:00
Hoan HL 89007ba7cc wip 2026-01-17 15:13:47 +07:00
Hoan HL 29edccd5bf wip 2026-01-17 15:10:17 +07:00
Hoan HL ca9538030f wip 2026-01-14 17:37:38 +07:00
Hoan HL 417d58080e wip 2026-01-14 17:34:07 +07:00
Hoan HL 62fec6f5e4 wip 2026-01-14 16:00:43 +07:00
Hoan HL 490d23a703 wip 2026-01-14 15:46:29 +07:00
Hoan HL b91e6cc145 wip 2026-01-14 13:58:38 +07:00
Hoan HL 0e33358254 wip 2026-01-14 13:51:32 +07:00
Hoan HL 6b2b35acfb wip 2026-01-12 18:26:47 +07:00
Hoan HL ded60cb7aa wip 2026-01-12 17:18:11 +07:00
Hoan HL 460ca70ea7 wip 2026-01-12 12:55:43 +07:00
Hoan HL 0fbbbb8053 add testcases 2026-01-12 09:21:41 +07:00
Hoan HL 0ea7082da2 support gemini tts 2026-01-11 07:30:18 +07:00
Dave Horton 5f7e7458bb 0.2.27 2025-11-17 07:26:45 -05:00
Dave Horton f6714fb9e1 Merge pull request #134 from jambonz/fix/1628
fixed cartesia collect audio from stream
2025-11-13 07:15:51 -05:00
Hoan HL c04ef29f7c fixed cartesia collect audio from stream 2025-11-13 14:10:00 +07:00
Dave Horton 8154944252 0.2.26 2025-10-30 07:08:47 -04:00
Dave Horton 16fe8dce01 0.2.25 2025-10-30 07:08:10 -04:00
Dave Horton 11a955500d Merge pull request #132 from jambonz/feat/sonic_3
cartesia support volume for sonic3
2025-10-30 07:07:15 -04:00
Hoan HL c122129b55 wip 2025-10-30 12:29:09 +07:00
Hoan HL 41b26b966b wip 2025-10-30 12:08:38 +07:00
Hoan HL aa09c15b20 wip 2025-10-30 06:46:57 +07:00
Hoan HL 32d5f12638 cartesia support volume for sonic3 2025-10-30 05:57:38 +07:00
Dave Horton 4e336822a0 Merge pull request #131 from jambonz/feat/gh_fs_1384
support elevenlabs api_uri
2025-10-08 13:45:02 -04:00
Hoan HL b898d794b0 wip 2025-10-08 15:05:04 +07:00
Hoan HL fb754ca101 support elevenlabs api_uri 2025-10-08 10:26:17 +07:00
Dave Horton 8d8195be9a 0.2.24 2025-10-03 08:54:14 -04:00
Dave Horton eb4e1a773f Merge pull request #130 from jambonz/feat/disableTtsCache
set write_cache_file = 0 when disableTtsCache
2025-10-03 02:21:57 -04:00
Hoan HL 6fecb8755d set write_cache_file = 0 when disableTtsCache 2025-10-03 11:09:30 +07:00
Dave Horton fdb56cbc77 Merge pull request #128 from jambonz/feat/custom_tts_stream
support custom tts Stream
2025-09-11 09:24:18 -04:00
Quan HL 5328a60de8 support custom tts Stream 2025-09-11 09:18:42 -04:00
Dave Horton ea1ab301d5 0.2.23 2025-09-10 22:53:45 -04:00
Dave Horton 0a4994c5a3 Merge pull request #129 from jambonz/fix/fd_1372
remove optimize_streaming_latency as default option for elevenlabs
2025-09-10 22:53:27 -04:00
Quan HL cb71460189 remove optimize_streaming_latency as default option for elevenlabs 2025-09-11 09:44:53 +07:00
Dave Horton f08efa49b1 0.2.22 2025-08-20 18:03:08 -04:00
Dave Horton 80eda37ef6 Merge pull request #126 from jambonz/feat/revamp-playback-id
use synth key as playback id
2025-08-20 18:02:35 -04:00
Dave Horton 7768e58b19 use synth key as playback id 2025-08-20 17:11:45 -04:00
Dave Horton 6efc99ad83 0.2.21 2025-08-20 10:59:40 -04:00
Dave Horton 127d01a8c1 add playback_id setting to additional tts vendors 2025-08-20 10:58:28 -04:00
25 changed files with 4797 additions and 11435 deletions
+14 -12
View File
@@ -7,30 +7,32 @@ on:
jobs: jobs:
build: build:
runs-on: ubuntu-latest runs-on: ubuntu-latest
permissions:
id-token: write # required to request the GitHub OIDC token for AWS
contents: read
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
- uses: actions/setup-node@v4 - uses: actions/setup-node@v4
with: with:
node-version: lts/* node-version: '20'
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
aws-region: us-east-1
- run: npm install - run: npm install
- name: Install Docker Compose
run: |
sudo curl -L "https://github.com/docker/compose/releases/download/1.29.2/docker-compose-$(uname -s)-$(uname -m)" -o /usr/local/bin/docker-compose
sudo chmod +x /usr/local/bin/docker-compose
docker-compose --version
- run: npm run jslint - run: npm run jslint
- run: sudo apt update && sudo apt install -y squid - run: sudo apt update && sudo apt install -y squid
- run: sudo cp test/squid.conf /etc/squid/squid.conf - run: sudo cp test/squid.conf /etc/squid/squid.conf
- run: sudo systemctl start squid - run: sudo systemctl start squid
- run: npm test - run: npm test
env: env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }} # AWS_* are exported by configure-aws-credentials above as short-lived
AWS_REGION: ${{ secrets.AWS_REGION }} # OIDC credentials; no AWS keys are stored as repository secrets.
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }} GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }}
IBM_API_KEY: ${{ secrets.IBM_API_KEY }} # Enables the "AWS speech synth tests by RoleArn" test, which exercises the
IBM_TTS_API_KEY: ${{ secrets.IBM_TTS_API_KEY }} # AssumeRole credential path used in production.
IBM_TTS_REGION: ${{ secrets.IBM_TTS_REGION }} AWS_ROLE_ARN: ${{ secrets.AWS_ROLE_ARN }}
MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }} MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }}
MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }} MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
+2 -2
View File
@@ -12,8 +12,8 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v3 - uses: actions/checkout@v4
- uses: actions/setup-node@v3 - uses: actions/setup-node@v4
with: with:
node-version: lts/* node-version: lts/*
registry-url: 'https://registry.npmjs.org' registry-url: 'https://registry.npmjs.org'
+84
View File
@@ -0,0 +1,84 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Overview
`@jambonz/speech-utils` is a Node.js library providing TTS (Text-to-Speech) utilities for the jambonz CPaaS platform. It handles speech synthesis with caching through Redis and supports multiple TTS vendors.
## Commands
```bash
# Run tests (requires Docker for Redis)
npm test
# Run linter
npm run jslint
# Auto-fix lint issues
npm run jslint:fix
# Generate coverage report
npm run coverage
```
## Testing
Tests use `tape` and require Redis. The test harness automatically starts/stops Redis via Docker Compose (`test/docker-compose-testbed.yaml`).
Most tests are conditional based on environment variables for vendor credentials:
- `GCP_FILE` or `GCP_JSON_KEY` - Google TTS
- `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_REGION` - AWS Polly
- `MICROSOFT_API_KEY`, `MICROSOFT_REGION` - Azure TTS
- `ELEVENLABS_API_KEY`, `ELEVENLABS_VOICE_ID`, `ELEVENLABS_MODEL_ID` - ElevenLabs
- `OPENAI_API_KEY` - OpenAI Whisper TTS
- And others per vendor
Redis config is in `config/test.json` (port 3379).
## Architecture
### Entry Point
`index.js` exports a factory function that takes Redis options and a logger, returning an object with these methods:
- `synthAudio` - Main synthesis function
- `getTtsVoices` - List available voices for a vendor
- `purgeTtsCache` / `getTtsSize` / `addFileToCache` - Cache management
- `getAwsAuthToken` - Token management
### Core Module: `lib/synth-audio.js`
The `synthAudio` function handles synthesis for all vendors. Key behaviors:
1. **Cache check**: Generates SHA1 hash key from (vendor, language, voice, engine, model, text, instructions)
2. **Streaming vs non-streaming**: When `JAMBONES_DISABLE_TTS_STREAMING` is not set and `renderForCaching=false`, returns `say:{params}text` format for FreeSWITCH streaming playback instead of generating files
3. **Vendor dispatch**: Switch statement routes to vendor-specific synth functions (`synthGoogle`, `synthPolly`, `synthMicrosoft`, etc.)
4. **Caching**: Stores audio as base64 JSON in Redis with configurable TTL (default 4 hours)
### Supported Vendors
google, aws/polly, microsoft/azure, nvidia (Riva), wellsaid, elevenlabs, cartesia, inworld, rimelabs, whisper (OpenAI), deepgram, resemble, custom:*
### gRPC Stubs
`stubs/riva/` contains generated protobuf/gRPC code for NVIDIA Riva.
## Environment Variables
Key configuration via env vars (see `lib/config.js`):
- `JAMBONES_DISABLE_TTS_STREAMING` - Force non-streaming mode
- `JAMBONES_DISABLE_AZURE_TTS_STREAMING` - Azure-specific streaming disable
- `JAMBONES_TTS_CACHE_DURATION_MINS` - Cache TTL in minutes (default: 240)
- `JAMBONES_TTS_TRIM_SILENCE` - Trim trailing silence from audio
- `JAMBONES_TMP_FOLDER` - Temp folder for audio files (default: /tmp)
- `JAMBONES_HTTP_PROXY_IP`, `JAMBONES_HTTP_PROXY_PORT` - HTTP proxy for Azure
- `JAMBONES_AZURE_ENABLE_SSML` - Force SSML wrapper for Azure plain text
## Key Dependencies
- `@jambonz/realtimedb-helpers` - Redis client and hash utilities
- `@google-cloud/text-to-speech` - Google TTS
- `@aws-sdk/client-polly` - AWS Polly
- `microsoft-cognitiveservices-speech-sdk` - Azure TTS
- `@grpc/grpc-js` - gRPC for Riva
- `openai` - OpenAI Whisper TTS
- `bent` - HTTP client for REST-based vendors
-11
View File
@@ -1,11 +0,0 @@
#!/bin/sh
mkdir -p stubs/nuance
for FILE in ./protos/nuance/*; do
grpc_tools_node_protoc \
--js_out=import_style=commonjs,binary:./stubs/nuance \
--grpc_out=grpc_js:./stubs/nuance \
--proto_path=./protos/nuance \
$FILE
done
+1 -3
View File
@@ -14,9 +14,7 @@ module.exports = (opts, logger) => {
purgeTtsCache: require('./lib/purge-tts-cache').bind(null, client, logger), purgeTtsCache: require('./lib/purge-tts-cache').bind(null, client, logger),
addFileToCache: require('./lib/add-file-to-cache').bind(null, client, logger), addFileToCache: require('./lib/add-file-to-cache').bind(null, client, logger),
synthAudio: require('./lib/synth-audio').bind(null, client, createHash, retrieveHash, logger), synthAudio: require('./lib/synth-audio').bind(null, client, createHash, retrieveHash, logger),
getVerbioAccessToken: require('./lib/get-verbio-token').bind(null, client, logger),
getNuanceAccessToken: require('./lib/get-nuance-access-token').bind(null, client, logger),
getIbmAccessToken: require('./lib/get-ibm-access-token').bind(null, client, logger),
getAwsAuthToken: require('./lib/get-aws-sts-token').bind(null, logger, createHash, retrieveHash), getAwsAuthToken: require('./lib/get-aws-sts-token').bind(null, logger, createHash, retrieveHash),
getTtsVoices: require('./lib/get-tts-voices').bind(null, client, createHash, retrieveHash, logger), getTtsVoices: require('./lib/get-tts-voices').bind(null, client, createHash, retrieveHash, logger),
}; };
-48
View File
@@ -1,48 +0,0 @@
const formurlencoded = require('form-urlencoded');
const {Pool} = require('undici');
const pool = new Pool('https://iam.cloud.ibm.com');
const {makeIbmKey, noopLogger} = require('./utils');
const { HTTP_TIMEOUT } = require('./config');
const debug = require('debug')('jambonz:realtimedb-helpers');
async function getIbmAccessToken(client, logger, apiKey) {
logger = logger || noopLogger;
try {
const key = makeIbmKey(apiKey);
const access_token = await client.get(key);
if (access_token) return {access_token, servedFromCache: true};
/* access token not found in cache, so fetch it from Ibm */
const payload = {
grant_type: 'urn:ibm:params:oauth:grant-type:apikey',
apikey: apiKey
};
const {statusCode, headers, body} = await pool.request({
path: '/identity/token',
method: 'POST',
headers: {
'Content-Type': 'application/x-www-form-urlencoded'
},
body: formurlencoded(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
const json = await body.json();
logger.debug({statusCode, headers, body: json}, 'error fetching access token from Ibm');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
await client.set(key, json.access_token, 'EX', json.expires_in - 30);
return {...json, servedFromCache: false};
} catch (err) {
debug(err, 'getIbmAccessToken: Error retrieving Ibm access token');
logger.error(err, 'getIbmAccessToken: Error retrieving Ibm access token for client_id ${clientId}');
throw err;
}
}
module.exports = getIbmAccessToken;
-49
View File
@@ -1,49 +0,0 @@
const formurlencoded = require('form-urlencoded');
const {Pool} = require('undici');
const pool = new Pool('https://auth.crt.nuance.com');
const {makeNuanceKey, makeBasicAuthHeader, noopLogger} = require('./utils');
const { HTTP_TIMEOUT } = require('./config');
const debug = require('debug')('jambonz:realtimedb-helpers');
async function getNuanceAccessToken(client, logger, clientId, secret, scope) {
logger = logger || noopLogger;
try {
const key = makeNuanceKey(clientId, secret, scope);
const access_token = await client.get(key);
if (access_token) return {access_token, servedFromCache: true};
/* access token not found in cache, so fetch it from Nuance */
const payload = {
grant_type: 'client_credentials',
scope
};
const auth = makeBasicAuthHeader(clientId, secret);
const {statusCode, headers, body} = await pool.request({
path: '/oauth2/token',
method: 'POST',
headers: {
...auth,
'Content-Type': 'application/x-www-form-urlencoded'
},
body: formurlencoded(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
logger.debug({statusCode, headers, body: body.text()}, 'error fetching access token from Nuance');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
await client.set(key, json.access_token, 'EX', json.expires_in - 30);
return {...json, servedFromCache: false};
} catch (err) {
debug(err, `getNuanceAccessToken: Error retrieving Nuance access token for client_id ${clientId}`);
logger.error(err, `getNuanceAccessToken: Error retrieving Nuance access token for client_id ${clientId}`);
throw err;
}
}
module.exports = getNuanceAccessToken;
+2 -111
View File
@@ -1,91 +1,8 @@
const assert = require('assert'); const assert = require('assert');
const {noopLogger, createNuanceClient, createKryptonClient} = require('./utils'); const {noopLogger} = require('./utils');
const getNuanceAccessToken = require('./get-nuance-access-token');
const getVerbioAccessToken = require('./get-verbio-token');
const {GetVoicesRequest, Voice} = require('../stubs/nuance/synthesizer_pb');
const TextToSpeechV1 = require('ibm-watson/text-to-speech/v1');
const { IamAuthenticator } = require('ibm-watson/auth');
const ttsGoogle = require('@google-cloud/text-to-speech'); const ttsGoogle = require('@google-cloud/text-to-speech');
const { PollyClient, DescribeVoicesCommand } = require('@aws-sdk/client-polly'); const { PollyClient, DescribeVoicesCommand } = require('@aws-sdk/client-polly');
const getAwsAuthToken = require('./get-aws-sts-token'); const getAwsAuthToken = require('./get-aws-sts-token');
const {Pool} = require('undici');
const { HTTP_TIMEOUT } = require('./config');
const verbioVoicePool = new Pool('https://us.rest.speechcenter.verbio.com');
const getIbmVoices = async(client, logger, credentials) => {
const {tts_region, tts_api_key} = credentials;
console.log(`region: ${tts_region}, api_key: ${tts_api_key}`);
const textToSpeech = new TextToSpeechV1({
authenticator: new IamAuthenticator({
apikey: tts_api_key,
}),
serviceUrl: `https://api.${tts_region}.text-to-speech.watson.cloud.ibm.com`
});
const voices = await textToSpeech.listVoices();
return voices;
};
const getNuanceVoices = async(client, logger, credentials) => {
const {client_id: clientId, secret: secret, nuance_tts_uri} = credentials;
return new Promise(async(resolve, reject) => {
/* get a nuance access token */
let token, nuanceClient;
try {
if (nuance_tts_uri) {
nuanceClient = await createKryptonClient(nuance_tts_uri);
}
else {
const access_token = await getNuanceAccessToken(client, logger, clientId, secret, 'tts');
token = access_token.access_token;
nuanceClient = await createNuanceClient(token);
}
} catch (err) {
logger.error({err}, 'getTtsVoices: error retrieving access token');
return reject(err);
}
/* retrieve all voices */
const v = new Voice();
const request = new GetVoicesRequest();
request.setVoice(v);
nuanceClient.getVoices(request, (err, response) => {
if (err) {
logger.error({err, clientId, secret, token}, 'getTtsVoices: error retrieving voices');
return reject(err);
}
/* return all the voices that are not restricted and eliminate duplicates */
const voices = response.getVoicesList()
.map((v) => {
return {
language: v.getLanguage(),
name: v.getName(),
model: v.getModel(),
gender: v.getGender() === 1 ? 'male' : 'female',
restricted: v.getRestricted()
};
});
const v = voices
.filter((v) => v.restricted === false)
.map((v) => {
delete v.restricted;
return v;
})
.sort((a, b) => {
if (a.language < b.language) return -1;
if (a.language > b.language) return 1;
if (a.name < b.name) return -1;
return 1;
});
const arr = [...new Set(v.map((v) => JSON.stringify(v)))]
.map((v) => JSON.parse(v));
resolve(arr);
});
});
};
const getGoogleVoices = async(_client, logger, credentials) => { const getGoogleVoices = async(_client, logger, credentials) => {
const client = new ttsGoogle.TextToSpeechClient({credentials}); const client = new ttsGoogle.TextToSpeechClient({credentials});
@@ -126,26 +43,6 @@ const getAwsVoices = async(_client, createHash, retrieveHash, logger, credential
} }
}; };
const getVerbioVoices = async(client, logger, credentials) => {
try {
const access_token = await getVerbioAccessToken(client, logger, credentials);
const { body} = await verbioVoicePool.request({
path: '/api/v1/voices',
method: 'GET',
headers: {
'Authorization': `Bearer ${access_token.access_token}`,
'User-Agent': 'jambonz'
},
timeout: HTTP_TIMEOUT,
followRedirects: false
});
return await body.json();
} catch (err) {
logger.info({err}, 'getVerbioVoices - failed to list voices for Verbio');
throw err;
}
};
/** /**
* Synthesize speech to an mp3 file, and also cache the generated speech * Synthesize speech to an mp3 file, and also cache the generated speech
* in redis (base64 format) for 24 hours so as to avoid unnecessarily paying * in redis (base64 format) for 24 hours so as to avoid unnecessarily paying
@@ -165,21 +62,15 @@ const getVerbioVoices = async(client, logger, credentials) => {
async function getTtsVoices(client, createHash, retrieveHash, logger, {vendor, credentials}) { async function getTtsVoices(client, createHash, retrieveHash, logger, {vendor, credentials}) {
logger = logger || noopLogger; logger = logger || noopLogger;
assert.ok(['nuance', 'ibm', 'google', 'aws', 'polly', 'verbio'].includes(vendor), assert.ok(['google', 'aws', 'polly'].includes(vendor),
`getTtsVoices not supported for vendor ${vendor}`); `getTtsVoices not supported for vendor ${vendor}`);
switch (vendor) { switch (vendor) {
case 'nuance':
return getNuanceVoices(client, logger, credentials);
case 'ibm':
return getIbmVoices(client, logger, credentials);
case 'google': case 'google':
return getGoogleVoices(client, logger, credentials); return getGoogleVoices(client, logger, credentials);
case 'aws': case 'aws':
case 'polly': case 'polly':
return getAwsVoices(client, createHash, retrieveHash, logger, credentials); return getAwsVoices(client, createHash, retrieveHash, logger, credentials);
case 'verbio':
return getVerbioVoices(client, logger, credentials);
default: default:
break; break;
} }
-51
View File
@@ -1,51 +0,0 @@
const {Pool} = require('undici');
const { noopLogger, makeVerbioKey } = require('./utils');
const { HTTP_TIMEOUT } = require('./config');
const pool = new Pool('https://auth.speechcenter.verbio.com:444');
const debug = require('debug')('jambonz:realtimedb-helpers');
async function getVerbioAccessToken(client, logger, credentials) {
logger = logger || noopLogger;
const { client_id, client_secret } = credentials;
try {
const key = makeVerbioKey(client_id);
const access_token = await client.get(key);
if (access_token) {
return {access_token, servedFromCache: true};
}
const payload = {
client_id,
client_secret
};
const {statusCode, headers, body} = await pool.request({
path: '/api/v1/token',
method: 'POST',
headers: {
'Content-Type': 'application/json',
'User-Agent': 'jambonz'
},
body: JSON.stringify(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
logger.debug({statusCode, headers, body: await body.text()}, 'error fetching access token from Verbio');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
const expiry = Math.floor(json.expiration_time - Date.now() / 1000 - 30);
await client.set(key, json.access_token, 'EX', expiry);
return {...json, servedFromCache: false};
} catch (err) {
debug(err, `getVerbioAccessToken: Error retrieving Verbio access token for client_id ${client_id}`);
logger.error(err, `getVerbioAccessToken: Error retrieving Verbio access token for client_id ${client_id}`);
throw err;
}
}
module.exports = getVerbioAccessToken;
+869 -386
View File
File diff suppressed because it is too large Load Diff
+1 -107
View File
@@ -1,20 +1,7 @@
const crypto = require('crypto'); const crypto = require('crypto');
const {SynthesizerClient} = require('../stubs/nuance/synthesizer_grpc_pb');
const {RivaSpeechSynthesisClient} = require('../stubs/riva/proto/riva_tts_grpc_pb'); const {RivaSpeechSynthesisClient} = require('../stubs/riva/proto/riva_tts_grpc_pb');
const {Pool} = require('undici');
const pool = new Pool('https://auth.crt.nuance.com');
const NUANCE_AUTH_ENDPOINT = 'tts.api.nuance.com:443';
const grpc = require('@grpc/grpc-js'); const grpc = require('@grpc/grpc-js');
const formurlencoded = require('form-urlencoded'); const { TMP_FOLDER } = require('./config');
const { TMP_FOLDER, HTTP_TIMEOUT } = require('./config');
const debug = require('debug')('jambonz:realtimedb-helpers');
/**
* Future TODO: cache recently used connections to providers
* to avoid connection overhead during a call.
* Will need to periodically age them out to avoid memory leaks.
*/
//const nuanceClientMap = new Map();
function makeSynthKey({ function makeSynthKey({
account_sid = '', account_sid = '',
@@ -45,96 +32,12 @@ const noopLogger = {
error: () => {} error: () => {}
}; };
const toBase64 = (str) => Buffer.from(str || '', 'utf8').toString('base64');
function makeBasicAuthHeader(username, password) {
if (!username || !password) return {};
const creds = `${encodeURIComponent(username)}:${password || ''}`;
const header = `Basic ${toBase64(creds)}`;
return {Authorization: header};
}
function makeIbmKey(apiKey) {
const hash = crypto.createHash('sha1');
hash.update(apiKey);
return `ibm:${hash.digest('hex')}`;
}
function makeAwsKey(awsAccessKeyId) { function makeAwsKey(awsAccessKeyId) {
const hash = crypto.createHash('sha1'); const hash = crypto.createHash('sha1');
hash.update(awsAccessKeyId); hash.update(awsAccessKeyId);
return `aws:${hash.digest('hex')}`; return `aws:${hash.digest('hex')}`;
} }
function makePlayhtKey(apiKey) {
const hash = crypto.createHash('sha1');
hash.update(apiKey);
return `playht:${hash.digest('hex')}`;
}
function makeVerbioKey(client_id) {
const hash = crypto.createHash('sha1');
hash.update(client_id);
return `verbio:${hash.digest('hex')}`;
}
function makeNuanceKey(clientId, secret, scope) {
const hash = crypto.createHash('sha1');
hash.update(`${clientId}:${secret}:${scope}`);
return `nuance:${hash.digest('hex')}`;
}
const getNuanceAccessToken = async(clientId, secret, scope = 'asr tts') => {
const payload = {
grant_type: 'client_credentials',
scope
};
const auth = makeBasicAuthHeader(clientId, secret);
const {statusCode, headers, body} = await pool.request({
path: '/oauth2/token',
method: 'POST',
headers: {
...auth,
'Content-Type': 'application/x-www-form-urlencoded'
},
body: formurlencoded(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
debug({statusCode, headers, body: body.text()}, 'error fetching access token from Nuance');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
return json.access_token;
};
const createKryptonClient = async(uri) => {
const client = new SynthesizerClient(uri, grpc.credentials.createInsecure());
return client;
};
const createNuanceClient = async(access_token) => {
//if (nuanceClientMap.has(access_token)) return nuanceClientMap.get(access_token);
const generateMetadata = (params, callback) => {
var metadata = new grpc.Metadata();
metadata.add('authorization', `Bearer ${access_token}`);
callback(null, metadata);
};
const sslCreds = grpc.credentials.createSsl();
const authCreds = grpc.credentials.createFromMetadataGenerator(generateMetadata);
const combined_creds = grpc.credentials.combineChannelCredentials(sslCreds, authCreds);
const client = new SynthesizerClient(NUANCE_AUTH_ENDPOINT, combined_creds);
//if (process.env.NUANCE_CACHE_TTS_CONNECTIONS) nuanceClientMap.set(access_token, client);
return client;
};
const createRivaClient = async(rivaUri) => { const createRivaClient = async(rivaUri) => {
const client = new RivaSpeechSynthesisClient(rivaUri, grpc.credentials.createInsecure()); const client = new RivaSpeechSynthesisClient(rivaUri, grpc.credentials.createInsecure());
return client; return client;
@@ -142,17 +45,8 @@ const createRivaClient = async(rivaUri) => {
module.exports = { module.exports = {
makeSynthKey, makeSynthKey,
makeNuanceKey,
makeIbmKey,
makePlayhtKey,
makeAwsKey, makeAwsKey,
makeVerbioKey,
getNuanceAccessToken,
createNuanceClient,
createKryptonClient,
createRivaClient, createRivaClient,
makeBasicAuthHeader,
NUANCE_AUTH_ENDPOINT,
noopLogger, noopLogger,
makeFilePath makeFilePath
}; };
+2969 -2838
View File
File diff suppressed because it is too large Load Diff
+6 -11
View File
@@ -1,6 +1,6 @@
{ {
"name": "@jambonz/speech-utils", "name": "@jambonz/speech-utils",
"version": "0.2.20", "version": "1.0.21",
"description": "TTS-related speech utilities for jambonz", "description": "TTS-related speech utilities for jambonz",
"main": "index.js", "main": "index.js",
"author": "Dave Horton", "author": "Dave Horton",
@@ -13,7 +13,6 @@
"coverage": "nyc --reporter html --report-dir ./coverage npm run test", "coverage": "nyc --reporter html --report-dir ./coverage npm run test",
"jslint": "eslint index.js lib", "jslint": "eslint index.js lib",
"jslint:fix": "npm run jslint --fix", "jslint:fix": "npm run jslint --fix",
"build": "./build_stubs.sh",
"prepare": "husky" "prepare": "husky"
}, },
"repository": { "repository": {
@@ -26,24 +25,20 @@
}, },
"homepage": "https://github.com/jambonz/speech-utils#readme", "homepage": "https://github.com/jambonz/speech-utils#readme",
"dependencies": { "dependencies": {
"23": "^0.0.0",
"@aws-sdk/client-polly": "^3.496.0", "@aws-sdk/client-polly": "^3.496.0",
"@aws-sdk/client-sts": "^3.496.0", "@aws-sdk/client-sts": "^3.496.0",
"@cartesia/cartesia-js": "^2.1.0", "@cartesia/cartesia-js": "^2.2.7",
"@google-cloud/text-to-speech": "^5.5.0", "@google-cloud/text-to-speech": "^6.4.0",
"@grpc/grpc-js": "^1.9.14", "@grpc/grpc-js": "^1.9.14",
"@jambonz/realtimedb-helpers": "^0.8.7", "@jambonz/realtimedb-helpers": "^0.8.7",
"bent": "^7.3.12", "bent": "^7.3.12",
"debug": "^4.3.4", "debug": "^4.3.4",
"form-urlencoded": "^6.1.4",
"google-protobuf": "^3.21.2", "google-protobuf": "^3.21.2",
"ibm-watson": "^11.0.0", "microsoft-cognitiveservices-speech-sdk": "^1.51.0",
"microsoft-cognitiveservices-speech-sdk": "1.38.0", "openai": "^4.98.0"
"openai": "^4.98.0",
"undici": "^7.5.0"
}, },
"devDependencies": { "devDependencies": {
"config": "^3.3.11", "config": "^4.2.0",
"eslint": "^9.3.0", "eslint": "^9.3.0",
"eslint-plugin-promise": "^6.2.0", "eslint-plugin-promise": "^6.2.0",
"husky": "^9.0.11", "husky": "^9.0.11",
-332
View File
@@ -1,332 +0,0 @@
syntax = "proto3";
package nuance.tts.v1;
/*
The Synthesizer service offers these functionalities:
* - GetVoices: Queries the list of available voices, with filters to reduce the search space.
* - Synthesize: Synthesizes audio from input text and parameters, and returns an audio stream.
* - UnarySynthesize: Synthesizes audio from input text and parameters, and returns a single audio response.
*/
service Synthesizer {
rpc GetVoices(GetVoicesRequest) returns (GetVoicesResponse) {}
rpc Synthesize(SynthesisRequest) returns (stream SynthesisResponse) {}
rpc UnarySynthesize(SynthesisRequest) returns (UnarySynthesisResponse) {}
}
/*
Input message for [Synthesizer](#synthesizer) - GetVoices, to query voices available to the client.
*/
message GetVoicesRequest {
Voice voice = 1; // Optionally filter the voices to retrieve, e.g. set language to en-US to return only American English voices.
}
/*
Output message for [Synthesizer](#synthesizer) - GetVoices. Includes a list of voices that matched the input criteria, if any.
*/
message GetVoicesResponse {
repeated Voice voices = 1; // Repeated. Voices and characteristics returned.
}
/*
Input message for [Synthesizer](#synthesizer) - Synthesize. Specifies input text, audio parameters, and events to subscribe to, in exchange for synthesized audio.
*/
message SynthesisRequest {
Voice voice = 1; // The voice to use for audio synthesis. Mandatory.
AudioParameters audio_params = 2; // Output audio parameters, such as encoding and volume.
Input input = 3; // Input text to synthesize, tuning data, etc. Mandatory.
EventParameters event_params = 4; // Markers and other info to include in server events returned during synthesis.
map<string, string> client_data = 5; // Repeated. Client-supplied key-value pairs to inject into the call log.
string user_id = 6; // Identifies a particular user within an application.
}
/*
Input or output message for voices. When sent as input:
* - In [GetVoicesRequest](#getvoicesrequest), it filters the list of available voices.
* - In [SynthesisRequest](#synthesisrequest), it specifies the voice to use for synthesis.
When received as output in [GetVoicesResponse](#getvoicesresponse), it returns the list of available voices.
*/
message Voice {
string name = 1; // The voice's name, e.g. 'Evan'. Mandatory for SynthesisRequest.
string model = 2; // The voice's quality model, e.g. 'enhanced' or 'standard'. Mandatory for SynthesisRequest.
string language = 3; // IETF language code, e.g. 'en-US'. Used only for GetVoicesRequest and GetVoicesResponse, to search for voices with a certain mother tongue. Ignored otherwise.
EnumAgeGroup age_group = 4; // Used only for GetVoicesRequest and GetVoicesResponse, to search for adult or child voices. Ignored otherwise.
EnumGender gender = 5; // Used only for GetVoicesRequest and GetVoicesResponse, to search for voices with a certain gender. Ignored otherwise.
uint32 sample_rate_hz = 6; // Used only for GetVoicesRequest and GetVoicesResponse, to search for a certain native sample rate. Ignored otherwise.
string language_tlw = 7; // Used only for GetVoicesRequest and GetVoicesResponse. Three-letter language code (e.g. 'enu' for American English), for configuring language identification in [Input](#input).
bool restricted = 8; // Used only in GetVoicesResponse, to identify restricted voices. These are custom voices available only to specific customers. Ignored otherwise.
string version = 9; // Used only in GetVoicesResponse, to return the voice's version. Ignored otherwise.
}
/*
Input or output field specifying whether the voice uses its adult or child version, if available. Included in [Voice](#voice).
*/
enum EnumAgeGroup {
ADULT = 0; // Adult voice. Default for GetVoicesRequest.
CHILD = 1; // Child voice.
}
/*
Input or output field, specifying gender for voices that support multiple genders. Included in [Voice](#voice).
*/
enum EnumGender {
ANY = 0; // Any gender voice. Default for GetVoicesRequest.
MALE = 1; // Male voice.
FEMALE = 2; // Female voice.
NEUTRAL = 3; // Neutral gender voice.
}
/*
Input message for audio-related parameters during synthesis, including encoding, volume, and audio length. Included in [SynthesisRequest](#synthesisrequest).
*/
message AudioParameters {
AudioFormat audio_format = 1; // Audio encoding. Default PCM 22.5kHz.
uint32 volume_percentage = 2; // Volume amplitude, from 0 to 100. Default 80.
float speaking_rate_factor = 3; // Speaking rate, from 0 to 2.0. Default 1.0.
uint32 audio_chunk_duration_ms = 4; // Maximum duration, in ms, of an audio chunk delivered to the client, from 1 to 60000. Default is 20000 (20 seconds). When this parameter is large enough (for example, 20 or 30 seconds), each audio chunk contains an audible segment surrounded by silence.
uint32 target_audio_length_ms = 5; // Maximum duration, in ms, of synthesized audio. When greater than 0, the server stops ongoing synthesis at the first sentence end, or silence, closest to the value.
bool disable_early_emission = 6; // By default, audio segments are emitted as soon as possible, even if they are not audible. This behavior may be disabled.
}
/*
Input message for audio encoding of synthesized text. Included in [AudioParameters](#audioparameters).
*/
message AudioFormat {
oneof audio_format {
PCM pcm = 1; // Signed 16-bit little endian PCM.
ALaw alaw = 2; // G.711 A-law, 8kHz.
ULaw ulaw = 3; // G.711 Mu-law, 8kHz.
OggOpus ogg_opus = 4; // Ogg Opus, 8kHz, 16kHz or 24kHz.
Opus opus = 5; // Opus, 8kHz, 16kHZ or 24 kHz. The audio will be sent one Opus packet at a time.
}
}
/*
Input message defining PCM sample rate. Included in [Audioformat](#audioformat).
*/
message PCM {
uint32 sample_rate_hz = 1; // Output sample rate: 8000, 16000, 22050 (default), 24000.
}
/*
Input message defining A-law audio format. G.711 audio formats are set to 8kHz. Included in [Audioformat](#audioformat).
*/
message ALaw {}
/*
Input message defining Mu-law audio format. G.711 audio formats are set to 8kHz. Included in [Audioformat](#audioformat).
*/
message ULaw {}
/*
Input message defining Opus output rate. Included in [Audioformat](#audioformat).
*/
message Opus {
uint32 sample_rate_hz = 1; // Output sample rate. Supported values: 8000, 16000, 24000 Hz.
uint32 bit_rate_bps = 2; // Valid range is 500 to 256000 bps. Default 28000 bps.
float max_frame_duration_ms = 3; // Opus frame size, in ms: 2.5, 5, 10, 20, 40, 60. Default 20.
uint32 complexity = 4; // Computational complexity. A complexity of 0 means the codec default.
EnumVariableBitrate vbr = 5; // Variable bitrate. On by default.
}
/*
Input message defining Ogg Opus output rate. Included in [Audioformat](#audioformat).
*/
message OggOpus {
uint32 sample_rate_hz = 1; // Output sample rate. Supported values: 8000, 16000, 24000 Hz.
uint32 bit_rate_bps = 2; // Valid range is 500 to 256000 bps. Default 28000 bps.
float max_frame_duration_ms = 3; // Opus frame size, in ms: 2.5, 5, 10, 20, 40, 60. Default 20.
uint32 complexity = 4; // Computational complexity. A complexity of 0 means the codec default.
EnumVariableBitrate vbr = 5; // Variable bitrate. On by default.
}
/*
Settings for variable bitrate. Included in [OggOpus](#oggopus). Turned on by default.
*/
enum EnumVariableBitrate {
VARIABLE_BITRATE_ON = 0; // Use variable bitrate. Default.
VARIABLE_BITRATE_OFF = 1; // Do not use variable bitrate.
VARIABLE_BITRATE_CONSTRAINED = 2; // Use constrained variable bitrate.
}
/*
Input message containing text to synthesize and synthesis parameters, including tuning data, etc. Included in [SynthesisRequest](#synthesisrequest). The type of input may be:
* - Plain text.
* - An SSML document.
* - An alternating sequence of plain text and Nuance control codes.
*/
message Input {
oneof input_data {
Text text = 1; // Text input.
SSML ssml = 2; // SSML input.
TokenizedSequence tokenized_sequence = 3; // Sequence of text and Nuance control codes.
}
repeated SynthesisResource resources = 4; // Repeated. Synthesis resources (user dictionaries, rulesets, etc.) to tune synthesized audio.
LanguageIdentificationParameters lid_params = 5; // LID parameters.
DownloadParameters download_params = 6; // Remote file download parameters.
}
/*
Input message for synthesizing plain text. The encoding must be in UTF-8.
*/
message Text {
oneof text_data {
string text = 1; // Plain input text in UTF-8 encoding.
string uri = 2; // Remote URI to the plain input text. Disabled for Nuance-hosted TTS.
}
}
/*
Input message for synthesizing an SSML document.
*/
message SSML {
oneof ssml_data {
string text = 1; // SSML input text.
string uri = 2; // Remote URI to the SSML input text. Disabled for Nuance-hosted TTS.
}
EnumSSMLValidationMode ssml_validation_mode = 3; // SSML validation mode. Default STRICT.
}
/*
Input message for synthesizing a sequence of plain text and Nuance control codes.
*/
message TokenizedSequence {
repeated Token tokens = 1;
}
/*
The unit when using TokenizedSequence for input. Each token can either be plain text or a Nuance control code.
*/
message Token {
oneof token_data {
string text = 1; // Plain input text.
ControlCode control_code = 2; // Nuance control code.
}
}
/*
A Nuance control code allows the user to control how text is spoken, similarly to SSML.
*/
message ControlCode {
string key = 1; // Name of the control code, e.g. "pause".
string value = 2; // Value of the control code.
}
/*
Input message specifying the type of file to tune the synthesized output and its location or contents. Included in [Input](#input).
*/
message SynthesisResource {
EnumResourceType type = 1; // Resource type, e.g. user dictionary, etc. Default USER_DICTIONARY.
oneof resource_data {
string uri = 2; // URI to the remote resource, or
bytes body = 3; // For EnumResourceType USER_DICTIONARY, the contents of the file.
}
}
/*
The type of synthesis resource to tune the output. Included in [SynthesisResource](#synthesisresource). User dictionaries provide custom pronunciations, rulesets apply search-and-replace rules to input text and ActivePrompt databases help tune synthesized audio under certain conditions, using Nuance Vocalizer Studio.
*/
enum EnumResourceType {
USER_DICTIONARY = 0; // User dictionary (application/edct-bin-dictionary). Default.
TEXT_USER_RULESET = 1; // Text user ruleset (application/x-vocalizer-rettt+text).
BINARY_USER_RULESET = 2; // Binary user ruleset (application/x-vocalizer-rettt+bin).
ACTIVEPROMPT_DB = 3; // ActivePrompt database (application/x-vocalizer/activeprompt-db).
ACTIVEPROMPT_DB_AUTO = 4; // ActivePrompt database with automatic insertion (application/x-vocalizer/activeprompt-db;mode=automatic).
SYSTEM_DICTIONARY = 5; // Nuance system dictionary (application/sdct-bin-dictionary).
}
/*
SSML validation mode when using SSML input. Included in [Input](#input). Strict by default but can be relaxed.
*/
enum EnumSSMLValidationMode {
STRICT = 0; // Strict SSL validation. Default.
WARN = 1; // Give warning only.
NONE = 2; // Do not validate.
}
/*
Input message controlling the language identifier. Included in [Input](#input). The language identifier runs on input blocks labeled with the <ESC>\lang=unknown\ control sequence or SSML xml:lang="unknown". The language identifier automatically restricts the matched languages to the installed voices. This limits the permissible languages, and also sets the order of precedence (first to last) when they have equal confidence scores.
*/
message LanguageIdentificationParameters {
bool disable = 1; // Whether to disable language identification. Turned on by default.
repeated string languages = 2; // Repeated. List of three-letter language codes (e.g. enu, frc, spm) to restrict language identification results, in order of precedence. Use GetVoicesRequest - Voice - language_tlw to obtain the three-letter codes. Default blank.
bool always_use_highest_confidence = 3; // If enabled, language identification always chooses the language with the highest confidence score, even if the score is low. Default false, meaning use language with any confidence.
}
/*
Input message containing parameters for remote file download, whether for input text (Input.uri) or a SynthesisResource (SynthesisResource.uri). Included in [Input](#input).
*/
message DownloadParameters {
map <string,string> headers = 1; // HTTP headers to include in outgoing requests. Only whitelisted headers will actually be sent.
bool refuse_cookies = 2; // Whether to disable cookies. By default, HTTP requests accept cookies.
oneof optional_download_parameter_request_timeout_ms {
uint32 request_timeout_ms = 3; // Request timeout in ms. Default (0) means server default, usually 30000 (30 seconds).
}
}
/*
Input message that defines event subscription parameters. Included in [SynthesisRequest](#synthesisrequest). Events that are requested are sent throughout the SynthesisResponse stream, when generated. Marker events can send events as certain parts of the synthesized audio are reached, for example, at the end of a word, sentence, or user-defined bookmark.
* Log events are produced throughout a synthesis request for events such as a voice loaded by the server or an audio chunk being ready to send.
*/
message EventParameters {
bool send_sentence_marker_events = 1; // Sentence marker. Default: do not send.
bool send_word_marker_events = 2; // Word marker. Default: do not send.
bool send_phoneme_marker_events = 3; // Phoneme marker. Default: do not send.
bool send_bookmark_marker_events = 4; // Bookmark marker. Default: do not send.
bool send_paragraph_marker_events = 5; // Paragraph marker. Default: do not send.
bool send_visemes = 6; // Lipsync information. Default: do not send.
bool send_log_events = 7; // Whether to log events during synthesis. By default, logging is turned off.
bool suppress_input = 8; // Whether to omit input text and URIs from log events. By default, these items are included.
}
/*
The [Synthesizer](#synthesizer) - Synthesize RPC call returns a stream of SynthesisResponse messages. The response contains one of:
* - A status response, indicating completion or failure of the request. This is received only once and signifies the end of a Synthesize call.
* - A list of events the client has requested. This can be received many times. See EventParameters for details.
* - An audio buffer. This may be received many times.
*/
message SynthesisResponse {
oneof response {
Status status = 1; // A status response, indicating completion or failure of the request.
Events events = 2; // A list of events. See EventParameters for details.
bytes audio = 3; // The latest audio buffer.
}
}
/*
The [Synthesizer](#synthesizer) - UnarySynthesize RPC call returns a single UnarySynthesisResponse message. It is similar to a SynthesisResponse message but includes all the information instead of a single type of response. The response contains:
* - A status response, indicating completion or failure of the request.
* - A list of events the client has requested. See EventParameters for details.
* - The complete audio buffer of the synthesized text.
*/
message UnarySynthesisResponse {
Status status = 1; // A status response, indicating completion or failure of the request.
Events events = 2; // A list of events. See EventParameters for details.
bytes audio = 3; // Audio buffer of the synthesized text.
}
/*
Output message containing a status response, indicating completion or failure of a SynthesisRequest. Included in [SynthesisResponse](#synthesisresponse) and [UnarySynthesisResponse](#unarysynthesisresponse).
*/
message Status {
uint32 code = 1; // HTTP-style return code: 200, 4xx, or 5xx as appropriate.
string message = 2; // Brief description of the status.
string details = 3; // Longer description if available.
}
/*
Output message defining a container for a list of events. This container is needed because oneof does not allow repeated parameters in Protobuf. Included in [SynthesisResponse](#synthesisresponse) and [UnarySynthesisResponse](#unarysynthesisresponse).
*/
message Events {
repeated Event events = 1; // Repeated. One or more events.
}
/*
Output message defining an event message. Included in [Events](#events). See EventParameters for details.
*/
message Event {
string name = 1; // Either "Markers" or the name of the event in the case of a Log Event.
map<string, string> values = 2; // Repeated. Key-value data relevant to the current event.
}
-104
View File
@@ -1,104 +0,0 @@
// GENERATED CODE -- DO NOT EDIT!
'use strict';
var grpc = require('@grpc/grpc-js');
var synthesizer_pb = require('./synthesizer_pb.js');
function serialize_nuance_tts_v1_GetVoicesRequest(arg) {
if (!(arg instanceof synthesizer_pb.GetVoicesRequest)) {
throw new Error('Expected argument of type nuance.tts.v1.GetVoicesRequest');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_GetVoicesRequest(buffer_arg) {
return synthesizer_pb.GetVoicesRequest.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_GetVoicesResponse(arg) {
if (!(arg instanceof synthesizer_pb.GetVoicesResponse)) {
throw new Error('Expected argument of type nuance.tts.v1.GetVoicesResponse');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_GetVoicesResponse(buffer_arg) {
return synthesizer_pb.GetVoicesResponse.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_SynthesisRequest(arg) {
if (!(arg instanceof synthesizer_pb.SynthesisRequest)) {
throw new Error('Expected argument of type nuance.tts.v1.SynthesisRequest');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_SynthesisRequest(buffer_arg) {
return synthesizer_pb.SynthesisRequest.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_SynthesisResponse(arg) {
if (!(arg instanceof synthesizer_pb.SynthesisResponse)) {
throw new Error('Expected argument of type nuance.tts.v1.SynthesisResponse');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_SynthesisResponse(buffer_arg) {
return synthesizer_pb.SynthesisResponse.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_UnarySynthesisResponse(arg) {
if (!(arg instanceof synthesizer_pb.UnarySynthesisResponse)) {
throw new Error('Expected argument of type nuance.tts.v1.UnarySynthesisResponse');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_UnarySynthesisResponse(buffer_arg) {
return synthesizer_pb.UnarySynthesisResponse.deserializeBinary(new Uint8Array(buffer_arg));
}
//
// The Synthesizer service offers these functionalities:
// - GetVoices: Queries the list of available voices, with filters to reduce the search space.
// - Synthesize: Synthesizes audio from input text and parameters, and returns an audio stream.
// - UnarySynthesize: Synthesizes audio from input text and parameters, and returns a single audio response.
var SynthesizerService = exports.SynthesizerService = {
getVoices: {
path: '/nuance.tts.v1.Synthesizer/GetVoices',
requestStream: false,
responseStream: false,
requestType: synthesizer_pb.GetVoicesRequest,
responseType: synthesizer_pb.GetVoicesResponse,
requestSerialize: serialize_nuance_tts_v1_GetVoicesRequest,
requestDeserialize: deserialize_nuance_tts_v1_GetVoicesRequest,
responseSerialize: serialize_nuance_tts_v1_GetVoicesResponse,
responseDeserialize: deserialize_nuance_tts_v1_GetVoicesResponse,
},
synthesize: {
path: '/nuance.tts.v1.Synthesizer/Synthesize',
requestStream: false,
responseStream: true,
requestType: synthesizer_pb.SynthesisRequest,
responseType: synthesizer_pb.SynthesisResponse,
requestSerialize: serialize_nuance_tts_v1_SynthesisRequest,
requestDeserialize: deserialize_nuance_tts_v1_SynthesisRequest,
responseSerialize: serialize_nuance_tts_v1_SynthesisResponse,
responseDeserialize: deserialize_nuance_tts_v1_SynthesisResponse,
},
unarySynthesize: {
path: '/nuance.tts.v1.Synthesizer/UnarySynthesize',
requestStream: false,
responseStream: false,
requestType: synthesizer_pb.SynthesisRequest,
responseType: synthesizer_pb.UnarySynthesisResponse,
requestSerialize: serialize_nuance_tts_v1_SynthesisRequest,
requestDeserialize: deserialize_nuance_tts_v1_SynthesisRequest,
responseSerialize: serialize_nuance_tts_v1_UnarySynthesisResponse,
responseDeserialize: deserialize_nuance_tts_v1_UnarySynthesisResponse,
},
};
exports.SynthesizerClient = grpc.makeGenericClientConstructor(SynthesizerService);
File diff suppressed because it is too large Load Diff
+23
View File
@@ -0,0 +1,23 @@
/**
* Resolve AWS credentials for the test suite.
*
* Returns null when AWS is not configured, so the caller skips.
*
* When AWS_SESSION_TOKEN is set the credentials are temporary -- GitHub OIDC in CI, or
* `aws sso login` locally -- and the key pair must not be passed through. Doing so sends
* lib/get-aws-sts-token.js down its accessKeyId branch, which calls GetSessionToken, and
* AWS rejects GetSessionToken when it is called with session credentials. Returning the
* region alone routes to the SDK's default credential provider chain instead, which
* handles temporary credentials correctly.
*/
module.exports = () => {
const region = process.env.AWS_REGION;
if (!region) return null;
const accessKeyId = process.env.AWS_ACCESS_KEY_ID;
const secretAccessKey = process.env.AWS_SECRET_ACCESS_KEY;
if (process.env.AWS_SESSION_TOKEN) return {region};
if (accessKeyId && secretAccessKey) return {accessKeyId, secretAccessKey, region};
return {region};
};
+11 -11
View File
@@ -6,33 +6,33 @@ process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason); console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
}); });
const awsCredentials = require('./aws-credentials');
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
test('AWS - create and cache auth token', async(t) => { test('AWS - create and cache auth token', async(t) => {
const fn = require('..'); const fn = require('..');
const {client, getAwsAuthToken} = fn(opts, logger); const {client, getAwsAuthToken} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) { const credentials = awsCredentials();
if (!credentials) {
t.pass('skipping AWS auth token tests since no AWS credentials provided'); t.pass('skipping AWS auth token tests since no AWS credentials provided');
t.end(); t.end();
client.quit(); client.quit();
return; return;
} }
// getAwsAuthToken derives its cache key from roleArn || accessKeyId || speech_credential_sid.
// With temporary credentials none of the first two are passed, so supply a stable sid --
// which is what production does for instance-profile credentials.
const args = {...credentials, speech_credential_sid: 'test-aws-speech-credential'};
try { try {
let obj = await getAwsAuthToken({ let obj = await getAwsAuthToken(args);
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
//console.log({obj}, 'received auth token from AWS'); //console.log({obj}, 'received auth token from AWS');
t.ok(obj.securityToken && !obj.servedFromCache, 'successfullY generated auth token from AWS'); t.ok(obj.securityToken && !obj.servedFromCache, 'successfullY generated auth token from AWS');
await sleep(250); await sleep(250);
obj = await getAwsAuthToken({ obj = await getAwsAuthToken(args);
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
//console.log({obj}, 'received auth token from AWS - second request'); //console.log({obj}, 'received auth token from AWS - second request');
t.ok(obj.securityToken && obj.servedFromCache, 'successfully received access token from cache'); t.ok(obj.securityToken && obj.servedFromCache, 'successfully received access token from cache');
+1 -1
View File
@@ -2,7 +2,7 @@ const test = require('tape').test ;
const exec = require('child_process').exec ; const exec = require('child_process').exec ;
test('starting docker network..', (t) => { test('starting docker network..', (t) => {
exec(`docker-compose -f ${__dirname}/docker-compose-testbed.yaml up -d`, (err, stdout, stderr) => { exec(`docker compose -f ${__dirname}/docker-compose-testbed.yaml up -d`, (err, stdout, stderr) => {
setTimeout(() => { setTimeout(() => {
t.end(err); t.end(err);
}, 2000); }, 2000);
+1 -1
View File
@@ -3,7 +3,7 @@ const exec = require('child_process').exec ;
test('stopping docker network..', (t) => { test('stopping docker network..', (t) => {
t.timeoutAfter(10000); t.timeoutAfter(10000);
exec(`docker-compose -f ${__dirname}/docker-compose-testbed.yaml down`, (err, stdout, stderr) => { exec(`docker compose -f ${__dirname}/docker-compose-testbed.yaml down`, (err, stdout, stderr) => {
//console.log(`stderr: ${stderr}`); //console.log(`stderr: ${stderr}`);
process.exit(0); process.exit(0);
}); });
-78
View File
@@ -1,78 +0,0 @@
const test = require('tape').test ;
const config = require('config');
const opts = config.get('redis');
const fs = require('fs');
const logger = require('pino')({level: 'error'});
process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
});
const stats = {
increment: () => {},
histogram: () => {}
};
test('IBM - create access key', async(t) => {
const fn = require('..');
const {client, getIbmAccessToken} = fn(opts, logger);
if (!process.env.IBM_API_KEY ) {
t.pass('skipping IBM test since no IBM api_key provided');
t.end();
client.quit();
return;
}
try {
let obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM');
t.ok(obj.access_token && !obj.servedFromCache, 'successfull received access token from IBM');
obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM - second request');
t.ok(obj.access_token && obj.servedFromCache, 'successfully received access token from cache');
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('IBM - retrieve tts voices test', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
t.pass('skipping IBM test since no IBM api_key and/or region provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'ibm',
credentials: {
tts_api_key: process.env.IBM_TTS_API_KEY,
tts_region: process.env.IBM_TTS_REGION
}
};
const obj = await getTtsVoices(opts);
const {voices} = obj.result;
//console.log(JSON.stringify(voices));
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from IBM`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
+1 -2
View File
@@ -2,6 +2,5 @@ require('./docker_start');
require('./synth'); require('./synth');
require('./list-voices'); require('./list-voices');
require('./aws'); require('./aws');
require('./ibm');
require('./nuance');
require('./docker_stop'); require('./docker_stop');
+5 -165
View File
@@ -1,4 +1,5 @@
const test = require('tape').test ; const test = require('tape').test ;
const awsCredentials = require('./aws-credentials');
const config = require('config'); const config = require('config');
const opts = config.get('redis'); const opts = config.get('redis');
const fs = require('fs'); const fs = require('fs');
@@ -12,164 +13,6 @@ const stats = {
histogram: () => {} histogram: () => {}
}; };
test('Verbio - get Access key and voices', async(t) => {
const fn = require('..');
const {client, getTtsVoices, getVerbioAccessToken} = fn(opts, logger);
if (!process.env.VERBIO_CLIENT_ID || !process.env.VERBIO_CLIENT_SECRET) {
t.pass('skipping Verbio test since no Verbio Keys provided');
t.end();
client.quit();
return;
}
try {
const credentials = {
client_id: process.env.VERBIO_CLIENT_ID,
client_secret: process.env.VERBIO_CLIENT_SECRET
};
let obj = await getVerbioAccessToken(credentials);
t.ok(obj.access_token , 'successfully received access token not from cache');
const voices = await getTtsVoices({vendor: 'verbio', credentials});
t.ok(voices && voices.length != 0, 'successfully received verbio voices');
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('IBM - create access key', async(t) => {
const fn = require('..');
const {client, getIbmAccessToken} = fn(opts, logger);
if (!process.env.IBM_API_KEY ) {
t.pass('skipping IBM test since no IBM api_key provided');
t.end();
client.quit();
return;
}
try {
let obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM');
t.ok(obj.access_token && !obj.servedFromCache, 'successfull received access token from IBM');
obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM - second request');
t.ok(obj.access_token && obj.servedFromCache, 'successfully received access token from cache');
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('IBM - retrieve tts voices test', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
t.pass('skipping IBM test since no IBM api_key and/or region provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'ibm',
credentials: {
tts_api_key: process.env.IBM_TTS_API_KEY,
tts_region: process.env.IBM_TTS_REGION
}
};
const obj = await getTtsVoices(opts);
const {voices} = obj.result;
//console.log(JSON.stringify(voices));
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from IBM`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance hosted tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET ) {
t.pass('skipping Nuance hosted test since no Nuance client_id and secret provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
client_id: process.env.NUANCE_CLIENT_ID,
secret: process.env.NUANCE_SECRET
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance on-prem tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_TTS_URI ) {
t.pass('skipping Nuance on-prem test since no Nuance uri provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
nuance_tts_uri: process.env.NUANCE_TTS_URI
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Google tests', async(t) => { test('Google tests', async(t) => {
const fn = require('..'); const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger); const {client, getTtsVoices} = fn(opts, logger);
@@ -203,18 +46,15 @@ test('AWS tests', async(t) => {
const fn = require('..'); const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger); const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.AWS_ACCESS_KEY_ID || !process.env.AWS_SECRET_ACCESS_KEY || !process.env.AWS_REGION) { const credentials = awsCredentials();
t.pass('skipping AWS speech synth tests since AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, or AWS_REGION not provided'); if (!credentials) {
t.pass('skipping AWS speech synth tests since AWS_REGION not provided');
return t.end(); return t.end();
} }
try { try {
const opts = { const opts = {
vendor: 'aws', vendor: 'aws',
credentials: { credentials
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION,
}
}; };
let result = await getTtsVoices(opts); let result = await getTtsVoices(opts);
t.ok(result?.Voices?.length > 0, `GetVoices: successfully retrieved ${result.Voices.length} voices from AWS`); t.ok(result?.Voices?.length > 0, `GetVoices: successfully retrieved ${result.Voices.length} voices from AWS`);
-81
View File
@@ -1,81 +0,0 @@
const test = require('tape').test ;
const config = require('config');
const opts = config.get('redis');
const fs = require('fs');
const logger = require('pino')({level: 'error'});
process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
});
const stats = {
increment: () => {},
histogram: () => {}
};
test('Nuance hosted tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET ) {
t.pass('skipping Nuance hosted test since no Nuance client_id and secret provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
client_id: process.env.NUANCE_CLIENT_ID,
secret: process.env.NUANCE_SECRET
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance on-prem tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_TTS_URI ) {
t.pass('skipping Nuance on-prem test since no Nuance uri provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
nuance_tts_uri: process.env.NUANCE_TTS_URI
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
+807 -216
View File
File diff suppressed because it is too large Load Diff