mirror of
https://github.com/jambonz/speech-utils.git
synced 2026-10-03 23:33:59 +00:00
Compare commits
36
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d89457f67d | ||
|
|
3c15976669 | ||
|
|
70abb04941 | ||
|
|
724af7fcab | ||
|
|
c0a79090c2 | ||
|
|
7a9a3e2468 | ||
|
|
ab2b3f7fb2 | ||
|
|
e848d0276e | ||
|
|
d2447d477d | ||
|
|
ec8bbacc2c | ||
|
|
9694714a7f | ||
|
|
6a6de9b7a1 | ||
|
|
04d1fb548b | ||
|
|
bbf0167b40 | ||
|
|
4cfa286242 | ||
|
|
42ac63adc6 | ||
|
|
12a121672d | ||
|
|
7b94a5a969 | ||
|
|
c47b4883c7 | ||
|
|
7d076bb8b4 | ||
|
|
142323a151 | ||
|
|
d176a644fe | ||
|
|
644f2918dc | ||
|
|
4f430b9785 | ||
|
|
18c658d20c | ||
|
|
062608cf13 | ||
|
|
a2ea94c14a | ||
|
|
04bb85ef81 | ||
|
|
c123f19898 | ||
|
|
305695d068 | ||
|
|
1477752d40 | ||
|
|
fbedfe947f | ||
|
|
ab88facd52 | ||
|
|
a2be64da89 | ||
|
|
c5e0f256e6 | ||
|
|
d0378751ba |
@@ -13,11 +13,6 @@ jobs:
|
||||
with:
|
||||
node-version: '20'
|
||||
- run: npm install
|
||||
- name: Install Docker Compose
|
||||
run: |
|
||||
sudo curl -L "https://github.com/docker/compose/releases/download/1.29.2/docker-compose-$(uname -s)-$(uname -m)" -o /usr/local/bin/docker-compose
|
||||
sudo chmod +x /usr/local/bin/docker-compose
|
||||
docker-compose --version
|
||||
- run: npm run jslint
|
||||
- run: sudo apt update && sudo apt install -y squid
|
||||
- run: sudo cp test/squid.conf /etc/squid/squid.conf
|
||||
@@ -28,9 +23,7 @@ jobs:
|
||||
AWS_REGION: ${{ secrets.AWS_REGION }}
|
||||
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
|
||||
GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }}
|
||||
IBM_API_KEY: ${{ secrets.IBM_API_KEY }}
|
||||
IBM_TTS_API_KEY: ${{ secrets.IBM_TTS_API_KEY }}
|
||||
IBM_TTS_REGION: ${{ secrets.IBM_TTS_REGION }}
|
||||
|
||||
MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }}
|
||||
MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }}
|
||||
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
||||
|
||||
@@ -12,8 +12,8 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v3
|
||||
- uses: actions/setup-node@v3
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: lts/*
|
||||
registry-url: 'https://registry.npmjs.org'
|
||||
|
||||
@@ -0,0 +1,84 @@
|
||||
# CLAUDE.md
|
||||
|
||||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||
|
||||
## Overview
|
||||
|
||||
`@jambonz/speech-utils` is a Node.js library providing TTS (Text-to-Speech) utilities for the jambonz CPaaS platform. It handles speech synthesis with caching through Redis and supports multiple TTS vendors.
|
||||
|
||||
## Commands
|
||||
|
||||
```bash
|
||||
# Run tests (requires Docker for Redis)
|
||||
npm test
|
||||
|
||||
# Run linter
|
||||
npm run jslint
|
||||
|
||||
# Auto-fix lint issues
|
||||
npm run jslint:fix
|
||||
|
||||
# Generate coverage report
|
||||
npm run coverage
|
||||
```
|
||||
|
||||
## Testing
|
||||
|
||||
Tests use `tape` and require Redis. The test harness automatically starts/stops Redis via Docker Compose (`test/docker-compose-testbed.yaml`).
|
||||
|
||||
Most tests are conditional based on environment variables for vendor credentials:
|
||||
- `GCP_FILE` or `GCP_JSON_KEY` - Google TTS
|
||||
- `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_REGION` - AWS Polly
|
||||
- `MICROSOFT_API_KEY`, `MICROSOFT_REGION` - Azure TTS
|
||||
- `ELEVENLABS_API_KEY`, `ELEVENLABS_VOICE_ID`, `ELEVENLABS_MODEL_ID` - ElevenLabs
|
||||
- `OPENAI_API_KEY` - OpenAI Whisper TTS
|
||||
- And others per vendor
|
||||
|
||||
Redis config is in `config/test.json` (port 3379).
|
||||
|
||||
## Architecture
|
||||
|
||||
### Entry Point
|
||||
|
||||
`index.js` exports a factory function that takes Redis options and a logger, returning an object with these methods:
|
||||
- `synthAudio` - Main synthesis function
|
||||
- `getTtsVoices` - List available voices for a vendor
|
||||
- `purgeTtsCache` / `getTtsSize` / `addFileToCache` - Cache management
|
||||
- `getAwsAuthToken` - Token management
|
||||
|
||||
### Core Module: `lib/synth-audio.js`
|
||||
|
||||
The `synthAudio` function handles synthesis for all vendors. Key behaviors:
|
||||
1. **Cache check**: Generates SHA1 hash key from (vendor, language, voice, engine, model, text, instructions)
|
||||
2. **Streaming vs non-streaming**: When `JAMBONES_DISABLE_TTS_STREAMING` is not set and `renderForCaching=false`, returns `say:{params}text` format for FreeSWITCH streaming playback instead of generating files
|
||||
3. **Vendor dispatch**: Switch statement routes to vendor-specific synth functions (`synthGoogle`, `synthPolly`, `synthMicrosoft`, etc.)
|
||||
4. **Caching**: Stores audio as base64 JSON in Redis with configurable TTL (default 4 hours)
|
||||
|
||||
### Supported Vendors
|
||||
|
||||
google, aws/polly, microsoft/azure, nvidia (Riva), wellsaid, elevenlabs, cartesia, inworld, rimelabs, whisper (OpenAI), deepgram, resemble, custom:*
|
||||
|
||||
### gRPC Stubs
|
||||
|
||||
`stubs/riva/` contains generated protobuf/gRPC code for NVIDIA Riva.
|
||||
|
||||
## Environment Variables
|
||||
|
||||
Key configuration via env vars (see `lib/config.js`):
|
||||
- `JAMBONES_DISABLE_TTS_STREAMING` - Force non-streaming mode
|
||||
- `JAMBONES_DISABLE_AZURE_TTS_STREAMING` - Azure-specific streaming disable
|
||||
- `JAMBONES_TTS_CACHE_DURATION_MINS` - Cache TTL in minutes (default: 240)
|
||||
- `JAMBONES_TTS_TRIM_SILENCE` - Trim trailing silence from audio
|
||||
- `JAMBONES_TMP_FOLDER` - Temp folder for audio files (default: /tmp)
|
||||
- `JAMBONES_HTTP_PROXY_IP`, `JAMBONES_HTTP_PROXY_PORT` - HTTP proxy for Azure
|
||||
- `JAMBONES_AZURE_ENABLE_SSML` - Force SSML wrapper for Azure plain text
|
||||
|
||||
## Key Dependencies
|
||||
|
||||
- `@jambonz/realtimedb-helpers` - Redis client and hash utilities
|
||||
- `@google-cloud/text-to-speech` - Google TTS
|
||||
- `@aws-sdk/client-polly` - AWS Polly
|
||||
- `microsoft-cognitiveservices-speech-sdk` - Azure TTS
|
||||
- `@grpc/grpc-js` - gRPC for Riva
|
||||
- `openai` - OpenAI Whisper TTS
|
||||
- `bent` - HTTP client for REST-based vendors
|
||||
@@ -1,11 +0,0 @@
|
||||
#!/bin/sh
|
||||
|
||||
mkdir -p stubs/nuance
|
||||
|
||||
for FILE in ./protos/nuance/*; do
|
||||
grpc_tools_node_protoc \
|
||||
--js_out=import_style=commonjs,binary:./stubs/nuance \
|
||||
--grpc_out=grpc_js:./stubs/nuance \
|
||||
--proto_path=./protos/nuance \
|
||||
$FILE
|
||||
done
|
||||
@@ -14,9 +14,7 @@ module.exports = (opts, logger) => {
|
||||
purgeTtsCache: require('./lib/purge-tts-cache').bind(null, client, logger),
|
||||
addFileToCache: require('./lib/add-file-to-cache').bind(null, client, logger),
|
||||
synthAudio: require('./lib/synth-audio').bind(null, client, createHash, retrieveHash, logger),
|
||||
getVerbioAccessToken: require('./lib/get-verbio-token').bind(null, client, logger),
|
||||
getNuanceAccessToken: require('./lib/get-nuance-access-token').bind(null, client, logger),
|
||||
getIbmAccessToken: require('./lib/get-ibm-access-token').bind(null, client, logger),
|
||||
|
||||
getAwsAuthToken: require('./lib/get-aws-sts-token').bind(null, logger, createHash, retrieveHash),
|
||||
getTtsVoices: require('./lib/get-tts-voices').bind(null, client, createHash, retrieveHash, logger),
|
||||
};
|
||||
|
||||
@@ -1,48 +0,0 @@
|
||||
const formurlencoded = require('form-urlencoded');
|
||||
const {Pool} = require('undici');
|
||||
const pool = new Pool('https://iam.cloud.ibm.com');
|
||||
const {makeIbmKey, noopLogger} = require('./utils');
|
||||
const { HTTP_TIMEOUT } = require('./config');
|
||||
const debug = require('debug')('jambonz:realtimedb-helpers');
|
||||
|
||||
async function getIbmAccessToken(client, logger, apiKey) {
|
||||
logger = logger || noopLogger;
|
||||
try {
|
||||
const key = makeIbmKey(apiKey);
|
||||
const access_token = await client.get(key);
|
||||
if (access_token) return {access_token, servedFromCache: true};
|
||||
|
||||
/* access token not found in cache, so fetch it from Ibm */
|
||||
const payload = {
|
||||
grant_type: 'urn:ibm:params:oauth:grant-type:apikey',
|
||||
apikey: apiKey
|
||||
};
|
||||
const {statusCode, headers, body} = await pool.request({
|
||||
path: '/identity/token',
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'Content-Type': 'application/x-www-form-urlencoded'
|
||||
},
|
||||
body: formurlencoded(payload),
|
||||
timeout: HTTP_TIMEOUT,
|
||||
followRedirects: false
|
||||
});
|
||||
|
||||
if (200 !== statusCode) {
|
||||
const json = await body.json();
|
||||
logger.debug({statusCode, headers, body: json}, 'error fetching access token from Ibm');
|
||||
const err = new Error();
|
||||
err.statusCode = statusCode;
|
||||
throw err;
|
||||
}
|
||||
const json = await body.json();
|
||||
await client.set(key, json.access_token, 'EX', json.expires_in - 30);
|
||||
return {...json, servedFromCache: false};
|
||||
} catch (err) {
|
||||
debug(err, 'getIbmAccessToken: Error retrieving Ibm access token');
|
||||
logger.error(err, 'getIbmAccessToken: Error retrieving Ibm access token for client_id ${clientId}');
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = getIbmAccessToken;
|
||||
@@ -1,49 +0,0 @@
|
||||
const formurlencoded = require('form-urlencoded');
|
||||
const {Pool} = require('undici');
|
||||
const pool = new Pool('https://auth.crt.nuance.com');
|
||||
const {makeNuanceKey, makeBasicAuthHeader, noopLogger} = require('./utils');
|
||||
const { HTTP_TIMEOUT } = require('./config');
|
||||
const debug = require('debug')('jambonz:realtimedb-helpers');
|
||||
|
||||
async function getNuanceAccessToken(client, logger, clientId, secret, scope) {
|
||||
logger = logger || noopLogger;
|
||||
try {
|
||||
const key = makeNuanceKey(clientId, secret, scope);
|
||||
const access_token = await client.get(key);
|
||||
if (access_token) return {access_token, servedFromCache: true};
|
||||
|
||||
/* access token not found in cache, so fetch it from Nuance */
|
||||
const payload = {
|
||||
grant_type: 'client_credentials',
|
||||
scope
|
||||
};
|
||||
const auth = makeBasicAuthHeader(clientId, secret);
|
||||
const {statusCode, headers, body} = await pool.request({
|
||||
path: '/oauth2/token',
|
||||
method: 'POST',
|
||||
headers: {
|
||||
...auth,
|
||||
'Content-Type': 'application/x-www-form-urlencoded'
|
||||
},
|
||||
body: formurlencoded(payload),
|
||||
timeout: HTTP_TIMEOUT,
|
||||
followRedirects: false
|
||||
});
|
||||
|
||||
if (200 !== statusCode) {
|
||||
logger.debug({statusCode, headers, body: body.text()}, 'error fetching access token from Nuance');
|
||||
const err = new Error();
|
||||
err.statusCode = statusCode;
|
||||
throw err;
|
||||
}
|
||||
const json = await body.json();
|
||||
await client.set(key, json.access_token, 'EX', json.expires_in - 30);
|
||||
return {...json, servedFromCache: false};
|
||||
} catch (err) {
|
||||
debug(err, `getNuanceAccessToken: Error retrieving Nuance access token for client_id ${clientId}`);
|
||||
logger.error(err, `getNuanceAccessToken: Error retrieving Nuance access token for client_id ${clientId}`);
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = getNuanceAccessToken;
|
||||
+2
-111
@@ -1,91 +1,8 @@
|
||||
const assert = require('assert');
|
||||
const {noopLogger, createNuanceClient, createKryptonClient} = require('./utils');
|
||||
const getNuanceAccessToken = require('./get-nuance-access-token');
|
||||
const getVerbioAccessToken = require('./get-verbio-token');
|
||||
const {GetVoicesRequest, Voice} = require('../stubs/nuance/synthesizer_pb');
|
||||
const TextToSpeechV1 = require('ibm-watson/text-to-speech/v1');
|
||||
const { IamAuthenticator } = require('ibm-watson/auth');
|
||||
const {noopLogger} = require('./utils');
|
||||
const ttsGoogle = require('@google-cloud/text-to-speech');
|
||||
const { PollyClient, DescribeVoicesCommand } = require('@aws-sdk/client-polly');
|
||||
const getAwsAuthToken = require('./get-aws-sts-token');
|
||||
const {Pool} = require('undici');
|
||||
const { HTTP_TIMEOUT } = require('./config');
|
||||
const verbioVoicePool = new Pool('https://us.rest.speechcenter.verbio.com');
|
||||
|
||||
const getIbmVoices = async(client, logger, credentials) => {
|
||||
const {tts_region, tts_api_key} = credentials;
|
||||
console.log(`region: ${tts_region}, api_key: ${tts_api_key}`);
|
||||
|
||||
const textToSpeech = new TextToSpeechV1({
|
||||
authenticator: new IamAuthenticator({
|
||||
apikey: tts_api_key,
|
||||
}),
|
||||
serviceUrl: `https://api.${tts_region}.text-to-speech.watson.cloud.ibm.com`
|
||||
});
|
||||
|
||||
const voices = await textToSpeech.listVoices();
|
||||
return voices;
|
||||
};
|
||||
|
||||
const getNuanceVoices = async(client, logger, credentials) => {
|
||||
const {client_id: clientId, secret: secret, nuance_tts_uri} = credentials;
|
||||
|
||||
return new Promise(async(resolve, reject) => {
|
||||
/* get a nuance access token */
|
||||
let token, nuanceClient;
|
||||
try {
|
||||
if (nuance_tts_uri) {
|
||||
nuanceClient = await createKryptonClient(nuance_tts_uri);
|
||||
}
|
||||
else {
|
||||
const access_token = await getNuanceAccessToken(client, logger, clientId, secret, 'tts');
|
||||
token = access_token.access_token;
|
||||
nuanceClient = await createNuanceClient(token);
|
||||
}
|
||||
} catch (err) {
|
||||
logger.error({err}, 'getTtsVoices: error retrieving access token');
|
||||
return reject(err);
|
||||
}
|
||||
/* retrieve all voices */
|
||||
const v = new Voice();
|
||||
const request = new GetVoicesRequest();
|
||||
request.setVoice(v);
|
||||
|
||||
nuanceClient.getVoices(request, (err, response) => {
|
||||
if (err) {
|
||||
logger.error({err, clientId, secret, token}, 'getTtsVoices: error retrieving voices');
|
||||
return reject(err);
|
||||
}
|
||||
|
||||
/* return all the voices that are not restricted and eliminate duplicates */
|
||||
const voices = response.getVoicesList()
|
||||
.map((v) => {
|
||||
return {
|
||||
language: v.getLanguage(),
|
||||
name: v.getName(),
|
||||
model: v.getModel(),
|
||||
gender: v.getGender() === 1 ? 'male' : 'female',
|
||||
restricted: v.getRestricted()
|
||||
};
|
||||
});
|
||||
const v = voices
|
||||
.filter((v) => v.restricted === false)
|
||||
.map((v) => {
|
||||
delete v.restricted;
|
||||
return v;
|
||||
})
|
||||
.sort((a, b) => {
|
||||
if (a.language < b.language) return -1;
|
||||
if (a.language > b.language) return 1;
|
||||
if (a.name < b.name) return -1;
|
||||
return 1;
|
||||
});
|
||||
const arr = [...new Set(v.map((v) => JSON.stringify(v)))]
|
||||
.map((v) => JSON.parse(v));
|
||||
resolve(arr);
|
||||
});
|
||||
});
|
||||
};
|
||||
|
||||
const getGoogleVoices = async(_client, logger, credentials) => {
|
||||
const client = new ttsGoogle.TextToSpeechClient({credentials});
|
||||
@@ -126,26 +43,6 @@ const getAwsVoices = async(_client, createHash, retrieveHash, logger, credential
|
||||
}
|
||||
};
|
||||
|
||||
const getVerbioVoices = async(client, logger, credentials) => {
|
||||
try {
|
||||
const access_token = await getVerbioAccessToken(client, logger, credentials);
|
||||
const { body} = await verbioVoicePool.request({
|
||||
path: '/api/v1/voices',
|
||||
method: 'GET',
|
||||
headers: {
|
||||
'Authorization': `Bearer ${access_token.access_token}`,
|
||||
'User-Agent': 'jambonz'
|
||||
},
|
||||
timeout: HTTP_TIMEOUT,
|
||||
followRedirects: false
|
||||
});
|
||||
return await body.json();
|
||||
} catch (err) {
|
||||
logger.info({err}, 'getVerbioVoices - failed to list voices for Verbio');
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
/**
|
||||
* Synthesize speech to an mp3 file, and also cache the generated speech
|
||||
* in redis (base64 format) for 24 hours so as to avoid unnecessarily paying
|
||||
@@ -165,21 +62,15 @@ const getVerbioVoices = async(client, logger, credentials) => {
|
||||
async function getTtsVoices(client, createHash, retrieveHash, logger, {vendor, credentials}) {
|
||||
logger = logger || noopLogger;
|
||||
|
||||
assert.ok(['nuance', 'ibm', 'google', 'aws', 'polly', 'verbio'].includes(vendor),
|
||||
assert.ok(['google', 'aws', 'polly'].includes(vendor),
|
||||
`getTtsVoices not supported for vendor ${vendor}`);
|
||||
|
||||
switch (vendor) {
|
||||
case 'nuance':
|
||||
return getNuanceVoices(client, logger, credentials);
|
||||
case 'ibm':
|
||||
return getIbmVoices(client, logger, credentials);
|
||||
case 'google':
|
||||
return getGoogleVoices(client, logger, credentials);
|
||||
case 'aws':
|
||||
case 'polly':
|
||||
return getAwsVoices(client, createHash, retrieveHash, logger, credentials);
|
||||
case 'verbio':
|
||||
return getVerbioVoices(client, logger, credentials);
|
||||
default:
|
||||
break;
|
||||
}
|
||||
|
||||
@@ -1,51 +0,0 @@
|
||||
const {Pool} = require('undici');
|
||||
const { noopLogger, makeVerbioKey } = require('./utils');
|
||||
const { HTTP_TIMEOUT } = require('./config');
|
||||
const pool = new Pool('https://auth.speechcenter.verbio.com:444');
|
||||
const debug = require('debug')('jambonz:realtimedb-helpers');
|
||||
|
||||
async function getVerbioAccessToken(client, logger, credentials) {
|
||||
logger = logger || noopLogger;
|
||||
const { client_id, client_secret } = credentials;
|
||||
try {
|
||||
const key = makeVerbioKey(client_id);
|
||||
const access_token = await client.get(key);
|
||||
if (access_token) {
|
||||
return {access_token, servedFromCache: true};
|
||||
}
|
||||
|
||||
const payload = {
|
||||
client_id,
|
||||
client_secret
|
||||
};
|
||||
|
||||
const {statusCode, headers, body} = await pool.request({
|
||||
path: '/api/v1/token',
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'Content-Type': 'application/json',
|
||||
'User-Agent': 'jambonz'
|
||||
},
|
||||
body: JSON.stringify(payload),
|
||||
timeout: HTTP_TIMEOUT,
|
||||
followRedirects: false
|
||||
});
|
||||
|
||||
if (200 !== statusCode) {
|
||||
logger.debug({statusCode, headers, body: await body.text()}, 'error fetching access token from Verbio');
|
||||
const err = new Error();
|
||||
err.statusCode = statusCode;
|
||||
throw err;
|
||||
}
|
||||
const json = await body.json();
|
||||
const expiry = Math.floor(json.expiration_time - Date.now() / 1000 - 30);
|
||||
await client.set(key, json.access_token, 'EX', expiry);
|
||||
return {...json, servedFromCache: false};
|
||||
} catch (err) {
|
||||
debug(err, `getVerbioAccessToken: Error retrieving Verbio access token for client_id ${client_id}`);
|
||||
logger.error(err, `getVerbioAccessToken: Error retrieving Verbio access token for client_id ${client_id}`);
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = getVerbioAccessToken;
|
||||
+399
-285
@@ -6,8 +6,7 @@ const { PollyClient, SynthesizeSpeechCommand } = require('@aws-sdk/client-polly'
|
||||
const { CartesiaClient } = require('@cartesia/cartesia-js');
|
||||
|
||||
const sdk = require('microsoft-cognitiveservices-speech-sdk');
|
||||
const TextToSpeechV1 = require('ibm-watson/text-to-speech/v1');
|
||||
const { IamAuthenticator } = require('ibm-watson/auth');
|
||||
|
||||
const {
|
||||
ResultReason,
|
||||
SpeechConfig,
|
||||
@@ -17,26 +16,10 @@ const {
|
||||
} = sdk;
|
||||
const {
|
||||
makeSynthKey,
|
||||
createNuanceClient,
|
||||
createKryptonClient,
|
||||
createRivaClient,
|
||||
noopLogger,
|
||||
makeFilePath,
|
||||
makePlayhtKey
|
||||
makeFilePath
|
||||
} = require('./utils');
|
||||
const getNuanceAccessToken = require('./get-nuance-access-token');
|
||||
const getVerbioAccessToken = require('./get-verbio-token');
|
||||
const {
|
||||
SynthesisRequest,
|
||||
Voice,
|
||||
AudioFormat,
|
||||
AudioParameters,
|
||||
PCM,
|
||||
Input,
|
||||
Text,
|
||||
SSML,
|
||||
EventParameters
|
||||
} = require('../stubs/nuance/synthesizer_pb');
|
||||
const {SynthesizeSpeechRequest} = require('../stubs/riva/proto/riva_tts_pb');
|
||||
const {AudioEncoding} = require('../stubs/riva/proto/riva_audio_pb');
|
||||
const debug = require('debug')('jambonz:realtimedb-helpers');
|
||||
@@ -96,10 +79,12 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
let rtt;
|
||||
logger = logger || noopLogger;
|
||||
|
||||
assert.ok(['google', 'aws', 'polly', 'microsoft', 'wellsaid', 'nuance', 'nvidia', 'ibm', 'elevenlabs',
|
||||
'whisper', 'deepgram', 'playht', 'rimelabs', 'verbio', 'cartesia', 'inworld', 'resemble'].includes(vendor) ||
|
||||
assert.ok(['google', 'aws', 'polly', 'microsoft', 'wellsaid', 'nvidia', 'elevenlabs',
|
||||
'whisper', 'deepgram', 'deepgramflux', 'rimelabs', 'cartesia', 'gradium', 'nineninesix', 'inworld', 'resemble',
|
||||
'murf', 'xai']
|
||||
.includes(vendor) ||
|
||||
vendor.startsWith('custom'),
|
||||
`synthAudio supported vendors are google, aws, microsoft, nuance, nvidia and wellsaid ..etc, not ${vendor}`);
|
||||
`synthAudio supported vendors are google, aws, microsoft, nvidia and wellsaid ..etc, not ${vendor}`);
|
||||
if ('google' === vendor) {
|
||||
assert.ok(language, 'synthAudio requires language when google is used');
|
||||
}
|
||||
@@ -110,23 +95,11 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
assert.ok(language || deploymentId, 'synthAudio requires language when microsoft is used');
|
||||
assert.ok(voice || deploymentId, 'synthAudio requires voice when microsoft is used');
|
||||
}
|
||||
else if ('nuance' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when nuance is used');
|
||||
if (!credentials.nuance_tts_uri) {
|
||||
assert.ok(credentials.client_id, 'synthAudio requires client_id in credentials when nuance is used');
|
||||
assert.ok(credentials.secret, 'synthAudio requires client_id in credentials when nuance is used');
|
||||
}
|
||||
}
|
||||
else if ('nvidia' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when nvidia is used');
|
||||
assert.ok(language, 'synthAudio requires language when nvidia is used');
|
||||
assert.ok(credentials.riva_server_uri, 'synthAudio requires riva_server_uri in credentials when nvidia is used');
|
||||
}
|
||||
else if ('ibm' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when ibm is used');
|
||||
assert.ok(credentials.tts_region, 'synthAudio requires tts_region in credentials when ibm watson is used');
|
||||
assert.ok(credentials.tts_api_key, 'synthAudio requires tts_api_key in credentials when nuance is used');
|
||||
}
|
||||
else if ('wellsaid' === vendor) {
|
||||
language = 'en-US'; // WellSaid only supports English atm
|
||||
assert.ok(voice, 'synthAudio requires voice when wellsaid is used');
|
||||
@@ -135,11 +108,6 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
assert.ok(voice, 'synthAudio requires voice when elevenlabs is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when elevenlabs is used');
|
||||
assert.ok(credentials.model_id, 'synthAudio requires model_id when elevenlabs is used');
|
||||
} else if ('playht' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when playht is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when playht is used');
|
||||
assert.ok(credentials.user_id, 'synthAudio requires user_id when playht is used');
|
||||
assert.ok(credentials.voice_engine, 'synthAudio requires voice_engine when playht is used');
|
||||
} else if ('inworld' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when inworld is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when inworld is used');
|
||||
@@ -154,17 +122,30 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when whisper is used');
|
||||
} else if (vendor.startsWith('custom')) {
|
||||
assert.ok(credentials.custom_tts_url, `synthAudio requires custom_tts_url in credentials when ${vendor} is used`);
|
||||
} else if ('verbio' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when verbio is used');
|
||||
assert.ok(credentials.client_id, 'synthAudio requires client_id when verbio is used');
|
||||
assert.ok(credentials.client_secret, 'synthAudio requires client_secret when verbio is used');
|
||||
} else if ('deepgram' === vendor) {
|
||||
if (!credentials.deepgram_tts_uri) {
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when deepgram is used');
|
||||
}
|
||||
} else if ('deepgramflux' === vendor) {
|
||||
// Deepgram Flux TTS (/v2/speak); the flux model rides on `voice`/`model`
|
||||
if (!credentials.deepgram_tts_uri) {
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when deepgramflux is used');
|
||||
}
|
||||
} else if ('xai' === vendor) {
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when xai is used');
|
||||
} else if ('cartesia' === vendor) {
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when cartesia is used');
|
||||
assert.ok(credentials.model_id, 'synthAudio requires model_id when cartesia is used');
|
||||
} else if ('gradium' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when gradium is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when gradium is used');
|
||||
} else if ('nineninesix' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when nineninesix is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when nineninesix is used');
|
||||
assert.ok(credentials.model_id, 'synthAudio requires model_id when nineninesix is used');
|
||||
} else if ('murf' === vendor) {
|
||||
assert.ok(voice, 'synthAudio requires voice when murf is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when murf is used');
|
||||
} else if (vendor === 'resemble') {
|
||||
assert.ok(voice, 'synthAudio requires voice when resemble is used');
|
||||
assert.ok(credentials.api_key, 'synthAudio requires api_key when resemble is used');
|
||||
@@ -222,17 +203,10 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
audioData = await synthMicrosoft(logger, {credentials, stats, language, voice, key, text, deploymentId,
|
||||
renderForCaching, disableTtsStreaming, disableTtsCache});
|
||||
break;
|
||||
case 'nuance':
|
||||
model = model || 'enhanced';
|
||||
audioData = await synthNuance(client, logger, {credentials, stats, voice, model, key, text});
|
||||
break;
|
||||
case 'nvidia':
|
||||
audioData = await synthNvidia(client, logger, {credentials, stats, language, voice, model, key, text,
|
||||
renderForCaching, disableTtsStreaming, disableTtsCache});
|
||||
break;
|
||||
case 'ibm':
|
||||
audioData = await synthIbm(logger, {credentials, stats, voice, key, text});
|
||||
break;
|
||||
case 'wellsaid':
|
||||
audioData = await synthWellSaid(logger, {credentials, stats, language, voice, key, text});
|
||||
break;
|
||||
@@ -241,16 +215,21 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'playht':
|
||||
audioData = await synthPlayHT(client, logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'cartesia':
|
||||
audioData = await synthCartesia(logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'gradium':
|
||||
audioData = await synthGradium(logger, {
|
||||
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'nineninesix':
|
||||
audioData = await synthNineninesix(logger, {
|
||||
credentials, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'inworld':
|
||||
audioData = await synthInworld(logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
@@ -261,20 +240,29 @@ async function synthAudio(client, createHash, retrieveHash, logger, stats, { acc
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'murf':
|
||||
audioData = await synthMurf(logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'whisper':
|
||||
audioData = await synthWhisper(logger, {
|
||||
credentials, stats, voice, key, text, instructions, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'verbio':
|
||||
audioData = await synthVerbio(client, logger, {
|
||||
credentials, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache});
|
||||
if (audioData?.filePath) return audioData;
|
||||
break;
|
||||
case 'deepgram':
|
||||
audioData = await synthDeepgram(logger, {credentials, stats, model, key, text,
|
||||
renderForCaching, disableTtsStreaming, disableTtsCache});
|
||||
break;
|
||||
case 'deepgramflux':
|
||||
audioData = await synthDeepgramFlux(logger, {credentials, stats, model: model || voice, key, text,
|
||||
renderForCaching, disableTtsStreaming, disableTtsCache});
|
||||
break;
|
||||
case 'xai':
|
||||
audioData = await synthXai(logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming,
|
||||
disableTtsCache});
|
||||
break;
|
||||
case 'resemble':
|
||||
audioData = await synthResemble(logger, {
|
||||
credentials, stats, voice, key, text, options, renderForCaching, disableTtsStreaming, disableTtsCache});
|
||||
@@ -325,6 +313,7 @@ const synthPolly = async(createHash, retrieveHash, logger,
|
||||
params += `,playback_id=${key}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
params += ',vendor=aws';
|
||||
params += `,voice=${voice}`;
|
||||
if (accessKeyId && secretAccessKey) {
|
||||
if (accessKeyId) params += `,accessKeyId=${accessKeyId}`;
|
||||
if (secretAccessKey) params += `,secretAccessKey=${secretAccessKey}`;
|
||||
@@ -412,6 +401,32 @@ const synthPolly = async(createHash, retrieveHash, logger,
|
||||
}
|
||||
};
|
||||
|
||||
/* google AudioConfig settings we support, as [google camelCase name, freeswitch param name] */
|
||||
const GOOGLE_AUDIO_SETTINGS = [
|
||||
['speakingRate', 'speaking_rate'],
|
||||
['pitch', 'pitch'],
|
||||
['volumeGainDb', 'volume_gain_db']
|
||||
];
|
||||
|
||||
/**
|
||||
* Extract google AudioConfig settings from the synthesizer options. They may be supplied
|
||||
* nested under an audioConfig property (mirroring google's AudioConfig object) or flat at
|
||||
* the top level, and either google's camelCase or snake_case names are accepted.
|
||||
* @see https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig
|
||||
* @returns object keyed by google's camelCase names, holding only valid numeric settings
|
||||
*/
|
||||
const googleAudioConfig = (options) => {
|
||||
const provided = {...options, ...(options?.audioConfig || {})};
|
||||
const audioConfig = {};
|
||||
for (const [name, snakeName] of GOOGLE_AUDIO_SETTINGS) {
|
||||
const value = provided[name] ?? provided[snakeName];
|
||||
/* note: 0 is meaningful for pitch and volumeGainDb, so check for absence explicitly */
|
||||
if (value === undefined || value === null || value === '') continue;
|
||||
const num = Number(value);
|
||||
if (Number.isFinite(num)) audioConfig[name] = num;
|
||||
}
|
||||
return audioConfig;
|
||||
};
|
||||
|
||||
const synthGoogle = async(logger, {
|
||||
credentials, stats, language, voice, gender, key, text, model, options, instructions,
|
||||
@@ -439,15 +454,23 @@ const synthGoogle = async(logger, {
|
||||
params += `,voice=${voice}`;
|
||||
params += `,language_code=${language || 'en-US'}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
const useLiveApi = options?.useLiveApi ?? isHDVoice;
|
||||
const useGeminiTts = options?.useGeminiTts ?? isGemini;
|
||||
params += `,use_live_api=${useLiveApi ? 1 : 0}`;
|
||||
params += `,use_gemini_tts=${useGeminiTts ? 1 : 0}`;
|
||||
// api_mode: tts (standard), live (HD voices), gemini (Gemini TTS)
|
||||
const apiMode = options?.apiMode || (isGemini ? 'gemini' : (isHDVoice ? 'live' : 'tts'));
|
||||
params += `,api_mode=${apiMode}`;
|
||||
if (model) params += `,model_name=${model}`;
|
||||
if (gender) params += `,gender=${gender}`;
|
||||
// comma is used to separate parameters in freeswitch tts module
|
||||
const prompt = options?.prompt || instructions;
|
||||
if (prompt) params += `,prompt=${prompt.replace(/\n/g, ' ').replace(/,/g, ';')}`;
|
||||
/**
|
||||
* AudioConfig settings. Note google only honors these in some api modes:
|
||||
* tts applies all of them, live (HD voices) applies speakingRate only,
|
||||
* and gemini ignores them entirely (use prompt instead for style control).
|
||||
*/
|
||||
const audioSettings = googleAudioConfig(options);
|
||||
for (const [name, snakeName] of GOOGLE_AUDIO_SETTINGS) {
|
||||
if (name in audioSettings) params += `,${snakeName}=${audioSettings[name]}`;
|
||||
}
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
@@ -516,6 +539,9 @@ const synthGoogle = async(logger, {
|
||||
sampleRate = 8000;
|
||||
}
|
||||
|
||||
/* gemini voices do not support the AudioConfig settings; they use prompt for style control */
|
||||
if (!isGemini) Object.assign(audioConfig, googleAudioConfig(options));
|
||||
|
||||
const opts = { input, voice: voiceParams, audioConfig };
|
||||
|
||||
try {
|
||||
@@ -536,39 +562,6 @@ const synthGoogle = async(logger, {
|
||||
}
|
||||
};
|
||||
|
||||
const synthIbm = async(logger, {credentials, stats, voice, text}) => {
|
||||
const {tts_api_key, tts_region} = credentials;
|
||||
const params = {
|
||||
text,
|
||||
voice,
|
||||
accept: 'audio/mp3'
|
||||
};
|
||||
|
||||
try {
|
||||
const textToSpeech = new TextToSpeechV1({
|
||||
authenticator: new IamAuthenticator({
|
||||
apikey: tts_api_key,
|
||||
}),
|
||||
serviceUrl: `https://api.${tts_region}.text-to-speech.watson.cloud.ibm.com`
|
||||
});
|
||||
|
||||
const r = await textToSpeech.synthesize(params);
|
||||
const chunks = [];
|
||||
for await (const chunk of r.result) {
|
||||
chunks.push(chunk);
|
||||
}
|
||||
return {
|
||||
audioContent: Buffer.concat(chunks),
|
||||
extension: 'mp3',
|
||||
sampleRate: 8000
|
||||
};
|
||||
} catch (err) {
|
||||
logger.info({err, params}, 'synthAudio: Error synthesizing speech using ibm');
|
||||
stats.increment('tts.count', ['vendor:ibm', 'accepted:no']);
|
||||
throw new Error(err.statusText || err.message);
|
||||
}
|
||||
};
|
||||
|
||||
async function _synthOnPremMicrosoft(logger, {
|
||||
credentials,
|
||||
language,
|
||||
@@ -767,70 +760,6 @@ const synthWellSaid = async(logger, {credentials, stats, language, voice, gender
|
||||
}
|
||||
};
|
||||
|
||||
const synthNuance = async(client, logger, {credentials, stats, voice, model, text}) => {
|
||||
let nuanceClient;
|
||||
const {client_id, secret, nuance_tts_uri} = credentials;
|
||||
if (nuance_tts_uri) {
|
||||
nuanceClient = await createKryptonClient(nuance_tts_uri);
|
||||
}
|
||||
else {
|
||||
/* get a nuance access token */
|
||||
const {access_token} = await getNuanceAccessToken(client, logger, client_id, secret, 'tts');
|
||||
nuanceClient = await createNuanceClient(access_token);
|
||||
}
|
||||
|
||||
const v = new Voice();
|
||||
const p = new AudioParameters();
|
||||
const f = new AudioFormat();
|
||||
const pcm = new PCM();
|
||||
const params = new EventParameters();
|
||||
const request = new SynthesisRequest();
|
||||
const input = new Input();
|
||||
|
||||
if (text.startsWith('<speak')) {
|
||||
const ssml = new SSML();
|
||||
ssml.setText(text);
|
||||
input.setSsml(ssml);
|
||||
}
|
||||
else {
|
||||
const t = new Text();
|
||||
t.setText(text);
|
||||
input.setText(t);
|
||||
}
|
||||
const sampleRate = 8000;
|
||||
pcm.setSampleRateHz(sampleRate);
|
||||
f.setPcm(pcm);
|
||||
p.setAudioFormat(f);
|
||||
v.setName(voice);
|
||||
v.setModel(model);
|
||||
request.setVoice(v);
|
||||
request.setAudioParams(p);
|
||||
request.setInput(input);
|
||||
request.setEventParams(params);
|
||||
request.setUserId('jambonz');
|
||||
|
||||
return new Promise((resolve, reject) => {
|
||||
nuanceClient.unarySynthesize(request, (err, response) => {
|
||||
if (err) {
|
||||
console.error(err);
|
||||
return reject(err);
|
||||
}
|
||||
const status = response.getStatus();
|
||||
const code = status.getCode();
|
||||
if (code !== 200) {
|
||||
const message = status.getMessage();
|
||||
const details = status.getDetails();
|
||||
return reject({code, message, details});
|
||||
}
|
||||
resolve({
|
||||
audioContent: Buffer.from(response.getAudio()),
|
||||
extension: 'r8',
|
||||
sampleRate
|
||||
});
|
||||
});
|
||||
});
|
||||
};
|
||||
|
||||
const synthNvidia = async(client, logger, {
|
||||
credentials, stats, language, voice, model, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
@@ -970,7 +899,7 @@ const synthElevenlabs = async(logger, {
|
||||
const optimize_streaming_latency = opts.optimize_streaming_latency ?
|
||||
`?optimize_streaming_latency=${opts.optimize_streaming_latency}` : '';
|
||||
try {
|
||||
const post = bent(`https://${api_uri}`, 'POST', 'buffer', {
|
||||
const post = bent(`https://${api_uri || 'api.elevenlabs.io'}`, 'POST', 'buffer', {
|
||||
'xi-api-key': api_key,
|
||||
'Accept': 'audio/mpeg',
|
||||
'Content-Type': 'application/json'
|
||||
@@ -996,101 +925,6 @@ const synthElevenlabs = async(logger, {
|
||||
}
|
||||
};
|
||||
|
||||
const synthPlayHT = async(client, logger, {
|
||||
credentials, options, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
const {api_key, user_id, voice_engine, playht_tts_uri, options: credOpts} = credentials;
|
||||
const opts = !!options && Object.keys(options).length !== 0 ? options : JSON.parse(credOpts || '{}');
|
||||
|
||||
let synthesizeUrl = playht_tts_uri ? `${playht_tts_uri}/api/v2/tts/stream` : 'https://api.play.ht/api/v2/tts/stream';
|
||||
// If model is play3.0, the synthesizeUrl is got from authentication endpoint
|
||||
if (voice_engine === 'Play3.0') {
|
||||
try {
|
||||
const post = bent('https://api.play.ht', 'POST', 'json', 201, {
|
||||
'AUTHORIZATION': api_key,
|
||||
'X-USER-ID': user_id,
|
||||
'Accept': 'application/json'
|
||||
});
|
||||
const key = makePlayhtKey(api_key);
|
||||
const url = await client.get(key);
|
||||
if (!url) {
|
||||
const {inference_address, expires_at_ms} = await post('/api/v3/auth');
|
||||
synthesizeUrl = inference_address;
|
||||
const expiry = Math.floor((expires_at_ms - Date.now()) / 1000 - 30);
|
||||
await client.set(key, inference_address, 'EX', expiry);
|
||||
} else {
|
||||
// Use cached URL
|
||||
synthesizeUrl = url;
|
||||
}
|
||||
} catch (err) {
|
||||
logger.info({err}, 'synth PlayHT returned error for authentication version 3.0');
|
||||
stats.increment('tts.count', ['vendor:playht', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
|
||||
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
|
||||
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
let params = '{';
|
||||
params += `api_key=${api_key}`;
|
||||
params += `,playback_id=${key}`;
|
||||
params += `,user_id=${user_id}`;
|
||||
params += ',vendor=playht';
|
||||
params += `,voice=${voice}`;
|
||||
params += `,voice_engine=${voice_engine}`;
|
||||
params += `,synthesize_url=${synthesizeUrl}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
params += `,language=${language}`;
|
||||
if (opts.quality) params += `,quality=${opts.quality}`;
|
||||
if (opts.speed) params += `,speed=${opts.speed}`;
|
||||
if (opts.seed) params += `,style=${opts.seed}`;
|
||||
if (opts.temperature) params += `,temperature=${opts.temperature}`;
|
||||
if (opts.emotion) params += `,emotion=${opts.emotion}`;
|
||||
if (opts.voice_guidance) params += `,voice_guidance=${opts.voice_guidance}`;
|
||||
if (opts.style_guidance) params += `,style_guidance=${opts.style_guidance}`;
|
||||
if (opts.text_guidance) params += `,text_guidance=${opts.text_guidance}`;
|
||||
if (opts.top_p) params += `,top_p=${opts.top_p}`;
|
||||
if (opts.repetition_penalty) params += `,repetition_penalty=${opts.repetition_penalty}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
|
||||
servedFromCache: false,
|
||||
rtt: 0
|
||||
};
|
||||
}
|
||||
|
||||
try {
|
||||
const post = bent('POST', 'buffer', {
|
||||
...(voice_engine !== 'Play3.0' && {
|
||||
'AUTHORIZATION': api_key,
|
||||
'X-USER-ID': user_id,
|
||||
}),
|
||||
'Accept': 'audio/mpeg',
|
||||
'Content-Type': 'application/json'
|
||||
});
|
||||
|
||||
const audioContent = await post(synthesizeUrl, {
|
||||
text,
|
||||
...(voice_engine === 'Play3.0' && { language }),
|
||||
voice,
|
||||
voice_engine,
|
||||
output_format: 'mp3',
|
||||
sample_rate: 8000,
|
||||
...opts
|
||||
});
|
||||
return {
|
||||
audioContent,
|
||||
extension: 'mp3',
|
||||
sampleRate: 8000
|
||||
};
|
||||
} catch (err) {
|
||||
logger.info({err}, 'synth PlayHT returned error');
|
||||
stats.increment('tts.count', ['vendor:playht', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
const synthInworld = async(logger, {
|
||||
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
@@ -1179,6 +1013,8 @@ const synthRimelabs = async(logger, {
|
||||
if (opts.repetition_penalty) params += `,repetition_penalty=${opts.repetition_penalty}`;
|
||||
if (opts.top_p) params += `,top_p=${opts.top_p}`;
|
||||
if (opts.max_tokens) params += `,max_tokens=${opts.max_tokens}`;
|
||||
// Coda model parameters
|
||||
if (opts.timeScaleFactor) params += `,time_scale_factor=${opts.timeScaleFactor}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
@@ -1214,54 +1050,72 @@ const synthRimelabs = async(logger, {
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
const synthVerbio = async(client, logger, {
|
||||
credentials, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
|
||||
const synthMurf = async(logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
//https://doc.speechcenter.verbio.com/#tag/Text-To-Speech-REST-API
|
||||
if (text.length > 2000) {
|
||||
throw new Error('Verbio cannot synthesize for the text length larger than 2000 characters');
|
||||
}
|
||||
const token = await getVerbioAccessToken(client, logger, credentials);
|
||||
if (!process.env.JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
const {api_key, model_id, api_uri, options: credOpts} = credentials;
|
||||
const opts = !!options && Object.keys(options).length !== 0 ? options : JSON.parse(credOpts || '{}');
|
||||
|
||||
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
|
||||
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
/* param keys here must match mod_murf_tts's text_param handler */
|
||||
let params = '{';
|
||||
params += `access_token=${token.access_token}`;
|
||||
params += `api_key=${api_key}`;
|
||||
params += `,playback_id=${key}`;
|
||||
params += ',vendor=verbio';
|
||||
params += ',vendor=murf';
|
||||
params += `,voice=${voice}`;
|
||||
if (model_id) params += `,model_id=${model_id}`;
|
||||
if (language) params += `,language=${language}`;
|
||||
if (api_uri) params += `,api_uri=${api_uri}`;
|
||||
if (opts.style) params += `,style=${opts.style}`;
|
||||
if (opts.rate !== undefined && opts.rate !== null) params += `,rate=${opts.rate}`;
|
||||
if (opts.pitch !== undefined && opts.pitch !== null) params += `,pitch=${opts.pitch}`;
|
||||
if (opts.variation !== undefined && opts.variation !== null) params += `,variation=${opts.variation}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ')}`,
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
|
||||
servedFromCache: false,
|
||||
rtt: 0
|
||||
};
|
||||
}
|
||||
|
||||
try {
|
||||
const post = bent('https://us.rest.speechcenter.verbio.com', 'POST', 'buffer', {
|
||||
'Authorization': `Bearer ${token.access_token}`,
|
||||
'User-Agent': 'jambonz',
|
||||
const sampleRate = 8000;
|
||||
/* no Accept header: murf returns 406 if it doesn't match; the response
|
||||
container is selected by the `format` field in the body instead */
|
||||
const post = bent(api_uri || 'https://global.api.murf.ai', 'POST', 'buffer', {
|
||||
'api-key': api_key,
|
||||
'Content-Type': 'application/json'
|
||||
});
|
||||
const audioContent = await post('/api/v1/synthesize', {
|
||||
/* murf REST schema is documented loosely; field names follow the SDK params
|
||||
(voice_id/model/format/sample_rate) plus the websocket voice fields. */
|
||||
const audioContent = await post('/v1/speech/stream', {
|
||||
text,
|
||||
voice_id: voice,
|
||||
output_sample_rate: '8k',
|
||||
output_encoding: 'pcm16',
|
||||
text
|
||||
...(model_id && {model: model_id}),
|
||||
...(language && {locale: language}),
|
||||
...(opts.style && {style: opts.style}),
|
||||
...(opts.rate !== undefined && opts.rate !== null && {rate: opts.rate}),
|
||||
...(opts.pitch !== undefined && opts.pitch !== null && {pitch: opts.pitch}),
|
||||
...(opts.variation !== undefined && opts.variation !== null && {variation: opts.variation}),
|
||||
format: 'WAV',
|
||||
sample_rate: sampleRate,
|
||||
channel_type: 'MONO'
|
||||
});
|
||||
return {
|
||||
audioContent,
|
||||
extension: 'r8',
|
||||
sampleRate: 8000
|
||||
extension: 'wav',
|
||||
sampleRate
|
||||
};
|
||||
} catch (err) {
|
||||
logger.info({err}, 'synth Verbio returned error');
|
||||
stats.increment('tts.count', ['vendor:verbio', 'accepted:no']);
|
||||
logger.info({err}, 'synth murf returned error');
|
||||
stats.increment('tts.count', ['vendor:murf', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
const synthWhisper = async(logger, {credentials, stats, voice, key, text, instructions,
|
||||
renderForCaching, disableTtsStreaming, disableTtsCache}) => {
|
||||
const {api_key, model_id, baseURL, timeout, speed} = credentials;
|
||||
@@ -1352,6 +1206,116 @@ const synthDeepgram = async(logger, {credentials, stats, model, key, text, rende
|
||||
}
|
||||
};
|
||||
|
||||
// Deepgram Flux TTS — the conversation-native model served from /v2/speak.
|
||||
// Streaming rides the mediajam deepgramflux dialect via a say: filePath; the
|
||||
// batch/cache path POSTs to /v2/speak (mp3 is batch-only for Flux).
|
||||
const synthDeepgramFlux = async(logger, {credentials, stats, model, key, text, renderForCaching,
|
||||
disableTtsStreaming, disableTtsCache}) => {
|
||||
const {api_key, deepgram_tts_uri} = credentials;
|
||||
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
let params = '{';
|
||||
params += `api_key=${api_key}`;
|
||||
params += `,playback_id=${key}`;
|
||||
params += ',vendor=deepgramflux';
|
||||
params += `,voice=${model}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
if (deepgram_tts_uri) params += `,endpoint=${deepgram_tts_uri}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ')}`,
|
||||
servedFromCache: false,
|
||||
rtt: 0
|
||||
};
|
||||
}
|
||||
try {
|
||||
const post = bent(deepgram_tts_uri || 'https://api.deepgram.com', 'POST', 'buffer', {
|
||||
// on-premise deepgram does not require to have api_key
|
||||
...(api_key && {'Authorization': `Token ${api_key}`}),
|
||||
'Accept': 'audio/mpeg',
|
||||
'Content-Type': 'application/json'
|
||||
});
|
||||
const audioContent = await post(`/v2/speak?model=${model}`, {
|
||||
text
|
||||
});
|
||||
return {
|
||||
audioContent,
|
||||
extension: 'mp3',
|
||||
sampleRate: 8000
|
||||
};
|
||||
} catch (err) {
|
||||
logger.info({err}, 'synth Deepgram Flux returned error');
|
||||
stats.increment('tts.count', ['vendor:deepgramflux', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
const synthXai = async(logger, {
|
||||
credentials, options, stats, language, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
const {api_key, api_uri, options: credOpts} = credentials;
|
||||
const opts = !!options && Object.keys(options).length !== 0 ? options : JSON.parse(credOpts || '{}');
|
||||
const speed = opts.speed;
|
||||
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
let params = '{';
|
||||
params += `api_key=${api_key}`;
|
||||
params += `,playback_id=${key}`;
|
||||
params += ',vendor=xai';
|
||||
if (voice) params += `,voice=${voice}`;
|
||||
if (language) params += `,language=${language}`;
|
||||
if (speed !== null && speed !== undefined) params += `,speed=${speed}`;
|
||||
if (opts.optimize_streaming_latency != null) {
|
||||
params += `,optimize_streaming_latency=${opts.optimize_streaming_latency}`;
|
||||
}
|
||||
if (opts.text_normalization != null) params += `,text_normalization=${opts.text_normalization}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
if (api_uri) params += `,endpoint=${api_uri}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
|
||||
servedFromCache: false,
|
||||
rtt: 0
|
||||
};
|
||||
}
|
||||
try {
|
||||
const post = bent(`https://${api_uri || 'api.x.ai'}`, 'POST', 'buffer', {
|
||||
'Authorization': `Bearer ${api_key}`,
|
||||
'Content-Type': 'application/json'
|
||||
});
|
||||
const audioContent = await post('/v1/tts', {
|
||||
text,
|
||||
language: language || 'auto',
|
||||
...(voice && {voice_id: voice}),
|
||||
...(speed !== null && speed !== undefined && {speed}),
|
||||
...(opts.optimize_streaming_latency != null && {optimize_streaming_latency: opts.optimize_streaming_latency}),
|
||||
...(opts.text_normalization != null && {text_normalization: opts.text_normalization}),
|
||||
output_format: {
|
||||
codec: 'wav',
|
||||
sample_rate: 8000
|
||||
}
|
||||
});
|
||||
return {
|
||||
audioContent,
|
||||
extension: 'wav',
|
||||
sampleRate: 8000
|
||||
};
|
||||
} catch (err) {
|
||||
// xAI errors are JSON {code, error} - read the body so the surfaced message isn't 'undefined'
|
||||
if (err.name === 'StatusError' && typeof err.text === 'function') {
|
||||
try {
|
||||
const body = await err.text();
|
||||
if (body) err.message = body;
|
||||
} catch (readErr) {
|
||||
logger.info({readErr}, 'synth xai: failed to read error response body');
|
||||
}
|
||||
}
|
||||
logger.info({err}, 'synth xai returned error');
|
||||
stats.increment('tts.count', ['vendor:xai', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
const synthCartesia = async(logger, {
|
||||
credentials, options, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
@@ -1384,6 +1348,18 @@ const synthCartesia = async(logger, {
|
||||
try {
|
||||
const client = new CartesiaClient({ apiKey: api_key });
|
||||
const sampleRate = 48000;
|
||||
|
||||
// omit a control unless explicitly provided (0 is a valid value, so test for nullish only)
|
||||
const has = (v) => v !== null && v !== undefined;
|
||||
/* Voice controls are model-family specific:
|
||||
- sonic-2 takes `experimentalControls` (emotion is an array, no volume).
|
||||
- sonic-3 family (sonic-3, sonic-3.5, pinned sonic-3.x snapshots) takes
|
||||
`generationConfig` (emotion is a string, volume supported). Match the same
|
||||
"starts with sonic-3" predicate the freeswitch streaming module uses
|
||||
(mod_cartesia_tts_streaming: strncmp(model_id, "sonic-3", ...)), so cached
|
||||
and streamed synthesis behave identically.
|
||||
Older models (sonic, sonic-english, sonic-multilingual, sonic-2024-*) take neither. */
|
||||
const isSonic3 = model_id?.startsWith('sonic-3');
|
||||
const mp3Stream = await client.tts.bytes({
|
||||
modelId: model_id,
|
||||
transcript: text,
|
||||
@@ -1399,16 +1375,16 @@ const synthCartesia = async(logger, {
|
||||
),
|
||||
...(model_id === 'sonic-2' && (opts.speed || opts.emotion) && {
|
||||
experimentalControls: {
|
||||
...(opts.speed !== null && opts.speed !== undefined && {speed: opts.speed}),
|
||||
...(has(opts.speed) && {speed: opts.speed}),
|
||||
...(opts.emotion && {emotion: [opts.emotion]}),
|
||||
}
|
||||
}),
|
||||
},
|
||||
...(model_id === 'sonic-3' && (opts.speed || opts.emotion || opts.volume) && {
|
||||
...(isSonic3 && (has(opts.speed) || has(opts.emotion) || has(opts.volume)) && {
|
||||
generationConfig: {
|
||||
...(opts.volume !== null && opts.volume !== undefined && {volume: opts.volume}),
|
||||
...(opts.speed !== null && opts.speed !== undefined && {speed: opts.speed}),
|
||||
...(opts.emotion !== null && opts.emotion !== undefined && {emotion: opts.emotion}),
|
||||
...(has(opts.volume) && {volume: opts.volume}),
|
||||
...(has(opts.speed) && {speed: opts.speed}),
|
||||
...(has(opts.emotion) && {emotion: opts.emotion}),
|
||||
}
|
||||
}),
|
||||
language: language,
|
||||
@@ -1432,6 +1408,32 @@ const synthCartesia = async(logger, {
|
||||
sampleRate
|
||||
};
|
||||
} catch (err) {
|
||||
/* Cartesia's tts.bytes() uses a streaming response, so on an HTTP error the SDK
|
||||
throws a CartesiaError whose `body` is an unconsumed stream-wrapper object
|
||||
(async-iterable, yielding Uint8Array chunks) rather than the parsed error
|
||||
JSON. Read it so callers get a meaningful message instead of a serialized
|
||||
stream object. */
|
||||
if (err && err.body && typeof err.body !== 'string' && typeof err.body[Symbol.asyncIterator] === 'function') {
|
||||
try {
|
||||
const chunks = [];
|
||||
for await (const chunk of err.body) {
|
||||
chunks.push(Buffer.isBuffer(chunk) ? chunk : Buffer.from(chunk));
|
||||
}
|
||||
const text = Buffer.concat(chunks).toString('utf8');
|
||||
if (text) {
|
||||
let parsed;
|
||||
try {
|
||||
parsed = JSON.parse(text);
|
||||
} catch {
|
||||
parsed = null;
|
||||
}
|
||||
err.message = (parsed && (parsed.error || parsed.message)) || text;
|
||||
err.body = parsed || text;
|
||||
}
|
||||
} catch (readErr) {
|
||||
logger.info({readErr}, 'synth Cartesia: failed to read error response body');
|
||||
}
|
||||
}
|
||||
logger.info({err}, 'synth Cartesia returned error');
|
||||
stats.increment('tts.count', ['vendor:cartesia', 'accepted:no']);
|
||||
throw err;
|
||||
@@ -1439,6 +1441,118 @@ const synthCartesia = async(logger, {
|
||||
|
||||
};
|
||||
|
||||
/* gradium.ai — json websocket for streaming, and a POST endpoint for the cache
|
||||
render. only_audio:true + pcm_8000 returns bare little-endian 16-bit samples,
|
||||
which is exactly the r8 container, so we avoid gradium's streaming wav header
|
||||
(it carries 0xffffffff as the RIFF size, since length is unknown up front).
|
||||
*/
|
||||
const synthGradium = async(logger, {
|
||||
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
const {api_key, model_id} = credentials;
|
||||
const {json_config, pronunciation_id} = options || {};
|
||||
|
||||
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
|
||||
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
let params = '{';
|
||||
params += `api_key=${api_key}`;
|
||||
params += `,playback_id=${key}`;
|
||||
params += ',vendor=gradium';
|
||||
params += `,voice=${voice}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
if (model_id) params += `,model_id=${model_id}`;
|
||||
if (pronunciation_id) params += `,pronunciation_id=${pronunciation_id}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
|
||||
servedFromCache: false,
|
||||
rtt: 0
|
||||
};
|
||||
}
|
||||
|
||||
try {
|
||||
const sampleRate = 8000;
|
||||
const post = bent('https://api.gradium.ai', 'POST', 'buffer', {
|
||||
'x-api-key': api_key,
|
||||
'Content-Type': 'application/json'
|
||||
});
|
||||
const audioContent = await post('/api/post/speech/tts', {
|
||||
text,
|
||||
voice_id: voice,
|
||||
model_name: model_id || 'default',
|
||||
output_format: `pcm_${sampleRate}`,
|
||||
only_audio: true,
|
||||
...(json_config && {json_config}),
|
||||
...(pronunciation_id && {pronunciation_id})
|
||||
});
|
||||
return {
|
||||
audioContent,
|
||||
extension: 'r8',
|
||||
sampleRate
|
||||
};
|
||||
} catch (err) {
|
||||
logger.info({err}, 'synth gradium returned error');
|
||||
stats.increment('tts.count', ['vendor:gradium', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
/* nineninesix.ai — a Cartesia-compatible API, but only raw/wav come back
|
||||
(mp3 is rejected), so the cache render asks for wav rather than mp3. */
|
||||
const synthNineninesix = async(logger, {
|
||||
credentials, stats, voice, language, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
const {api_key, model_id} = credentials;
|
||||
|
||||
/* default to using the streaming interface, unless disabled by env var OR we want just a cache file */
|
||||
if (!JAMBONES_DISABLE_TTS_STREAMING && !renderForCaching && !disableTtsStreaming) {
|
||||
let params = '{';
|
||||
params += `api_key=${api_key}`;
|
||||
params += `,playback_id=${key}`;
|
||||
params += `,model_id=${model_id}`;
|
||||
params += ',vendor=nineninesix';
|
||||
params += `,voice=${voice}`;
|
||||
params += `,write_cache_file=${disableTtsCache ? 0 : 1}`;
|
||||
if (language) params += `,language=${language}`;
|
||||
params += '}';
|
||||
|
||||
return {
|
||||
filePath: `say:${params}${text.replace(/\n/g, ' ').replace(/\r/g, ' ')}`,
|
||||
servedFromCache: false,
|
||||
rtt: 0
|
||||
};
|
||||
}
|
||||
|
||||
try {
|
||||
const sampleRate = 8000;
|
||||
const post = bent('https://api.nineninesix.ai', 'POST', 'buffer', {
|
||||
'Authorization': `Bearer ${api_key}`,
|
||||
'Content-Type': 'application/json'
|
||||
});
|
||||
const audioContent = await post('/tts/bytes', {
|
||||
model_id,
|
||||
transcript: text,
|
||||
voice: {mode: 'id', id: voice},
|
||||
...(language && {language}),
|
||||
output_format: {
|
||||
container: 'wav',
|
||||
encoding: 'pcm_s16le',
|
||||
sample_rate: sampleRate
|
||||
}
|
||||
});
|
||||
return {
|
||||
audioContent,
|
||||
extension: 'wav',
|
||||
sampleRate
|
||||
};
|
||||
} catch (err) {
|
||||
logger.info({err}, 'synth nineninesix returned error');
|
||||
stats.increment('tts.count', ['vendor:nineninesix', 'accepted:no']);
|
||||
throw err;
|
||||
}
|
||||
};
|
||||
|
||||
const synthResemble = async(logger, {
|
||||
credentials, options, stats, voice, key, text, renderForCaching, disableTtsStreaming, disableTtsCache
|
||||
}) => {
|
||||
|
||||
+1
-107
@@ -1,20 +1,7 @@
|
||||
const crypto = require('crypto');
|
||||
const {SynthesizerClient} = require('../stubs/nuance/synthesizer_grpc_pb');
|
||||
const {RivaSpeechSynthesisClient} = require('../stubs/riva/proto/riva_tts_grpc_pb');
|
||||
const {Pool} = require('undici');
|
||||
const pool = new Pool('https://auth.crt.nuance.com');
|
||||
const NUANCE_AUTH_ENDPOINT = 'tts.api.nuance.com:443';
|
||||
const grpc = require('@grpc/grpc-js');
|
||||
const formurlencoded = require('form-urlencoded');
|
||||
const { TMP_FOLDER, HTTP_TIMEOUT } = require('./config');
|
||||
|
||||
const debug = require('debug')('jambonz:realtimedb-helpers');
|
||||
/**
|
||||
* Future TODO: cache recently used connections to providers
|
||||
* to avoid connection overhead during a call.
|
||||
* Will need to periodically age them out to avoid memory leaks.
|
||||
*/
|
||||
//const nuanceClientMap = new Map();
|
||||
const { TMP_FOLDER } = require('./config');
|
||||
|
||||
function makeSynthKey({
|
||||
account_sid = '',
|
||||
@@ -45,96 +32,12 @@ const noopLogger = {
|
||||
error: () => {}
|
||||
};
|
||||
|
||||
const toBase64 = (str) => Buffer.from(str || '', 'utf8').toString('base64');
|
||||
|
||||
function makeBasicAuthHeader(username, password) {
|
||||
if (!username || !password) return {};
|
||||
const creds = `${encodeURIComponent(username)}:${password || ''}`;
|
||||
const header = `Basic ${toBase64(creds)}`;
|
||||
return {Authorization: header};
|
||||
}
|
||||
|
||||
function makeIbmKey(apiKey) {
|
||||
const hash = crypto.createHash('sha1');
|
||||
hash.update(apiKey);
|
||||
return `ibm:${hash.digest('hex')}`;
|
||||
}
|
||||
|
||||
function makeAwsKey(awsAccessKeyId) {
|
||||
const hash = crypto.createHash('sha1');
|
||||
hash.update(awsAccessKeyId);
|
||||
return `aws:${hash.digest('hex')}`;
|
||||
}
|
||||
|
||||
function makePlayhtKey(apiKey) {
|
||||
const hash = crypto.createHash('sha1');
|
||||
hash.update(apiKey);
|
||||
return `playht:${hash.digest('hex')}`;
|
||||
}
|
||||
function makeVerbioKey(client_id) {
|
||||
const hash = crypto.createHash('sha1');
|
||||
hash.update(client_id);
|
||||
return `verbio:${hash.digest('hex')}`;
|
||||
}
|
||||
|
||||
function makeNuanceKey(clientId, secret, scope) {
|
||||
const hash = crypto.createHash('sha1');
|
||||
hash.update(`${clientId}:${secret}:${scope}`);
|
||||
return `nuance:${hash.digest('hex')}`;
|
||||
}
|
||||
|
||||
const getNuanceAccessToken = async(clientId, secret, scope = 'asr tts') => {
|
||||
const payload = {
|
||||
grant_type: 'client_credentials',
|
||||
scope
|
||||
};
|
||||
const auth = makeBasicAuthHeader(clientId, secret);
|
||||
const {statusCode, headers, body} = await pool.request({
|
||||
path: '/oauth2/token',
|
||||
method: 'POST',
|
||||
headers: {
|
||||
...auth,
|
||||
'Content-Type': 'application/x-www-form-urlencoded'
|
||||
},
|
||||
body: formurlencoded(payload),
|
||||
timeout: HTTP_TIMEOUT,
|
||||
followRedirects: false
|
||||
});
|
||||
|
||||
if (200 !== statusCode) {
|
||||
debug({statusCode, headers, body: body.text()}, 'error fetching access token from Nuance');
|
||||
const err = new Error();
|
||||
err.statusCode = statusCode;
|
||||
throw err;
|
||||
}
|
||||
const json = await body.json();
|
||||
return json.access_token;
|
||||
};
|
||||
|
||||
const createKryptonClient = async(uri) => {
|
||||
const client = new SynthesizerClient(uri, grpc.credentials.createInsecure());
|
||||
return client;
|
||||
};
|
||||
|
||||
const createNuanceClient = async(access_token) => {
|
||||
|
||||
//if (nuanceClientMap.has(access_token)) return nuanceClientMap.get(access_token);
|
||||
|
||||
const generateMetadata = (params, callback) => {
|
||||
var metadata = new grpc.Metadata();
|
||||
metadata.add('authorization', `Bearer ${access_token}`);
|
||||
callback(null, metadata);
|
||||
};
|
||||
|
||||
const sslCreds = grpc.credentials.createSsl();
|
||||
const authCreds = grpc.credentials.createFromMetadataGenerator(generateMetadata);
|
||||
const combined_creds = grpc.credentials.combineChannelCredentials(sslCreds, authCreds);
|
||||
const client = new SynthesizerClient(NUANCE_AUTH_ENDPOINT, combined_creds);
|
||||
|
||||
//if (process.env.NUANCE_CACHE_TTS_CONNECTIONS) nuanceClientMap.set(access_token, client);
|
||||
return client;
|
||||
};
|
||||
|
||||
const createRivaClient = async(rivaUri) => {
|
||||
const client = new RivaSpeechSynthesisClient(rivaUri, grpc.credentials.createInsecure());
|
||||
return client;
|
||||
@@ -142,17 +45,8 @@ const createRivaClient = async(rivaUri) => {
|
||||
|
||||
module.exports = {
|
||||
makeSynthKey,
|
||||
makeNuanceKey,
|
||||
makeIbmKey,
|
||||
makePlayhtKey,
|
||||
makeAwsKey,
|
||||
makeVerbioKey,
|
||||
getNuanceAccessToken,
|
||||
createNuanceClient,
|
||||
createKryptonClient,
|
||||
createRivaClient,
|
||||
makeBasicAuthHeader,
|
||||
NUANCE_AUTH_ENDPOINT,
|
||||
noopLogger,
|
||||
makeFilePath
|
||||
};
|
||||
|
||||
Generated
+1815
-2572
File diff suppressed because it is too large
Load Diff
+4
-9
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@jambonz/speech-utils",
|
||||
"version": "0.2.27",
|
||||
"version": "1.0.15",
|
||||
"description": "TTS-related speech utilities for jambonz",
|
||||
"main": "index.js",
|
||||
"author": "Dave Horton",
|
||||
@@ -13,7 +13,6 @@
|
||||
"coverage": "nyc --reporter html --report-dir ./coverage npm run test",
|
||||
"jslint": "eslint index.js lib",
|
||||
"jslint:fix": "npm run jslint --fix",
|
||||
"build": "./build_stubs.sh",
|
||||
"prepare": "husky"
|
||||
},
|
||||
"repository": {
|
||||
@@ -26,7 +25,6 @@
|
||||
},
|
||||
"homepage": "https://github.com/jambonz/speech-utils#readme",
|
||||
"dependencies": {
|
||||
"23": "^0.0.0",
|
||||
"@aws-sdk/client-polly": "^3.496.0",
|
||||
"@aws-sdk/client-sts": "^3.496.0",
|
||||
"@cartesia/cartesia-js": "^2.2.7",
|
||||
@@ -35,15 +33,12 @@
|
||||
"@jambonz/realtimedb-helpers": "^0.8.7",
|
||||
"bent": "^7.3.12",
|
||||
"debug": "^4.3.4",
|
||||
"form-urlencoded": "^6.1.4",
|
||||
"google-protobuf": "^3.21.2",
|
||||
"ibm-watson": "^11.0.0",
|
||||
"microsoft-cognitiveservices-speech-sdk": "1.38.0",
|
||||
"openai": "^4.98.0",
|
||||
"undici": "^7.5.0"
|
||||
"microsoft-cognitiveservices-speech-sdk": "^1.51.0",
|
||||
"openai": "^4.98.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"config": "^3.3.11",
|
||||
"config": "^4.2.0",
|
||||
"eslint": "^9.3.0",
|
||||
"eslint-plugin-promise": "^6.2.0",
|
||||
"husky": "^9.0.11",
|
||||
|
||||
@@ -1,332 +0,0 @@
|
||||
syntax = "proto3";
|
||||
|
||||
package nuance.tts.v1;
|
||||
|
||||
/*
|
||||
The Synthesizer service offers these functionalities:
|
||||
* - GetVoices: Queries the list of available voices, with filters to reduce the search space.
|
||||
* - Synthesize: Synthesizes audio from input text and parameters, and returns an audio stream.
|
||||
* - UnarySynthesize: Synthesizes audio from input text and parameters, and returns a single audio response.
|
||||
*/
|
||||
service Synthesizer {
|
||||
rpc GetVoices(GetVoicesRequest) returns (GetVoicesResponse) {}
|
||||
rpc Synthesize(SynthesisRequest) returns (stream SynthesisResponse) {}
|
||||
rpc UnarySynthesize(SynthesisRequest) returns (UnarySynthesisResponse) {}
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for [Synthesizer](#synthesizer) - GetVoices, to query voices available to the client.
|
||||
*/
|
||||
message GetVoicesRequest {
|
||||
Voice voice = 1; // Optionally filter the voices to retrieve, e.g. set language to en-US to return only American English voices.
|
||||
}
|
||||
|
||||
/*
|
||||
Output message for [Synthesizer](#synthesizer) - GetVoices. Includes a list of voices that matched the input criteria, if any.
|
||||
*/
|
||||
message GetVoicesResponse {
|
||||
repeated Voice voices = 1; // Repeated. Voices and characteristics returned.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for [Synthesizer](#synthesizer) - Synthesize. Specifies input text, audio parameters, and events to subscribe to, in exchange for synthesized audio.
|
||||
*/
|
||||
message SynthesisRequest {
|
||||
Voice voice = 1; // The voice to use for audio synthesis. Mandatory.
|
||||
AudioParameters audio_params = 2; // Output audio parameters, such as encoding and volume.
|
||||
Input input = 3; // Input text to synthesize, tuning data, etc. Mandatory.
|
||||
EventParameters event_params = 4; // Markers and other info to include in server events returned during synthesis.
|
||||
map<string, string> client_data = 5; // Repeated. Client-supplied key-value pairs to inject into the call log.
|
||||
string user_id = 6; // Identifies a particular user within an application.
|
||||
}
|
||||
|
||||
/*
|
||||
Input or output message for voices. When sent as input:
|
||||
* - In [GetVoicesRequest](#getvoicesrequest), it filters the list of available voices.
|
||||
* - In [SynthesisRequest](#synthesisrequest), it specifies the voice to use for synthesis.
|
||||
|
||||
When received as output in [GetVoicesResponse](#getvoicesresponse), it returns the list of available voices.
|
||||
*/
|
||||
message Voice {
|
||||
string name = 1; // The voice's name, e.g. 'Evan'. Mandatory for SynthesisRequest.
|
||||
string model = 2; // The voice's quality model, e.g. 'enhanced' or 'standard'. Mandatory for SynthesisRequest.
|
||||
string language = 3; // IETF language code, e.g. 'en-US'. Used only for GetVoicesRequest and GetVoicesResponse, to search for voices with a certain mother tongue. Ignored otherwise.
|
||||
EnumAgeGroup age_group = 4; // Used only for GetVoicesRequest and GetVoicesResponse, to search for adult or child voices. Ignored otherwise.
|
||||
EnumGender gender = 5; // Used only for GetVoicesRequest and GetVoicesResponse, to search for voices with a certain gender. Ignored otherwise.
|
||||
uint32 sample_rate_hz = 6; // Used only for GetVoicesRequest and GetVoicesResponse, to search for a certain native sample rate. Ignored otherwise.
|
||||
string language_tlw = 7; // Used only for GetVoicesRequest and GetVoicesResponse. Three-letter language code (e.g. 'enu' for American English), for configuring language identification in [Input](#input).
|
||||
bool restricted = 8; // Used only in GetVoicesResponse, to identify restricted voices. These are custom voices available only to specific customers. Ignored otherwise.
|
||||
string version = 9; // Used only in GetVoicesResponse, to return the voice's version. Ignored otherwise.
|
||||
}
|
||||
|
||||
/*
|
||||
Input or output field specifying whether the voice uses its adult or child version, if available. Included in [Voice](#voice).
|
||||
*/
|
||||
enum EnumAgeGroup {
|
||||
ADULT = 0; // Adult voice. Default for GetVoicesRequest.
|
||||
CHILD = 1; // Child voice.
|
||||
}
|
||||
|
||||
/*
|
||||
Input or output field, specifying gender for voices that support multiple genders. Included in [Voice](#voice).
|
||||
*/
|
||||
enum EnumGender {
|
||||
ANY = 0; // Any gender voice. Default for GetVoicesRequest.
|
||||
MALE = 1; // Male voice.
|
||||
FEMALE = 2; // Female voice.
|
||||
NEUTRAL = 3; // Neutral gender voice.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for audio-related parameters during synthesis, including encoding, volume, and audio length. Included in [SynthesisRequest](#synthesisrequest).
|
||||
*/
|
||||
message AudioParameters {
|
||||
AudioFormat audio_format = 1; // Audio encoding. Default PCM 22.5kHz.
|
||||
uint32 volume_percentage = 2; // Volume amplitude, from 0 to 100. Default 80.
|
||||
float speaking_rate_factor = 3; // Speaking rate, from 0 to 2.0. Default 1.0.
|
||||
uint32 audio_chunk_duration_ms = 4; // Maximum duration, in ms, of an audio chunk delivered to the client, from 1 to 60000. Default is 20000 (20 seconds). When this parameter is large enough (for example, 20 or 30 seconds), each audio chunk contains an audible segment surrounded by silence.
|
||||
uint32 target_audio_length_ms = 5; // Maximum duration, in ms, of synthesized audio. When greater than 0, the server stops ongoing synthesis at the first sentence end, or silence, closest to the value.
|
||||
bool disable_early_emission = 6; // By default, audio segments are emitted as soon as possible, even if they are not audible. This behavior may be disabled.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for audio encoding of synthesized text. Included in [AudioParameters](#audioparameters).
|
||||
*/
|
||||
message AudioFormat {
|
||||
oneof audio_format {
|
||||
PCM pcm = 1; // Signed 16-bit little endian PCM.
|
||||
ALaw alaw = 2; // G.711 A-law, 8kHz.
|
||||
ULaw ulaw = 3; // G.711 Mu-law, 8kHz.
|
||||
OggOpus ogg_opus = 4; // Ogg Opus, 8kHz, 16kHz or 24kHz.
|
||||
Opus opus = 5; // Opus, 8kHz, 16kHZ or 24 kHz. The audio will be sent one Opus packet at a time.
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
Input message defining PCM sample rate. Included in [Audioformat](#audioformat).
|
||||
*/
|
||||
message PCM {
|
||||
uint32 sample_rate_hz = 1; // Output sample rate: 8000, 16000, 22050 (default), 24000.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message defining A-law audio format. G.711 audio formats are set to 8kHz. Included in [Audioformat](#audioformat).
|
||||
*/
|
||||
message ALaw {}
|
||||
|
||||
/*
|
||||
Input message defining Mu-law audio format. G.711 audio formats are set to 8kHz. Included in [Audioformat](#audioformat).
|
||||
*/
|
||||
message ULaw {}
|
||||
|
||||
/*
|
||||
Input message defining Opus output rate. Included in [Audioformat](#audioformat).
|
||||
*/
|
||||
message Opus {
|
||||
uint32 sample_rate_hz = 1; // Output sample rate. Supported values: 8000, 16000, 24000 Hz.
|
||||
uint32 bit_rate_bps = 2; // Valid range is 500 to 256000 bps. Default 28000 bps.
|
||||
float max_frame_duration_ms = 3; // Opus frame size, in ms: 2.5, 5, 10, 20, 40, 60. Default 20.
|
||||
uint32 complexity = 4; // Computational complexity. A complexity of 0 means the codec default.
|
||||
EnumVariableBitrate vbr = 5; // Variable bitrate. On by default.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message defining Ogg Opus output rate. Included in [Audioformat](#audioformat).
|
||||
*/
|
||||
message OggOpus {
|
||||
uint32 sample_rate_hz = 1; // Output sample rate. Supported values: 8000, 16000, 24000 Hz.
|
||||
uint32 bit_rate_bps = 2; // Valid range is 500 to 256000 bps. Default 28000 bps.
|
||||
float max_frame_duration_ms = 3; // Opus frame size, in ms: 2.5, 5, 10, 20, 40, 60. Default 20.
|
||||
uint32 complexity = 4; // Computational complexity. A complexity of 0 means the codec default.
|
||||
EnumVariableBitrate vbr = 5; // Variable bitrate. On by default.
|
||||
}
|
||||
|
||||
/*
|
||||
Settings for variable bitrate. Included in [OggOpus](#oggopus). Turned on by default.
|
||||
*/
|
||||
enum EnumVariableBitrate {
|
||||
VARIABLE_BITRATE_ON = 0; // Use variable bitrate. Default.
|
||||
VARIABLE_BITRATE_OFF = 1; // Do not use variable bitrate.
|
||||
VARIABLE_BITRATE_CONSTRAINED = 2; // Use constrained variable bitrate.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message containing text to synthesize and synthesis parameters, including tuning data, etc. Included in [SynthesisRequest](#synthesisrequest). The type of input may be:
|
||||
* - Plain text.
|
||||
* - An SSML document.
|
||||
* - An alternating sequence of plain text and Nuance control codes.
|
||||
*/
|
||||
message Input {
|
||||
oneof input_data {
|
||||
Text text = 1; // Text input.
|
||||
SSML ssml = 2; // SSML input.
|
||||
TokenizedSequence tokenized_sequence = 3; // Sequence of text and Nuance control codes.
|
||||
}
|
||||
repeated SynthesisResource resources = 4; // Repeated. Synthesis resources (user dictionaries, rulesets, etc.) to tune synthesized audio.
|
||||
LanguageIdentificationParameters lid_params = 5; // LID parameters.
|
||||
DownloadParameters download_params = 6; // Remote file download parameters.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for synthesizing plain text. The encoding must be in UTF-8.
|
||||
*/
|
||||
message Text {
|
||||
oneof text_data {
|
||||
string text = 1; // Plain input text in UTF-8 encoding.
|
||||
string uri = 2; // Remote URI to the plain input text. Disabled for Nuance-hosted TTS.
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for synthesizing an SSML document.
|
||||
*/
|
||||
message SSML {
|
||||
oneof ssml_data {
|
||||
string text = 1; // SSML input text.
|
||||
string uri = 2; // Remote URI to the SSML input text. Disabled for Nuance-hosted TTS.
|
||||
}
|
||||
EnumSSMLValidationMode ssml_validation_mode = 3; // SSML validation mode. Default STRICT.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message for synthesizing a sequence of plain text and Nuance control codes.
|
||||
*/
|
||||
message TokenizedSequence {
|
||||
repeated Token tokens = 1;
|
||||
}
|
||||
|
||||
/*
|
||||
The unit when using TokenizedSequence for input. Each token can either be plain text or a Nuance control code.
|
||||
*/
|
||||
message Token {
|
||||
oneof token_data {
|
||||
string text = 1; // Plain input text.
|
||||
ControlCode control_code = 2; // Nuance control code.
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
A Nuance control code allows the user to control how text is spoken, similarly to SSML.
|
||||
*/
|
||||
message ControlCode {
|
||||
string key = 1; // Name of the control code, e.g. "pause".
|
||||
string value = 2; // Value of the control code.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message specifying the type of file to tune the synthesized output and its location or contents. Included in [Input](#input).
|
||||
*/
|
||||
message SynthesisResource {
|
||||
EnumResourceType type = 1; // Resource type, e.g. user dictionary, etc. Default USER_DICTIONARY.
|
||||
oneof resource_data {
|
||||
string uri = 2; // URI to the remote resource, or
|
||||
bytes body = 3; // For EnumResourceType USER_DICTIONARY, the contents of the file.
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
The type of synthesis resource to tune the output. Included in [SynthesisResource](#synthesisresource). User dictionaries provide custom pronunciations, rulesets apply search-and-replace rules to input text and ActivePrompt databases help tune synthesized audio under certain conditions, using Nuance Vocalizer Studio.
|
||||
*/
|
||||
enum EnumResourceType {
|
||||
USER_DICTIONARY = 0; // User dictionary (application/edct-bin-dictionary). Default.
|
||||
TEXT_USER_RULESET = 1; // Text user ruleset (application/x-vocalizer-rettt+text).
|
||||
BINARY_USER_RULESET = 2; // Binary user ruleset (application/x-vocalizer-rettt+bin).
|
||||
ACTIVEPROMPT_DB = 3; // ActivePrompt database (application/x-vocalizer/activeprompt-db).
|
||||
ACTIVEPROMPT_DB_AUTO = 4; // ActivePrompt database with automatic insertion (application/x-vocalizer/activeprompt-db;mode=automatic).
|
||||
SYSTEM_DICTIONARY = 5; // Nuance system dictionary (application/sdct-bin-dictionary).
|
||||
}
|
||||
|
||||
/*
|
||||
SSML validation mode when using SSML input. Included in [Input](#input). Strict by default but can be relaxed.
|
||||
*/
|
||||
enum EnumSSMLValidationMode {
|
||||
STRICT = 0; // Strict SSL validation. Default.
|
||||
WARN = 1; // Give warning only.
|
||||
NONE = 2; // Do not validate.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message controlling the language identifier. Included in [Input](#input). The language identifier runs on input blocks labeled with the <ESC>\lang=unknown\ control sequence or SSML xml:lang="unknown". The language identifier automatically restricts the matched languages to the installed voices. This limits the permissible languages, and also sets the order of precedence (first to last) when they have equal confidence scores.
|
||||
*/
|
||||
message LanguageIdentificationParameters {
|
||||
bool disable = 1; // Whether to disable language identification. Turned on by default.
|
||||
repeated string languages = 2; // Repeated. List of three-letter language codes (e.g. enu, frc, spm) to restrict language identification results, in order of precedence. Use GetVoicesRequest - Voice - language_tlw to obtain the three-letter codes. Default blank.
|
||||
bool always_use_highest_confidence = 3; // If enabled, language identification always chooses the language with the highest confidence score, even if the score is low. Default false, meaning use language with any confidence.
|
||||
}
|
||||
|
||||
/*
|
||||
Input message containing parameters for remote file download, whether for input text (Input.uri) or a SynthesisResource (SynthesisResource.uri). Included in [Input](#input).
|
||||
*/
|
||||
message DownloadParameters {
|
||||
map <string,string> headers = 1; // HTTP headers to include in outgoing requests. Only whitelisted headers will actually be sent.
|
||||
bool refuse_cookies = 2; // Whether to disable cookies. By default, HTTP requests accept cookies.
|
||||
oneof optional_download_parameter_request_timeout_ms {
|
||||
uint32 request_timeout_ms = 3; // Request timeout in ms. Default (0) means server default, usually 30000 (30 seconds).
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
Input message that defines event subscription parameters. Included in [SynthesisRequest](#synthesisrequest). Events that are requested are sent throughout the SynthesisResponse stream, when generated. Marker events can send events as certain parts of the synthesized audio are reached, for example, at the end of a word, sentence, or user-defined bookmark.
|
||||
|
||||
* Log events are produced throughout a synthesis request for events such as a voice loaded by the server or an audio chunk being ready to send.
|
||||
*/
|
||||
message EventParameters {
|
||||
bool send_sentence_marker_events = 1; // Sentence marker. Default: do not send.
|
||||
bool send_word_marker_events = 2; // Word marker. Default: do not send.
|
||||
bool send_phoneme_marker_events = 3; // Phoneme marker. Default: do not send.
|
||||
bool send_bookmark_marker_events = 4; // Bookmark marker. Default: do not send.
|
||||
bool send_paragraph_marker_events = 5; // Paragraph marker. Default: do not send.
|
||||
bool send_visemes = 6; // Lipsync information. Default: do not send.
|
||||
bool send_log_events = 7; // Whether to log events during synthesis. By default, logging is turned off.
|
||||
bool suppress_input = 8; // Whether to omit input text and URIs from log events. By default, these items are included.
|
||||
}
|
||||
|
||||
/*
|
||||
The [Synthesizer](#synthesizer) - Synthesize RPC call returns a stream of SynthesisResponse messages. The response contains one of:
|
||||
* - A status response, indicating completion or failure of the request. This is received only once and signifies the end of a Synthesize call.
|
||||
* - A list of events the client has requested. This can be received many times. See EventParameters for details.
|
||||
* - An audio buffer. This may be received many times.
|
||||
*/
|
||||
message SynthesisResponse {
|
||||
oneof response {
|
||||
Status status = 1; // A status response, indicating completion or failure of the request.
|
||||
Events events = 2; // A list of events. See EventParameters for details.
|
||||
bytes audio = 3; // The latest audio buffer.
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
The [Synthesizer](#synthesizer) - UnarySynthesize RPC call returns a single UnarySynthesisResponse message. It is similar to a SynthesisResponse message but includes all the information instead of a single type of response. The response contains:
|
||||
* - A status response, indicating completion or failure of the request.
|
||||
* - A list of events the client has requested. See EventParameters for details.
|
||||
* - The complete audio buffer of the synthesized text.
|
||||
*/
|
||||
message UnarySynthesisResponse {
|
||||
Status status = 1; // A status response, indicating completion or failure of the request.
|
||||
Events events = 2; // A list of events. See EventParameters for details.
|
||||
bytes audio = 3; // Audio buffer of the synthesized text.
|
||||
}
|
||||
|
||||
/*
|
||||
Output message containing a status response, indicating completion or failure of a SynthesisRequest. Included in [SynthesisResponse](#synthesisresponse) and [UnarySynthesisResponse](#unarysynthesisresponse).
|
||||
*/
|
||||
message Status {
|
||||
uint32 code = 1; // HTTP-style return code: 200, 4xx, or 5xx as appropriate.
|
||||
string message = 2; // Brief description of the status.
|
||||
string details = 3; // Longer description if available.
|
||||
}
|
||||
|
||||
/*
|
||||
Output message defining a container for a list of events. This container is needed because oneof does not allow repeated parameters in Protobuf. Included in [SynthesisResponse](#synthesisresponse) and [UnarySynthesisResponse](#unarysynthesisresponse).
|
||||
*/
|
||||
message Events {
|
||||
repeated Event events = 1; // Repeated. One or more events.
|
||||
}
|
||||
|
||||
/*
|
||||
Output message defining an event message. Included in [Events](#events). See EventParameters for details.
|
||||
*/
|
||||
message Event {
|
||||
string name = 1; // Either "Markers" or the name of the event in the case of a Log Event.
|
||||
map<string, string> values = 2; // Repeated. Key-value data relevant to the current event.
|
||||
}
|
||||
@@ -1,104 +0,0 @@
|
||||
// GENERATED CODE -- DO NOT EDIT!
|
||||
|
||||
'use strict';
|
||||
var grpc = require('@grpc/grpc-js');
|
||||
var synthesizer_pb = require('./synthesizer_pb.js');
|
||||
|
||||
function serialize_nuance_tts_v1_GetVoicesRequest(arg) {
|
||||
if (!(arg instanceof synthesizer_pb.GetVoicesRequest)) {
|
||||
throw new Error('Expected argument of type nuance.tts.v1.GetVoicesRequest');
|
||||
}
|
||||
return Buffer.from(arg.serializeBinary());
|
||||
}
|
||||
|
||||
function deserialize_nuance_tts_v1_GetVoicesRequest(buffer_arg) {
|
||||
return synthesizer_pb.GetVoicesRequest.deserializeBinary(new Uint8Array(buffer_arg));
|
||||
}
|
||||
|
||||
function serialize_nuance_tts_v1_GetVoicesResponse(arg) {
|
||||
if (!(arg instanceof synthesizer_pb.GetVoicesResponse)) {
|
||||
throw new Error('Expected argument of type nuance.tts.v1.GetVoicesResponse');
|
||||
}
|
||||
return Buffer.from(arg.serializeBinary());
|
||||
}
|
||||
|
||||
function deserialize_nuance_tts_v1_GetVoicesResponse(buffer_arg) {
|
||||
return synthesizer_pb.GetVoicesResponse.deserializeBinary(new Uint8Array(buffer_arg));
|
||||
}
|
||||
|
||||
function serialize_nuance_tts_v1_SynthesisRequest(arg) {
|
||||
if (!(arg instanceof synthesizer_pb.SynthesisRequest)) {
|
||||
throw new Error('Expected argument of type nuance.tts.v1.SynthesisRequest');
|
||||
}
|
||||
return Buffer.from(arg.serializeBinary());
|
||||
}
|
||||
|
||||
function deserialize_nuance_tts_v1_SynthesisRequest(buffer_arg) {
|
||||
return synthesizer_pb.SynthesisRequest.deserializeBinary(new Uint8Array(buffer_arg));
|
||||
}
|
||||
|
||||
function serialize_nuance_tts_v1_SynthesisResponse(arg) {
|
||||
if (!(arg instanceof synthesizer_pb.SynthesisResponse)) {
|
||||
throw new Error('Expected argument of type nuance.tts.v1.SynthesisResponse');
|
||||
}
|
||||
return Buffer.from(arg.serializeBinary());
|
||||
}
|
||||
|
||||
function deserialize_nuance_tts_v1_SynthesisResponse(buffer_arg) {
|
||||
return synthesizer_pb.SynthesisResponse.deserializeBinary(new Uint8Array(buffer_arg));
|
||||
}
|
||||
|
||||
function serialize_nuance_tts_v1_UnarySynthesisResponse(arg) {
|
||||
if (!(arg instanceof synthesizer_pb.UnarySynthesisResponse)) {
|
||||
throw new Error('Expected argument of type nuance.tts.v1.UnarySynthesisResponse');
|
||||
}
|
||||
return Buffer.from(arg.serializeBinary());
|
||||
}
|
||||
|
||||
function deserialize_nuance_tts_v1_UnarySynthesisResponse(buffer_arg) {
|
||||
return synthesizer_pb.UnarySynthesisResponse.deserializeBinary(new Uint8Array(buffer_arg));
|
||||
}
|
||||
|
||||
|
||||
//
|
||||
// The Synthesizer service offers these functionalities:
|
||||
// - GetVoices: Queries the list of available voices, with filters to reduce the search space.
|
||||
// - Synthesize: Synthesizes audio from input text and parameters, and returns an audio stream.
|
||||
// - UnarySynthesize: Synthesizes audio from input text and parameters, and returns a single audio response.
|
||||
var SynthesizerService = exports.SynthesizerService = {
|
||||
getVoices: {
|
||||
path: '/nuance.tts.v1.Synthesizer/GetVoices',
|
||||
requestStream: false,
|
||||
responseStream: false,
|
||||
requestType: synthesizer_pb.GetVoicesRequest,
|
||||
responseType: synthesizer_pb.GetVoicesResponse,
|
||||
requestSerialize: serialize_nuance_tts_v1_GetVoicesRequest,
|
||||
requestDeserialize: deserialize_nuance_tts_v1_GetVoicesRequest,
|
||||
responseSerialize: serialize_nuance_tts_v1_GetVoicesResponse,
|
||||
responseDeserialize: deserialize_nuance_tts_v1_GetVoicesResponse,
|
||||
},
|
||||
synthesize: {
|
||||
path: '/nuance.tts.v1.Synthesizer/Synthesize',
|
||||
requestStream: false,
|
||||
responseStream: true,
|
||||
requestType: synthesizer_pb.SynthesisRequest,
|
||||
responseType: synthesizer_pb.SynthesisResponse,
|
||||
requestSerialize: serialize_nuance_tts_v1_SynthesisRequest,
|
||||
requestDeserialize: deserialize_nuance_tts_v1_SynthesisRequest,
|
||||
responseSerialize: serialize_nuance_tts_v1_SynthesisResponse,
|
||||
responseDeserialize: deserialize_nuance_tts_v1_SynthesisResponse,
|
||||
},
|
||||
unarySynthesize: {
|
||||
path: '/nuance.tts.v1.Synthesizer/UnarySynthesize',
|
||||
requestStream: false,
|
||||
responseStream: false,
|
||||
requestType: synthesizer_pb.SynthesisRequest,
|
||||
responseType: synthesizer_pb.UnarySynthesisResponse,
|
||||
requestSerialize: serialize_nuance_tts_v1_SynthesisRequest,
|
||||
requestDeserialize: deserialize_nuance_tts_v1_SynthesisRequest,
|
||||
responseSerialize: serialize_nuance_tts_v1_UnarySynthesisResponse,
|
||||
responseDeserialize: deserialize_nuance_tts_v1_UnarySynthesisResponse,
|
||||
},
|
||||
};
|
||||
|
||||
exports.SynthesizerClient = grpc.makeGenericClientConstructor(SynthesizerService);
|
||||
File diff suppressed because it is too large
Load Diff
@@ -2,7 +2,7 @@ const test = require('tape').test ;
|
||||
const exec = require('child_process').exec ;
|
||||
|
||||
test('starting docker network..', (t) => {
|
||||
exec(`docker-compose -f ${__dirname}/docker-compose-testbed.yaml up -d`, (err, stdout, stderr) => {
|
||||
exec(`docker compose -f ${__dirname}/docker-compose-testbed.yaml up -d`, (err, stdout, stderr) => {
|
||||
setTimeout(() => {
|
||||
t.end(err);
|
||||
}, 2000);
|
||||
|
||||
+1
-1
@@ -3,7 +3,7 @@ const exec = require('child_process').exec ;
|
||||
|
||||
test('stopping docker network..', (t) => {
|
||||
t.timeoutAfter(10000);
|
||||
exec(`docker-compose -f ${__dirname}/docker-compose-testbed.yaml down`, (err, stdout, stderr) => {
|
||||
exec(`docker compose -f ${__dirname}/docker-compose-testbed.yaml down`, (err, stdout, stderr) => {
|
||||
//console.log(`stderr: ${stderr}`);
|
||||
process.exit(0);
|
||||
});
|
||||
|
||||
-78
@@ -1,78 +0,0 @@
|
||||
const test = require('tape').test ;
|
||||
const config = require('config');
|
||||
const opts = config.get('redis');
|
||||
const fs = require('fs');
|
||||
const logger = require('pino')({level: 'error'});
|
||||
process.on('unhandledRejection', (reason, p) => {
|
||||
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
|
||||
});
|
||||
|
||||
const stats = {
|
||||
increment: () => {},
|
||||
histogram: () => {}
|
||||
};
|
||||
|
||||
test('IBM - create access key', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getIbmAccessToken} = fn(opts, logger);
|
||||
|
||||
if (!process.env.IBM_API_KEY ) {
|
||||
t.pass('skipping IBM test since no IBM api_key provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
let obj = await getIbmAccessToken(process.env.IBM_API_KEY);
|
||||
//console.log({obj}, 'received access token from IBM');
|
||||
t.ok(obj.access_token && !obj.servedFromCache, 'successfull received access token from IBM');
|
||||
|
||||
obj = await getIbmAccessToken(process.env.IBM_API_KEY);
|
||||
//console.log({obj}, 'received access token from IBM - second request');
|
||||
t.ok(obj.access_token && obj.servedFromCache, 'successfully received access token from cache');
|
||||
|
||||
await client.flushall();
|
||||
t.end();
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('IBM - retrieve tts voices test', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
|
||||
t.pass('skipping IBM test since no IBM api_key and/or region provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const opts = {
|
||||
vendor: 'ibm',
|
||||
credentials: {
|
||||
tts_api_key: process.env.IBM_TTS_API_KEY,
|
||||
tts_region: process.env.IBM_TTS_REGION
|
||||
}
|
||||
};
|
||||
const obj = await getTtsVoices(opts);
|
||||
const {voices} = obj.result;
|
||||
//console.log(JSON.stringify(voices));
|
||||
t.ok(voices.length > 0 && voices[0].language,
|
||||
`GetVoices: successfully retrieved ${voices.length} voices from IBM`);
|
||||
|
||||
await client.flushall();
|
||||
|
||||
t.end();
|
||||
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
+1
-2
@@ -2,6 +2,5 @@ require('./docker_start');
|
||||
require('./synth');
|
||||
require('./list-voices');
|
||||
require('./aws');
|
||||
require('./ibm');
|
||||
require('./nuance');
|
||||
|
||||
require('./docker_stop');
|
||||
|
||||
@@ -12,164 +12,6 @@ const stats = {
|
||||
histogram: () => {}
|
||||
};
|
||||
|
||||
test('Verbio - get Access key and voices', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices, getVerbioAccessToken} = fn(opts, logger);
|
||||
if (!process.env.VERBIO_CLIENT_ID || !process.env.VERBIO_CLIENT_SECRET) {
|
||||
t.pass('skipping Verbio test since no Verbio Keys provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
const credentials = {
|
||||
client_id: process.env.VERBIO_CLIENT_ID,
|
||||
client_secret: process.env.VERBIO_CLIENT_SECRET
|
||||
};
|
||||
let obj = await getVerbioAccessToken(credentials);
|
||||
t.ok(obj.access_token , 'successfully received access token not from cache');
|
||||
const voices = await getTtsVoices({vendor: 'verbio', credentials});
|
||||
t.ok(voices && voices.length != 0, 'successfully received verbio voices');
|
||||
} catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('IBM - create access key', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getIbmAccessToken} = fn(opts, logger);
|
||||
|
||||
if (!process.env.IBM_API_KEY ) {
|
||||
t.pass('skipping IBM test since no IBM api_key provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
let obj = await getIbmAccessToken(process.env.IBM_API_KEY);
|
||||
//console.log({obj}, 'received access token from IBM');
|
||||
t.ok(obj.access_token && !obj.servedFromCache, 'successfull received access token from IBM');
|
||||
|
||||
obj = await getIbmAccessToken(process.env.IBM_API_KEY);
|
||||
//console.log({obj}, 'received access token from IBM - second request');
|
||||
t.ok(obj.access_token && obj.servedFromCache, 'successfully received access token from cache');
|
||||
|
||||
await client.flushall();
|
||||
t.end();
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('IBM - retrieve tts voices test', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
|
||||
t.pass('skipping IBM test since no IBM api_key and/or region provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const opts = {
|
||||
vendor: 'ibm',
|
||||
credentials: {
|
||||
tts_api_key: process.env.IBM_TTS_API_KEY,
|
||||
tts_region: process.env.IBM_TTS_REGION
|
||||
}
|
||||
};
|
||||
const obj = await getTtsVoices(opts);
|
||||
const {voices} = obj.result;
|
||||
//console.log(JSON.stringify(voices));
|
||||
t.ok(voices.length > 0 && voices[0].language,
|
||||
`GetVoices: successfully retrieved ${voices.length} voices from IBM`);
|
||||
|
||||
await client.flushall();
|
||||
|
||||
t.end();
|
||||
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Nuance hosted tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET ) {
|
||||
t.pass('skipping Nuance hosted test since no Nuance client_id and secret provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const opts = {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
client_id: process.env.NUANCE_CLIENT_ID,
|
||||
secret: process.env.NUANCE_SECRET
|
||||
}
|
||||
};
|
||||
let voices = await getTtsVoices(opts);
|
||||
t.ok(voices.length > 0 && voices[0].language,
|
||||
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
|
||||
|
||||
await client.flushall();
|
||||
|
||||
t.end();
|
||||
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Nuance on-prem tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NUANCE_TTS_URI ) {
|
||||
t.pass('skipping Nuance on-prem test since no Nuance uri provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const opts = {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
nuance_tts_uri: process.env.NUANCE_TTS_URI
|
||||
}
|
||||
};
|
||||
let voices = await getTtsVoices(opts);
|
||||
t.ok(voices.length > 0 && voices[0].language,
|
||||
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
|
||||
|
||||
await client.flushall();
|
||||
|
||||
t.end();
|
||||
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Google tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
@@ -1,81 +0,0 @@
|
||||
const test = require('tape').test ;
|
||||
const config = require('config');
|
||||
const opts = config.get('redis');
|
||||
const fs = require('fs');
|
||||
const logger = require('pino')({level: 'error'});
|
||||
process.on('unhandledRejection', (reason, p) => {
|
||||
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
|
||||
});
|
||||
|
||||
const stats = {
|
||||
increment: () => {},
|
||||
histogram: () => {}
|
||||
};
|
||||
|
||||
test('Nuance hosted tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET ) {
|
||||
t.pass('skipping Nuance hosted test since no Nuance client_id and secret provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const opts = {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
client_id: process.env.NUANCE_CLIENT_ID,
|
||||
secret: process.env.NUANCE_SECRET
|
||||
}
|
||||
};
|
||||
let voices = await getTtsVoices(opts);
|
||||
t.ok(voices.length > 0 && voices[0].language,
|
||||
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
|
||||
|
||||
await client.flushall();
|
||||
|
||||
t.end();
|
||||
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Nuance on-prem tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {client, getTtsVoices} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NUANCE_TTS_URI ) {
|
||||
t.pass('skipping Nuance on-prem test since no Nuance uri provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const opts = {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
nuance_tts_uri: process.env.NUANCE_TTS_URI
|
||||
}
|
||||
};
|
||||
let voices = await getTtsVoices(opts);
|
||||
t.ok(voices.length > 0 && voices[0].language,
|
||||
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
|
||||
|
||||
await client.flushall();
|
||||
|
||||
t.end();
|
||||
|
||||
}
|
||||
catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
+265
-222
@@ -304,14 +304,13 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
});
|
||||
t.ok(result.filePath.startsWith('say:'), 'Standard voice returns streaming say: path');
|
||||
t.ok(result.filePath.includes('vendor=google'), 'Standard voice streaming path contains vendor=google');
|
||||
t.ok(result.filePath.includes('use_live_api=0'), 'Standard voice uses use_live_api=0');
|
||||
t.ok(result.filePath.includes('use_gemini_tts=0'), 'Standard voice uses use_gemini_tts=0');
|
||||
t.ok(result.filePath.includes('api_mode=tts'), 'Standard voice uses api_mode=tts');
|
||||
t.ok(result.filePath.includes('voice=en-US-Wavenet-D'), 'Standard voice streaming path contains voice');
|
||||
// Verify credentials are base64 encoded (no raw JSON braces that would break FreeSWitch parsing)
|
||||
t.ok(result.filePath.includes('credentials='), 'Standard voice streaming path contains credentials');
|
||||
t.ok(!result.filePath.includes('credentials={'), 'Credentials are not raw JSON (base64 encoded)');
|
||||
|
||||
// Test 2: HD voice streaming (use_live_api=1)
|
||||
// Test 2: HD voice streaming (api_mode=live)
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: {
|
||||
@@ -327,11 +326,10 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
});
|
||||
t.ok(result.filePath.startsWith('say:'), 'HD voice returns streaming say: path');
|
||||
t.ok(result.filePath.includes('vendor=google'), 'HD voice streaming path contains vendor=google');
|
||||
t.ok(result.filePath.includes('use_live_api=1'), 'HD voice uses use_live_api=1 (Live API)');
|
||||
t.ok(result.filePath.includes('use_gemini_tts=0'), 'HD voice uses use_gemini_tts=0');
|
||||
t.ok(result.filePath.includes('api_mode=live'), 'HD voice uses api_mode=live');
|
||||
t.ok(result.filePath.includes('voice=en-US-Chirp3-HD-Charon'), 'HD voice streaming path contains voice');
|
||||
|
||||
// Test 3: Gemini TTS streaming (use_live_api=1)
|
||||
// Test 3: Gemini TTS streaming (api_mode=gemini)
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: {
|
||||
@@ -349,8 +347,7 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
});
|
||||
t.ok(result.filePath.startsWith('say:'), 'Gemini TTS returns streaming say: path');
|
||||
t.ok(result.filePath.includes('vendor=google'), 'Gemini TTS streaming path contains vendor=google');
|
||||
t.ok(result.filePath.includes('use_live_api=0'), 'Gemini TTS uses use_live_api=0');
|
||||
t.ok(result.filePath.includes('use_gemini_tts=1'), 'Gemini TTS uses use_gemini_tts=1');
|
||||
t.ok(result.filePath.includes('api_mode=gemini'), 'Gemini TTS uses api_mode=gemini');
|
||||
t.ok(result.filePath.includes(`model_name=${geminiModel}`), 'Gemini TTS streaming path contains model_name');
|
||||
t.ok(result.filePath.includes('prompt=Speak naturally.'), 'Gemini TTS streaming path contains prompt');
|
||||
|
||||
@@ -394,7 +391,7 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
// Commas in prompt should be replaced with semicolons
|
||||
t.ok(result.filePath.includes('prompt=Speak in a warm; friendly tone'), 'Commas in prompt are escaped to semicolons');
|
||||
|
||||
// Test 6: options.useLiveApi override (force live api on standard voice)
|
||||
// Test 6: options.apiMode override (force live on standard voice)
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: {
|
||||
@@ -405,14 +402,13 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing useLiveApi option override.',
|
||||
options: { useLiveApi: true },
|
||||
text: 'Testing apiMode option override to live.',
|
||||
options: { apiMode: 'live' },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes('use_live_api=1'), 'options.useLiveApi=true overrides default for standard voice');
|
||||
t.ok(result.filePath.includes('use_gemini_tts=0'), 'use_gemini_tts remains 0 for standard voice');
|
||||
t.ok(result.filePath.includes('api_mode=live'), 'options.apiMode=live overrides default for standard voice');
|
||||
|
||||
// Test 7: options.useGeminiTts override (force gemini tts without model)
|
||||
// Test 7: options.apiMode override (force gemini without model)
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: {
|
||||
@@ -423,14 +419,13 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'Kore',
|
||||
text: 'Testing useGeminiTts option override.',
|
||||
options: { useGeminiTts: true },
|
||||
text: 'Testing apiMode option override to gemini.',
|
||||
options: { apiMode: 'gemini' },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes('use_gemini_tts=1'), 'options.useGeminiTts=true overrides default');
|
||||
t.ok(result.filePath.includes('use_live_api=0'), 'use_live_api remains 0 without HD voice');
|
||||
t.ok(result.filePath.includes('api_mode=gemini'), 'options.apiMode=gemini overrides default');
|
||||
|
||||
// Test 8: Both options override together
|
||||
// Test 8: options.apiMode override (force tts on HD voice)
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: {
|
||||
@@ -440,13 +435,101 @@ test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) =>
|
||||
},
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing both options override.',
|
||||
options: { useLiveApi: true, useGeminiTts: true },
|
||||
voice: 'en-US-Chirp3-HD-Charon',
|
||||
text: 'Testing apiMode option override to tts on HD voice.',
|
||||
options: { apiMode: 'tts' },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes('use_live_api=1'), 'options.useLiveApi=true works with useGeminiTts');
|
||||
t.ok(result.filePath.includes('use_gemini_tts=1'), 'options.useGeminiTts=true works with useLiveApi');
|
||||
t.ok(result.filePath.includes('api_mode=tts'), 'options.apiMode=tts overrides HD voice default');
|
||||
|
||||
/* AudioConfig settings (speakingRate, pitch, volumeGainDb) */
|
||||
const googleCreds = {
|
||||
credentials: {
|
||||
client_email: creds.client_email,
|
||||
private_key: creds.private_key,
|
||||
},
|
||||
};
|
||||
|
||||
// Test 9: AudioConfig nested under options.audioConfig, as in google's API docs
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing nested audioConfig settings.',
|
||||
options: { audioConfig: { speakingRate: 1.4, pitch: -2.5, volumeGainDb: 6 } },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes(',speaking_rate=1.4'), 'nested audioConfig sets speaking_rate');
|
||||
t.ok(result.filePath.includes(',pitch=-2.5'), 'nested audioConfig sets pitch');
|
||||
t.ok(result.filePath.includes(',volume_gain_db=6'), 'nested audioConfig sets volume_gain_db');
|
||||
|
||||
// Test 10: AudioConfig flat at the top level of options, in camelCase or snake_case
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing flat audioConfig settings.',
|
||||
options: { speaking_rate: '0.8', volumeGainDb: 3 },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes(',speaking_rate=0.8'), 'flat snake_case speaking_rate is honored');
|
||||
t.ok(result.filePath.includes(',volume_gain_db=3'), 'flat camelCase volumeGainDb is honored');
|
||||
t.ok(!result.filePath.includes(',pitch='), 'unspecified audioConfig setting is omitted');
|
||||
|
||||
// Test 11: zero is a meaningful value for pitch and volumeGainDb, not an absent one
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing zero audioConfig settings.',
|
||||
options: { audioConfig: { pitch: 0, volumeGainDb: 0 } },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes(',pitch=0'), 'pitch=0 is passed through rather than dropped');
|
||||
t.ok(result.filePath.includes(',volume_gain_db=0'), 'volume_gain_db=0 is passed through rather than dropped');
|
||||
|
||||
// Test 12: non-numeric values are ignored, so they cannot corrupt the freeswitch param string
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing invalid audioConfig settings.',
|
||||
options: { audioConfig: { speakingRate: 'fast,evil=1', pitch: null, volumeGainDb: '' } },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(!result.filePath.includes('speaking_rate'), 'non-numeric speakingRate is ignored');
|
||||
t.ok(!result.filePath.includes('evil=1'), 'non-numeric value cannot inject extra params');
|
||||
t.ok(!result.filePath.includes('pitch='), 'null pitch is ignored');
|
||||
t.ok(!result.filePath.includes('volume_gain_db'), 'empty volumeGainDb is ignored');
|
||||
|
||||
// Test 13: no audioConfig supplied leaves the param string untouched
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'Testing absent audioConfig settings.',
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(!/speaking_rate|pitch=|volume_gain_db/.test(result.filePath),
|
||||
'no audioConfig params are added when none are supplied');
|
||||
|
||||
// Test 14: HD voice (api_mode=live) also carries speaking_rate, the one setting google streams support
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Chirp3-HD-Charon',
|
||||
text: 'Testing audioConfig on an HD voice.',
|
||||
options: { audioConfig: { speakingRate: 1.25 } },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.includes('api_mode=live'), 'HD voice with audioConfig still uses api_mode=live');
|
||||
t.ok(result.filePath.includes(',speaking_rate=1.25'), 'HD voice streaming path carries speaking_rate');
|
||||
|
||||
} catch (err) {
|
||||
console.error(err);
|
||||
@@ -532,6 +615,42 @@ test('Google TTS non-streaming tests (JAMBONES_DISABLE_TTS_STREAMING=true)', asy
|
||||
t.ok(!result.filePath.startsWith('say:'), 'Gemini TTS does NOT return streaming say: path when disabled');
|
||||
t.ok(result.filePath.endsWith('.mp3'), 'Gemini TTS returns mp3 file path');
|
||||
|
||||
const googleCreds = {
|
||||
credentials: {
|
||||
client_email: creds.client_email,
|
||||
private_key: creds.private_key,
|
||||
},
|
||||
};
|
||||
|
||||
/**
|
||||
* Test 4: AudioConfig settings are accepted by the synthesize API.
|
||||
* Google rejects out-of-range values with a 400, so a successful render also confirms
|
||||
* the settings reached the request rather than being silently dropped.
|
||||
*/
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'en-US-Wavenet-D',
|
||||
text: 'This is a test of audioConfig with streaming disabled.',
|
||||
options: { audioConfig: { speakingRate: 1.4, pitch: -2.5, volumeGainDb: 6 } },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.endsWith('.mp3'), 'standard voice renders mp3 with audioConfig settings applied');
|
||||
|
||||
/* Test 5: gemini voices ignore the AudioConfig settings rather than failing on them */
|
||||
result = await synthAudio(stats, {
|
||||
vendor: 'google',
|
||||
credentials: googleCreds,
|
||||
language: 'en-US',
|
||||
voice: 'Kore',
|
||||
model: geminiModel,
|
||||
text: 'This is a test of audioConfig on Gemini TTS.',
|
||||
options: { audioConfig: { speakingRate: 1.4, pitch: -2.5, volumeGainDb: 6 } },
|
||||
disableTtsCache: true
|
||||
});
|
||||
t.ok(result.filePath.endsWith('.mp3'), 'gemini voice renders mp3 with audioConfig settings skipped');
|
||||
|
||||
} catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
@@ -765,82 +884,6 @@ test('Azure custom voice speech synth tests', async(t) => {
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Nuance hosted speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET) {
|
||||
t.pass('skipping Nuance speech synth tests since NUANCE_CLIENT_ID or NUANCE_SECRET not provided');
|
||||
return t.end();
|
||||
}
|
||||
try {
|
||||
let opts = await synthAudio(stats, {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
client_id: process.env.NUANCE_CLIENT_ID,
|
||||
secret: process.env.NUANCE_SECRET,
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'Evan',
|
||||
text: 'This is a test. This is only a test',
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthesized nuance audio to ${opts.filePath}`);
|
||||
|
||||
opts = await synthAudio(stats, {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
client_id: process.env.NUANCE_CLIENT_ID,
|
||||
secret: process.env.NUANCE_SECRET,
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'Evan',
|
||||
text: 'This is a test. This is only a test',
|
||||
});
|
||||
t.ok(opts.servedFromCache, `successfully retrieved nuance audio from cache ${opts.filePath}`);
|
||||
} catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Nuance on-prem speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NUANCE_TTS_URI) {
|
||||
t.pass('skipping Nuance on prem speech synth tests since NUANCE_TTS_URI not provided');
|
||||
return t.end();
|
||||
}
|
||||
try {
|
||||
let opts = await synthAudio(stats, {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
nuance_tts_uri: process.env.NUANCE_TTS_URI
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'Evan',
|
||||
text: 'This is a test of on-prem. This is only a test',
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthesized nuance audio to ${opts.filePath}`);
|
||||
|
||||
opts = await synthAudio(stats, {
|
||||
vendor: 'nuance',
|
||||
credentials: {
|
||||
nuance_tts_uri: process.env.NUANCE_TTS_URI
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'Evan',
|
||||
text: 'This is a test of on-prem. This is only a test',
|
||||
});
|
||||
t.ok(opts.servedFromCache, `successfully retrieved nuance audio from cache ${opts.filePath}`);
|
||||
} catch (err) {
|
||||
console.error(err);
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Nvidia speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
@@ -878,46 +921,6 @@ test('Nvidia speech synth tests', async(t) => {
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('IBM watson speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
|
||||
t.pass('skipping IBM Watson speech synth tests since IBM_TTS_API_KEY or IBM_TTS_API_KEY not provided');
|
||||
return t.end();
|
||||
}
|
||||
const text = `<speak> Hi there and welcome to jambones! jambones is the <sub alias="seapass">CPaaS</sub> designed with the needs of communication service providers in mind. This is an example of simple text-to-speech, but there is so much more you can do. Try us out!</speak>`;
|
||||
try {
|
||||
let opts = await synthAudio(stats, {
|
||||
vendor: 'ibm',
|
||||
credentials: {
|
||||
tts_api_key: process.env.IBM_TTS_API_KEY,
|
||||
tts_region: process.env.IBM_TTS_REGION,
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'en-US_AllisonV2Voice',
|
||||
text,
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthesized ibm audio to ${opts.filePath}`);
|
||||
|
||||
opts = await synthAudio(stats, {
|
||||
vendor: 'ibm',
|
||||
credentials: {
|
||||
tts_api_key: process.env.IBM_TTS_API_KEY,
|
||||
tts_region: process.env.IBM_TTS_REGION,
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'en-US_AllisonV2Voice',
|
||||
text,
|
||||
});
|
||||
t.ok(opts.servedFromCache, `successfully retrieved ibm audio from cache ${opts.filePath}`);
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Custom Vendor speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
@@ -1021,55 +1024,6 @@ test('Elevenlabs speech synth tests', async(t) => {
|
||||
client.quit();
|
||||
});
|
||||
|
||||
const testPlayHT = async(t, voice_engine) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.PLAYHT_API_KEY || !process.env.PLAYHT_USER_ID) {
|
||||
t.pass('skipping PlayHT speech synth tests since PLAYHT_API_KEY or PLAYHT_USER_ID is/are not provided');
|
||||
return t.end();
|
||||
}
|
||||
const text = 'Hi there and welcome to jambones! ' + Date.now();
|
||||
try {
|
||||
const opts = await synthAudio(stats, {
|
||||
vendor: 'playht',
|
||||
credentials: {
|
||||
api_key: process.env.PLAYHT_API_KEY,
|
||||
user_id: process.env.PLAYHT_USER_ID,
|
||||
voice_engine,
|
||||
options: JSON.stringify({
|
||||
quality: 'medium',
|
||||
speed: 1,
|
||||
seed: 1,
|
||||
temperature: 1,
|
||||
emotion: 'female_happy',
|
||||
voice_guidance: 3,
|
||||
style_guidance: 20,
|
||||
text_guidance: 1,
|
||||
})
|
||||
},
|
||||
language: 'english',
|
||||
voice: 's3://voice-cloning-zero-shot/d9ff78ba-d016-47f6-b0ef-dd630f59414e/female-cs/manifest.json',
|
||||
text,
|
||||
renderForCaching: true
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully playht eleven audio to ${opts.filePath}`);
|
||||
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
};
|
||||
|
||||
test('PlayHT speech synth tests', async(t) => {
|
||||
await testPlayHT(t, 'PlayHT2.0-turbo');
|
||||
});
|
||||
|
||||
test('PlayHT3.0 speech synth tests', async(t) => {
|
||||
await testPlayHT(t, 'Play3.0');
|
||||
});
|
||||
|
||||
test('Cartesia speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
@@ -1104,6 +1058,65 @@ test('Cartesia speech synth tests', async(t) => {
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('gradium speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.GRADIUM_API_KEY) {
|
||||
t.pass('skipping gradium speech synth tests since GRADIUM_API_KEY is not provided');
|
||||
return t.end();
|
||||
}
|
||||
const text = 'Hi there and welcome to jambones! ' + Date.now();
|
||||
try {
|
||||
const opts = await synthAudio(stats, {
|
||||
vendor: 'gradium',
|
||||
credentials: {
|
||||
api_key: process.env.GRADIUM_API_KEY,
|
||||
model_id: 'default'
|
||||
},
|
||||
voice: 'YTpq7expH9539ERJ',
|
||||
text,
|
||||
renderForCaching: true
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthed gradium audio to ${opts.filePath}`);
|
||||
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('nineninesix speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.NINENINESIX_API_KEY) {
|
||||
t.pass('skipping nineninesix speech synth tests since NINENINESIX_API_KEY is not provided');
|
||||
return t.end();
|
||||
}
|
||||
const text = 'Hi there and welcome to jambones! ' + Date.now();
|
||||
try {
|
||||
const opts = await synthAudio(stats, {
|
||||
vendor: 'nineninesix',
|
||||
credentials: {
|
||||
api_key: process.env.NINENINESIX_API_KEY,
|
||||
model_id: 'gepard-1.0'
|
||||
},
|
||||
language: 'en',
|
||||
voice: '3ad7a827-7fd1-4954-bf35-47d4cc33d9ed',
|
||||
text,
|
||||
renderForCaching: true
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthed nineninesix audio to ${opts.filePath}`);
|
||||
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('inworld speech synth', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
@@ -1261,39 +1274,6 @@ test('whisper speech synth tests', async(t) => {
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('Verbio speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.VERBIO_CLIENT_ID || !process.env.VERBIO_CLIENT_SECRET) {
|
||||
t.pass('skipping Verbio Synthesize test since no Verbio Keys provided');
|
||||
t.end();
|
||||
client.quit();
|
||||
return;
|
||||
}
|
||||
|
||||
const text = 'Hi there and welcome to jambones!';
|
||||
try {
|
||||
let opts = await synthAudio(stats, {
|
||||
vendor: 'verbio',
|
||||
credentials: {
|
||||
client_id: process.env.VERBIO_CLIENT_ID,
|
||||
client_secret: process.env.VERBIO_CLIENT_SECRET
|
||||
},
|
||||
language: 'en-US',
|
||||
voice: 'tommy_en-us',
|
||||
text,
|
||||
renderForCaching: true
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthesized whisper audio to ${opts.filePath}`);
|
||||
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
})
|
||||
|
||||
test('Deepgram speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
@@ -1322,6 +1302,69 @@ test('Deepgram speech synth tests', async(t) => {
|
||||
client.quit();
|
||||
})
|
||||
|
||||
test('Deepgram Flux speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.DEEPGRAM_API_KEY) {
|
||||
t.pass('skipping Deepgram Flux speech synth tests since DEEPGRAM_API_KEY');
|
||||
return t.end();
|
||||
}
|
||||
const text = 'Hi there and welcome to jambones!';
|
||||
try {
|
||||
const opts = await synthAudio(stats, {
|
||||
vendor: 'deepgramflux',
|
||||
credentials: {
|
||||
api_key: process.env.DEEPGRAM_API_KEY
|
||||
},
|
||||
model: process.env.DEEPGRAM_FLUX_MODEL || 'flux-alexis-en',
|
||||
text,
|
||||
renderForCaching: true
|
||||
});
|
||||
t.ok(!opts.servedFromCache, `successfully synthesized deepgramflux audio to ${opts.filePath}`);
|
||||
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('xai speech synth tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {synthAudio, client} = fn(opts, logger);
|
||||
|
||||
if (!process.env.XAI_API_KEY) {
|
||||
t.pass('skipping xai speech synth tests - no XAI_API_KEY');
|
||||
return t.end();
|
||||
}
|
||||
const text = 'Hi there and welcome to jambones!';
|
||||
try {
|
||||
const opts = await synthAudio(stats, {
|
||||
vendor: 'xai',
|
||||
credentials: {
|
||||
api_key: process.env.XAI_API_KEY,
|
||||
options: JSON.stringify({
|
||||
voice: process.env.XAI_VOICE || 'eve',
|
||||
speed: 1.0,
|
||||
optimize_streaming_latency: 1,
|
||||
text_normalization: true
|
||||
})
|
||||
},
|
||||
language: 'en',
|
||||
voice: process.env.XAI_VOICE || 'eve',
|
||||
text,
|
||||
renderForCaching: true
|
||||
});
|
||||
t.ok(!opts.servedFromCache && opts.filePath, `successfully synthesized xai audio to ${opts.filePath}`);
|
||||
|
||||
} catch (err) {
|
||||
console.error(JSON.stringify(err));
|
||||
t.end(err);
|
||||
}
|
||||
client.quit();
|
||||
});
|
||||
|
||||
test('TTS Cache tests', async(t) => {
|
||||
const fn = require('..');
|
||||
const {purgeTtsCache, getTtsSize, client} = fn(opts, logger);
|
||||
|
||||
Reference in New Issue
Block a user