Compare commits

...
232 Commits
Author SHA1 Message Date
Dave Horton 04d1fb548b 1.0.9 2026-07-21 06:52:05 -04:00
Hoan Luu Huu bbf0167b40 support deepgramflux tts (#148) 2026-07-21 06:49:38 -04:00
Dave Horton 4cfa286242 1.0.8 2026-07-05 17:42:09 -04:00
Hoan Luu Huu 42ac63adc6 support xaiTTS (#147) 2026-07-05 17:41:36 -04:00
Dave Horton 12a121672d 1.0.7 2026-07-03 07:12:20 -04:00
Hoan Luu HuuandClaude Opus 4.8 7b94a5a969 feat(murf): add Murf.ai TTS support to synthAudio (#146)
* feat(murf): add Murf.ai TTS support to synthAudio

Add synthMurf() following the rimelabs/cartesia pattern:
- streaming path returns a say:{vendor=murf,...} filePath consumed by the
  FreeSWITCH mod_murf_tts module
- non-streaming path calls POST /v1/speech/stream (api-key header) and returns
  WAV audio for cache rendering

Register murf in the supported-vendor assert list and the synth switch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(murf): drop Accept: audio/basic header (caused 406 Not Acceptable)

Murf's /v1/speech/stream rejects an unmatched Accept header with 406; the
response container is chosen by the `format` body field instead. Verified a
WAV request now returns 200 (valid RIFF/WAVE).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 07:11:39 -04:00
Dave Horton c47b4883c7 1.0.6 2026-06-17 16:20:30 -04:00
Dave HortonandClaude Opus 4.8 7d076bb8b4 chore: deprecate + remove verbio, nuance, playht speech vendor support (#144)
* chore: deprecate and remove verbio, nuance speech vendor support

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: also deprecate and remove PlayHT speech vendor

PlayHT was acquired and no longer provides the service.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:20:00 -04:00
Dave HortonandClaude Opus 4.8 142323a151 fix(synth-audio): include the Polly VoiceId in the aws say: url
synthPolly received voice but never emitted it into the say:{...} url
(unlike synthMicrosoft, which includes voice=). Polly's SynthesizeSpeech
requires a VoiceId, so the media-server-native (say:-url) Polly path had no
way to pick a voice. Add voice= to the params.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 22:31:57 -04:00
Dave Horton d176a644fe 1.0.4 2026-06-06 10:44:11 +02:00
Hoan Luu Huu 644f2918dc support cartesia sonic3.5 (#143) 2026-06-06 10:42:40 +02:00
Dave HortonandClaude Opus 4.5 4f430b9785 update publish workflow to use actions v4
Fixes npm warning about deprecated always-auth config by updating
setup-node from v3 to v4.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-06-03 08:51:56 -04:00
Dave Horton 18c658d20c 1.0.3 2026-06-03 07:27:44 -04:00
Dave Horton 062608cf13 update lock file 2026-06-03 07:27:28 -04:00
Hoan Luu Huu a2ea94c14a support rimelabs coda (#142) 2026-06-03 07:24:13 -04:00
Dave Horton 04bb85ef81 1.0.1 2026-03-25 10:08:57 -04:00
Dave Horton c123f19898 remove ibm speech since it is not used (to my knowledge) and has dependencies with vulnerabilities (#141) 2026-03-25 10:08:30 -04:00
Dave Horton 305695d068 0.2.30 2026-01-22 08:03:44 -05:00
Dave Horton 1477752d40 update dep 2026-01-22 08:03:24 -05:00
Dave Horton fbedfe947f vuln 2026-01-22 07:58:06 -05:00
Dave Horton ab88facd52 Merge pull request #139 from jambonz/feat/google_gemini_tts
google tts support api_mode
2026-01-22 07:52:04 -05:00
Hoan HL a2be64da89 google tts support api_mode 2026-01-22 16:36:57 +07:00
Dave Horton c5e0f256e6 0.2.28 2026-01-17 21:39:59 -05:00
Dave Horton d0378751ba Merge pull request #135 from jambonz/feat/gemini_tts
support gemini tts
2026-01-17 21:39:28 -05:00
Hoan HL a7e391fcb2 wip 2026-01-18 08:45:14 +07:00
Hoan HL 10095de25d wip 2026-01-17 16:16:11 +07:00
Hoan HL 89007ba7cc wip 2026-01-17 15:13:47 +07:00
Hoan HL 29edccd5bf wip 2026-01-17 15:10:17 +07:00
Hoan HL ca9538030f wip 2026-01-14 17:37:38 +07:00
Hoan HL 417d58080e wip 2026-01-14 17:34:07 +07:00
Hoan HL 62fec6f5e4 wip 2026-01-14 16:00:43 +07:00
Hoan HL 490d23a703 wip 2026-01-14 15:46:29 +07:00
Hoan HL b91e6cc145 wip 2026-01-14 13:58:38 +07:00
Hoan HL 0e33358254 wip 2026-01-14 13:51:32 +07:00
Hoan HL 6b2b35acfb wip 2026-01-12 18:26:47 +07:00
Hoan HL ded60cb7aa wip 2026-01-12 17:18:11 +07:00
Hoan HL 460ca70ea7 wip 2026-01-12 12:55:43 +07:00
Hoan HL 0fbbbb8053 add testcases 2026-01-12 09:21:41 +07:00
Hoan HL 0ea7082da2 support gemini tts 2026-01-11 07:30:18 +07:00
Dave Horton 5f7e7458bb 0.2.27 2025-11-17 07:26:45 -05:00
Dave Horton f6714fb9e1 Merge pull request #134 from jambonz/fix/1628
fixed cartesia collect audio from stream
2025-11-13 07:15:51 -05:00
Hoan HL c04ef29f7c fixed cartesia collect audio from stream 2025-11-13 14:10:00 +07:00
Dave Horton 8154944252 0.2.26 2025-10-30 07:08:47 -04:00
Dave Horton 16fe8dce01 0.2.25 2025-10-30 07:08:10 -04:00
Dave Horton 11a955500d Merge pull request #132 from jambonz/feat/sonic_3
cartesia support volume for sonic3
2025-10-30 07:07:15 -04:00
Hoan HL c122129b55 wip 2025-10-30 12:29:09 +07:00
Hoan HL 41b26b966b wip 2025-10-30 12:08:38 +07:00
Hoan HL aa09c15b20 wip 2025-10-30 06:46:57 +07:00
Hoan HL 32d5f12638 cartesia support volume for sonic3 2025-10-30 05:57:38 +07:00
Dave Horton 4e336822a0 Merge pull request #131 from jambonz/feat/gh_fs_1384
support elevenlabs api_uri
2025-10-08 13:45:02 -04:00
Hoan HL b898d794b0 wip 2025-10-08 15:05:04 +07:00
Hoan HL fb754ca101 support elevenlabs api_uri 2025-10-08 10:26:17 +07:00
Dave Horton 8d8195be9a 0.2.24 2025-10-03 08:54:14 -04:00
Dave Horton eb4e1a773f Merge pull request #130 from jambonz/feat/disableTtsCache
set write_cache_file = 0 when disableTtsCache
2025-10-03 02:21:57 -04:00
Hoan HL 6fecb8755d set write_cache_file = 0 when disableTtsCache 2025-10-03 11:09:30 +07:00
Dave Horton fdb56cbc77 Merge pull request #128 from jambonz/feat/custom_tts_stream
support custom tts Stream
2025-09-11 09:24:18 -04:00
Quan HL 5328a60de8 support custom tts Stream 2025-09-11 09:18:42 -04:00
Dave Horton ea1ab301d5 0.2.23 2025-09-10 22:53:45 -04:00
Dave Horton 0a4994c5a3 Merge pull request #129 from jambonz/fix/fd_1372
remove optimize_streaming_latency as default option for elevenlabs
2025-09-10 22:53:27 -04:00
Quan HL cb71460189 remove optimize_streaming_latency as default option for elevenlabs 2025-09-11 09:44:53 +07:00
Dave Horton f08efa49b1 0.2.22 2025-08-20 18:03:08 -04:00
Dave Horton 80eda37ef6 Merge pull request #126 from jambonz/feat/revamp-playback-id
use synth key as playback id
2025-08-20 18:02:35 -04:00
Dave Horton 7768e58b19 use synth key as playback id 2025-08-20 17:11:45 -04:00
Dave Horton 6efc99ad83 0.2.21 2025-08-20 10:59:40 -04:00
Dave Horton 127d01a8c1 add playback_id setting to additional tts vendors 2025-08-20 10:58:28 -04:00
Dave Horton 1900b26d8b 0.2.20 2025-08-20 10:01:01 -04:00
Dave Horton aa25a6f6d9 Merge pull request #125 from jambonz/feat/add-playback-id
add playback_id to say metadata
2025-08-20 10:00:37 -04:00
Dave Horton 8dfe17b600 add playback_id to say metadata 2025-08-20 08:58:46 -04:00
Dave Horton 5463e9f56e 0.2.19 2025-08-17 09:35:07 -04:00
Dave Horton 2eb36af650 Merge pull request #124 from jambonz/fix/aws_tts
support mod_aws_tts with engine parameter
2025-08-17 09:23:58 -04:00
Quan HL b186cbc4f2 support mod_aws_tts with engine parameter 2025-08-17 06:44:40 +07:00
Dave Horton b1b9c182a9 0.2.18 2025-08-14 08:17:02 -04:00
Dave Horton 15c77626fc Merge pull request #123 from jambonz/feat/mod_aws_tts
support mod_aws_tts
2025-08-14 08:16:32 -04:00
Quan HL 76126ec0b3 wip 2025-08-14 18:49:29 +07:00
Quan HL 55431ee511 wip 2025-08-14 17:19:46 +07:00
Quan HL 1c76f74b4a wip 2025-08-14 17:13:05 +07:00
Quan HL 62ad8abb8e wip 2025-08-14 16:49:47 +07:00
Quan HL 92734aaedb wip 2025-08-14 16:38:36 +07:00
Quan HL fe4ccfe7d7 support mod_aws_tts 2025-08-14 16:34:40 +07:00
Dave Horton ab05976032 0.2.17 2025-08-13 07:41:24 -04:00
Dave Horton babd8abc51 Merge pull request #122 from jambonz/feat/resemble_tts_01
support resemble tts
2025-08-13 07:33:12 -04:00
Quan HL 91f85e16b2 support resemble tts 2025-08-13 14:00:24 +07:00
Dave Horton 8a0c5e3dbd tag 2025-08-10 20:11:52 -04:00
Dave Horton d40e012138 Merge pull request #120 from jambonz/feat/resemble_tts
support resemble ai
2025-08-10 19:12:26 -04:00
Quan HL 49b23f5240 support resemble ai 2025-08-10 15:29:01 +07:00
Dave Horton 043642ea5f 0.2.15 2025-07-13 10:27:02 -04:00
Dave Horton 91ea9820f9 Merge pull request #119 from vdharashive/main
TMP_FOLDER configurable
2025-07-10 07:32:16 -04:00
Vinod Dharashive fad0f57b13 TMP_FOLDER configurable
https://github.com/jambonz/speech-utils/issues/118
2025-07-10 15:01:08 +05:30
Dave Horton 2608e149a8 0.2.14 2025-07-09 13:43:38 -04:00
Dave Horton d21289acd6 sec fix 2025-07-09 13:43:12 -04:00
Dave Horton a7b19e1d8a Merge pull request #117 from sathishkumarpa-Kore/PLAT-41939_2
Fix security vulnerability by upgrading ibm-watson to latest version
2025-07-09 13:42:01 -04:00
sathishkumarpa-Kore 09c7aa9a5f Fix security vulnerability by upgrading ibm-watson 2025-07-03 18:22:34 +05:30
Dave Horton f311c21327 0.2.13 2025-06-26 08:02:24 -04:00
Dave Horton 535d0191da Merge pull request #115 from jambonz/feat/inworld_tts
support inworld tts
2025-06-26 08:01:23 -04:00
Quan HL c00d4f9be4 support inworld tts 2025-06-26 17:34:24 +07:00
Dave Horton db135ee5ad 0.2.12 2025-06-11 11:08:34 +02:00
Dave Horton 5eac4c2ad8 Merge pull request #114 from vasudevanubrolu/feat/893-azure-ssml
Feat/893 azure ssml
2025-06-11 11:05:48 +02:00
vasudevanubrolu e23a1a6d09 feat/893 azure ssml add namespace check 2025-06-04 12:53:37 +05:30
vasudevanubrolu 3a78300a08 feat/893 azure ssml only change free text 2025-06-04 11:18:11 +05:30
vasudevanubrolu 1d7390e7ae feat/893 azure ssml fix only on condition 2025-05-30 14:55:22 +05:30
vasudevanubrolu 3c0940f657 feat/893 azure ssml lang syntax fix 2025-05-30 14:55:22 +05:30
Dave Horton 6fb6195b16 0.2.11 2025-05-27 10:09:01 -04:00
Dave Horton 2ff8587601 Merge pull request #113 from vasudevanubrolu/feat/893-azure-ssml
Feat/893 azure ssml
2025-05-27 10:08:48 -04:00
vasudevanubrolu 925bd26a70 feat/893 add add lang tag for accent to be picked 2025-05-27 13:41:07 +05:30
vasudevanubrolu ddea485f5f feat/893 azure ssml lang for on prem 2025-05-26 15:13:45 +05:30
vasudevanubrolu 49de25feb8 feat/893 ssml config 2025-05-26 12:10:22 +05:30
vasudevanubrolu 08b55b8d79 feat/893 support default azure ssml config 2025-05-26 12:10:22 +05:30
vasudevanubrolu 59bca302b9 feat/893 azure ssml based on env config 2025-05-26 12:10:22 +05:30
Dave Horton 0d98f73c43 0.2.10 2025-05-13 09:57:01 -04:00
Dave Horton 36670e0080 Merge pull request #111 from vasudevanubrolu/feat/864-playht-onprem
feat/864 playht on prem
2025-05-13 09:50:46 -04:00
vasudevanubrolu 61672f9868 feat/864 playht on prem pr changes 2025-05-13 18:47:26 +05:30
vasudevan-Kore ddeca8eb99 feat/864 playht on prem 2025-05-13 18:47:26 +05:30
Dave Horton 6e8271b2f6 0.2.9 2025-05-13 07:47:03 -04:00
Dave Horton 74a4938eb6 Merge pull request #112 from jambonz/feat/whisper_instructions
support openai whisper instructions
2025-05-13 07:46:34 -04:00
Quan HL cb50f603cd wip 2025-05-13 18:02:11 +07:00
Quan HL 9a524f00bc wip 2025-05-13 17:58:57 +07:00
Quan HL 545df0b770 wip 2025-05-13 17:55:41 +07:00
Quan HL 9409405769 support openai whisper instructions 2025-05-13 15:23:59 +07:00
Dave Horton 7dc3bbdb01 0.2.8 2025-05-08 09:31:40 -04:00
Dave Horton f8f9de2645 Merge pull request #110 from jambonz/feat/rimlabs_arcana
support rimelabs arcana
2025-05-06 09:22:35 -04:00
Quan HL 467d7ede26 support rimelabs arcana 2025-05-06 16:16:18 +07:00
Dave Horton fc211ab2e7 0.2.7 2025-04-28 19:31:59 -04:00
Dave Horton 536f8aab31 Merge pull request #109 from jambonz/feat/riva_tts
support riva tts stream
2025-04-28 19:31:30 -04:00
Hoan Luu Huu 69b3fdffbe Merge branch 'main' into feat/riva_tts 2025-04-28 08:47:44 +07:00
Dave Horton 53fe72d89e 0.2.6 2025-04-23 07:12:54 -04:00
Dave Horton e79c15c5da update version 2025-04-23 07:12:25 -04:00
Hoan Luu Huu a4a427e174 Merge branch 'main' into feat/riva_tts 2025-04-23 18:11:19 +07:00
Dave Horton b697bc5268 Merge pull request #108 from jambonz/feat/ell_tts_new_params
elevenlabs tts speed and pronunciation_dictionary_locators
2025-04-23 07:08:49 -04:00
Quan HL d35d7f0aec support riva tts stream 2025-04-23 17:54:21 +07:00
Quan HL ed1c564fa2 wip 2025-04-04 16:15:45 +07:00
Quan HL 7189d471c1 elevenlabs tts speed and pronunciation_dictionary_locators 2025-04-04 15:44:29 +07:00
Dave Horton 04080cc5ec 0.2.4 2025-03-19 21:48:31 -04:00
Dave Horton 3e5ab4af27 Merge pull request #107 from jambonz/update-deps
update undici
2025-03-19 21:48:02 -04:00
Dave Horton f06dddd2f7 update undici 2025-03-19 21:46:21 -04:00
Dave Horton 7b4a71f55d 0.2.3 2025-02-07 07:20:12 -05:00
Dave Horton d552b65618 Merge pull request #106 from jambonz/feat/rimelabs_voices
rimelabs support multiple model and languages
2025-02-07 07:19:47 -05:00
Quan HL f701b50244 rimelabs support multiple model and languages 2025-02-07 14:20:52 +07:00
Dave Horton 199e502fbe 0.2.2 2025-02-03 08:14:15 -05:00
Dave Horton 6769779189 Merge pull request #105 from jambonz/fix/gh_1059
tts key should include model
2025-02-03 08:13:54 -05:00
Quan HL 4b43c3986c wip 2025-02-02 19:03:09 +07:00
Quan HL e1292772e6 tts key should include model 2025-02-02 18:54:16 +07:00
Dave Horton 37fb045431 0.2.1 2024-12-18 22:16:43 -05:00
Dave Horton c328df67a2 major semver bump 2024-12-18 22:15:28 -05:00
Dave Horton 877cba7065 Merge pull request #103 from jambonz/feat/refactor_synth
remove audio extension from audio key
2024-12-18 22:13:52 -05:00
Quan HL 35a178b468 remove audio extension from audio key 2024-12-18 16:58:54 +07:00
Quan HL 7327190471 remove audio extension from audio key 2024-12-18 16:46:24 +07:00
Dave Horton e5b5e3d0c6 0.1.24 2024-12-16 07:27:18 -05:00
Dave Horton 3760be088b update deps 2024-12-16 07:26:49 -05:00
Dave Horton 5cb52eb8ec Merge pull request #101 from jambonz/feat/tts_cartesia
support cartesia tts
2024-12-16 07:25:29 -05:00
Quan HL 8a390a8edf support cartesia tts 2024-12-16 15:58:18 +07:00
Dave Horton d96ab5cdf3 Merge pull request #99 from jambonz/feat/support_aws_instance_profile
support aws instance profile to get key
2024-11-29 21:58:22 -05:00
Quan HL 3bf74671c6 support aws instance profile to get key 2024-11-30 08:40:26 +07:00
Dave Horton a47ef6d7c4 0.1.22 2024-11-04 07:38:23 -05:00
Dave Horton 84089fa528 Merge pull request #98 from jambonz/fix/freshdesk_411
Fix custom tts vendor cached file can not be played
2024-11-04 07:37:49 -05:00
Quan HL 72be44eea2 adding testcase 2024-11-04 15:53:11 +07:00
Quan HL 05d6c4b32d fixed custom vendor cache audio stores file extension 2024-11-04 15:43:44 +07:00
Dave Horton 0c7e15d0a2 0.1.21 2024-10-31 09:48:18 -04:00
Dave Horton 63efecf9d9 Merge pull request #97 from jambonz/feat/google_voice_cloning
support google voice cloning
2024-10-31 09:47:12 -04:00
Quan HL 153ac3f1a4 fix review comment 2024-10-31 20:30:35 +07:00
Quan HL 115faa9f89 support google voice cloning 2024-10-31 20:23:11 +07:00
Dave Horton f183852961 0.1.20 2024-10-18 12:23:26 -04:00
Dave Horton 34c3e01729 Merge pull request #95 from jambonz/fix/rimelabs
fix rimelabs typo issue on getFileExtension function
2024-10-18 12:22:52 -04:00
Quan HL 9b2b16199e fix rimelabs typo issue on getFileExtension function 2024-10-18 22:44:17 +07:00
Dave Horton 9112c5f0ea 0.1.19 2024-10-16 07:23:40 -04:00
Dave Horton 50783dfd0a Merge pull request #94 from jambonz/fix/playht30_lang
add language to playht3.0
2024-10-16 07:23:04 -04:00
Quan HL ca0ef76fe1 add language to playht3.0 2024-10-16 08:14:38 +07:00
Dave Horton 7c91c537e4 0.1.18 2024-10-11 07:33:37 -04:00
Dave Horton 9747526664 Merge pull request #93 from jambonz/fix/playht_3.0
fixed playht3.0 cannot be played if credential is cached
2024-10-11 07:32:57 -04:00
Quan HL 31a0c7b02c fixed playht3.0 cannot be played if credential is cached 2024-10-11 12:05:38 +07:00
Dave Horton c18fbacd1b update playht3 2024-10-09 13:29:01 -04:00
Dave Horton b0fee6bbf1 Merge pull request #92 from jambonz/feat/playht30
support playht3.0
2024-10-09 13:26:53 -04:00
Quan HL f6cead6e92 add top_p and repetition_penalty to playht3.0 2024-10-03 19:24:23 +07:00
Quan HL 05fc96edc0 wip 2024-09-27 18:24:03 +07:00
Quan HL 6794a0b3be support playht3.0 2024-09-27 12:25:47 +07:00
Quan HL 1a04fd736c support playht3.0 2024-09-27 12:08:41 +07:00
Dave Horton 1846203807 0.1.16 2024-09-16 15:53:14 -04:00
Dave Horton 75be8658c1 Merge pull request #91 from jambonz/fix/diff_playht_voice_quality
fix playht has stream and cached audio differrent quality
2024-09-16 15:52:42 -04:00
Quan HL 91a5eebbaf fixed review comment 2024-09-16 18:46:10 +07:00
Quan HL c96f1e86ee fix review comment 2024-09-16 18:15:45 +07:00
Quan HL 8016c0886a fix playht has stream and cached audio differrent quality 2024-09-16 09:01:45 +07:00
Dave Horton 8f216e64d8 0.1.15 2024-08-12 09:30:09 -04:00
Dave Horton e1f4486e01 bump version 2024-08-12 09:27:12 -04:00
Dave Horton b0d6272974 Merge pull request #84 from jambonz/feat/precache_audio_with_tts_stream
support precache audio with tts stream enabled
2024-08-12 09:26:00 -04:00
Quan HL ef23b0807a add comment 2024-08-12 20:16:24 +07:00
Quan HL ab7e25243d improve on check precache 2024-08-12 20:10:43 +07:00
Quan HL bf0ea14423 install docker 2024-08-12 18:40:47 +07:00
Quan HL 305dabd84b wip 2024-08-12 18:35:48 +07:00
Quan HL b6a3fa5081 support precache audio with tts stream enabled 2024-08-12 18:29:01 +07:00
Dave Horton 73feadc4c4 0.1.13 2024-08-06 11:01:09 -04:00
Dave Horton aad0f4d62c Merge pull request #82 from jambonz/feat/deepgram_tts_endpoint
deepgram tts support endpoint for on-premise
2024-08-06 11:00:38 -04:00
Hoan Luu Huu 602b0cc60e Merge branch 'main' into feat/deepgram_tts_endpoint 2024-07-31 14:03:41 +07:00
Dave Horton 8511e762c8 version update 2024-07-30 07:29:25 -04:00
Dave Horton b75d3068c9 Merge pull request #81 from jambonz/feat/gh_fs_832
allow configure STS session expiry
2024-07-30 07:14:06 -04:00
Quan HL 461178e726 wip 2024-07-29 20:56:55 +07:00
Quan HL e0e4d47340 wip 2024-07-29 20:53:55 +07:00
Quan HL a595faa378 deepgram tts support endpoint for on-premise 2024-07-29 19:45:13 +07:00
Quan HL 7fd1e1a3c3 allow configure STS session expiry 2024-07-29 18:18:37 +07:00
Dave Horton 7f6a3d349c 0.1.11 2024-06-14 07:37:19 -04:00
Dave Horton 50429ff535 bump version 2024-06-14 07:36:53 -04:00
Dave Horton 3bf0ef8ea3 Merge pull request #80 from jambonz/fix/aws_arnrole
fix aws arnrole
2024-06-14 07:34:22 -04:00
Quan HL e9a5e83e36 wip 2024-06-14 15:04:19 +07:00
Quan HL 86a64ac091 wip 2024-06-14 15:00:56 +07:00
Quan HL 97e06b3ab3 wip 2024-06-14 10:23:22 +07:00
Quan HL 09e833d910 wip 2024-06-14 10:19:55 +07:00
Quan HL 8c4e12e54f wip 2024-06-14 10:18:41 +07:00
Quan HL 2642bd71a4 wip 2024-06-14 10:16:57 +07:00
Quan HL c4feac916f fix aws arnrole 2024-06-14 10:15:23 +07:00
Dave Horton 2ec56f564e 0.1.9 2024-06-06 12:31:13 -04:00
Dave Horton 9399ddcb58 Merge pull request #78 from jambonz/env/disable-ms-streaming
Env/disable ms streaming
2024-06-06 12:30:48 -04:00
Dave Horton aebf4eda30 lint 2024-06-06 12:25:53 -04:00
Dave Horton 6b0bdfdf2f add env JAMBONES_DISABLE_AZURE_TTS_STREAMING to disable Microsoft TTS streaming 2024-06-06 12:25:08 -04:00
Dave Horton d214c3184f 0.1.8 2024-06-05 06:48:09 -04:00
Dave Horton c335a508a3 bump version 2024-06-05 06:47:01 -04:00
Dave Horton d7797d5691 Merge pull request #77 from jambonz/feat/improve_elevenlabs
support elevenlabs previous_text, next_text
2024-06-05 06:43:48 -04:00
Quan HL b81f3313cb support elevenlabs previous_text, next_text 2024-06-05 15:53:26 +07:00
Dave Horton 49fe64744c 0.1.6 2024-05-30 07:13:00 -04:00
Dave Horton 106238cca5 bump version 2024-05-30 07:12:37 -04:00
Dave Horton 01735ccf2f Merge pull request #76 from jambonz/fix/verbio_cache
fixed verbio tts extension
2024-05-30 07:10:50 -04:00
Quan HL ce81cb8d24 fixed verbio tts extension 2024-05-30 16:13:05 +07:00
Dave Horton c9d7a9046f 0.1.4 2024-05-29 07:24:07 -04:00
Dave Horton a9ba383667 Merge pull request #75 from Catharsis68/feat/pre-commit-hook
Update eslint, add pre commit hook
2024-05-29 07:23:41 -04:00
Markus Frindt 06b97bcb0f Update eslint, add pre commit hook 2024-05-29 10:22:10 +02:00
Dave Horton 90152b36ef 0.1.3 2024-05-28 13:02:15 -04:00
Dave Horton 92484b7441 Merge pull request #74 from jambonz/fix/lint
lint
2024-05-28 13:01:35 -04:00
Dave Horton 398904785b lint 2024-05-28 12:58:37 -04:00
Dave Horton 531fa21f88 0.1.2 2024-05-28 12:55:00 -04:00
Dave Horton 904495d819 Merge pull request #73 from Catharsis68/feat/tts-cache-improvement
Improve handling of TTS cache by adding the file extension to the cac…
2024-05-28 12:54:31 -04:00
Markus Frindt f13fc84853 merge latest main into feature branch 2024-05-28 18:44:41 +02:00
Markus Frindt 39d54050cc simplify return for streaming responses 2024-05-28 15:23:53 +02:00
Markus Frindt 7618d334db add namespace for custom provider 2024-05-27 11:04:35 +02:00
Markus Frindt 71f20178d3 add test case 2024-05-24 14:09:07 +02:00
Markus Frindt cb6ab2479f Improve handling of TTS cache by adding the file extension to the cache key 2024-05-24 14:03:31 +02:00
35 changed files with 5838 additions and 12838 deletions
-1
View File
@@ -1 +0,0 @@
test/*
-126
View File
@@ -1,126 +0,0 @@
{
"env": {
"node": true,
"es6": true
},
"parserOptions": {
"ecmaFeatures": {
"jsx": false,
"modules": false
},
"ecmaVersion": 2020
},
"plugins": ["promise"],
"rules": {
"promise/always-return": "error",
"promise/no-return-wrap": "error",
"promise/param-names": "error",
"promise/catch-or-return": "error",
"promise/no-native": "off",
"promise/no-nesting": "warn",
"promise/no-promise-in-callback": "warn",
"promise/no-callback-in-promise": "warn",
"promise/no-return-in-finally": "warn",
// Possible Errors
// http://eslint.org/docs/rules/#possible-errors
"comma-dangle": [2, "only-multiline"],
"no-control-regex": 2,
"no-debugger": 2,
"no-dupe-args": 2,
"no-dupe-keys": 2,
"no-duplicate-case": 2,
"no-empty-character-class": 2,
"no-ex-assign": 2,
"no-extra-boolean-cast" : 2,
"no-extra-parens": [2, "functions"],
"no-extra-semi": 2,
"no-func-assign": 2,
"no-invalid-regexp": 2,
"no-irregular-whitespace": 2,
"no-negated-in-lhs": 2,
"no-obj-calls": 2,
"no-proto": 2,
"no-unexpected-multiline": 2,
"no-unreachable": 2,
"use-isnan": 2,
"valid-typeof": 2,
// Best Practices
// http://eslint.org/docs/rules/#best-practices
"no-fallthrough": 2,
"no-octal": 2,
"no-redeclare": 2,
"no-self-assign": 2,
"no-unused-labels": 2,
// Strict Mode
// http://eslint.org/docs/rules/#strict-mode
"strict": [2, "never"],
// Variables
// http://eslint.org/docs/rules/#variables
"no-delete-var": 2,
"no-undef": 2,
"no-unused-vars": [2, {"args": "none"}],
// Node.js and CommonJS
// http://eslint.org/docs/rules/#nodejs-and-commonjs
"no-mixed-requires": 2,
"no-new-require": 2,
"no-path-concat": 2,
"no-restricted-modules": [2, "sys", "_linklist"],
// Stylistic Issues
// http://eslint.org/docs/rules/#stylistic-issues
"comma-spacing": 2,
"eol-last": 2,
"indent": [2, 2, {"SwitchCase": 1}],
"keyword-spacing": 2,
"max-len": [2, 120, 2],
"new-parens": 2,
"no-mixed-spaces-and-tabs": 2,
"no-multiple-empty-lines": [2, {"max": 2}],
"no-trailing-spaces": [2, {"skipBlankLines": false }],
"quotes": [2, "single", "avoid-escape"],
"semi": 2,
"space-before-blocks": [2, "always"],
"space-before-function-paren": [2, "never"],
"space-in-parens": [2, "never"],
"space-infix-ops": 2,
"space-unary-ops": 2,
// ECMAScript 6
// http://eslint.org/docs/rules/#ecmascript-6
"arrow-parens": [2, "always"],
"arrow-spacing": [2, {"before": true, "after": true}],
"constructor-super": 2,
"no-class-assign": 2,
"no-confusing-arrow": 2,
"no-const-assign": 2,
"no-dupe-class-members": 2,
"no-new-symbol": 2,
"no-this-before-super": 2,
"prefer-const": 2
},
"globals": {
"DTRACE_HTTP_CLIENT_REQUEST" : false,
"LTTNG_HTTP_CLIENT_REQUEST" : false,
"COUNTER_HTTP_CLIENT_REQUEST" : false,
"DTRACE_HTTP_CLIENT_RESPONSE" : false,
"LTTNG_HTTP_CLIENT_RESPONSE" : false,
"COUNTER_HTTP_CLIENT_RESPONSE" : false,
"DTRACE_HTTP_SERVER_REQUEST" : false,
"LTTNG_HTTP_SERVER_REQUEST" : false,
"COUNTER_HTTP_SERVER_REQUEST" : false,
"DTRACE_HTTP_SERVER_RESPONSE" : false,
"LTTNG_HTTP_SERVER_RESPONSE" : false,
"COUNTER_HTTP_SERVER_RESPONSE" : false,
"DTRACE_NET_STREAM_END" : false,
"LTTNG_NET_STREAM_END" : false,
"COUNTER_NET_SERVER_CONNECTION_CLOSE" : false,
"DTRACE_NET_SERVER_CONNECTION" : false,
"LTTNG_NET_SERVER_CONNECTION" : false,
"COUNTER_NET_SERVER_CONNECTION" : false
}
}
+2 -4
View File
@@ -11,7 +11,7 @@ jobs:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: lts/*
node-version: '20'
- run: npm install
- run: npm run jslint
- run: sudo apt update && sudo apt install -y squid
@@ -23,9 +23,7 @@ jobs:
AWS_REGION: ${{ secrets.AWS_REGION }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
GCP_JSON_KEY: ${{ secrets.GCP_JSON_KEY }}
IBM_API_KEY: ${{ secrets.IBM_API_KEY }}
IBM_TTS_API_KEY: ${{ secrets.IBM_TTS_API_KEY }}
IBM_TTS_REGION: ${{ secrets.IBM_TTS_REGION }}
MICROSOFT_API_KEY: ${{ secrets.MICROSOFT_API_KEY }}
MICROSOFT_REGION: ${{ secrets.MICROSOFT_REGION }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
+2 -2
View File
@@ -12,8 +12,8 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- uses: actions/setup-node@v3
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: lts/*
registry-url: 'https://registry.npmjs.org'
+2
View File
@@ -39,3 +39,5 @@ node_modules
examples/*
.vscode
.env
+3
View File
@@ -0,0 +1,3 @@
#npm audit
npm run jslint:fix || true
npm test
-11
View File
@@ -1,11 +0,0 @@
#!/bin/sh
mkdir -p stubs/nuance
for FILE in ./protos/nuance/*; do
grpc_tools_node_protoc \
--js_out=import_style=commonjs,binary:./stubs/nuance \
--grpc_out=grpc_js:./stubs/nuance \
--proto_path=./protos/nuance \
$FILE
done
+140
View File
@@ -0,0 +1,140 @@
const promise = require('eslint-plugin-promise');
const globals = require('globals');
module.exports = [
{
files: ['**/*.js', '*.js'],
languageOptions: {
ecmaVersion: 2020,
parserOptions: {
'node': true,
'es6': true,
ecmaFeatures: {
jsx: false,
modules: false,
}
},
sourceType: 'module',
globals: {
...globals.node,
'DTRACE_HTTP_CLIENT_REQUEST': false,
'LTTNG_HTTP_CLIENT_REQUEST': false,
'COUNTER_HTTP_CLIENT_REQUEST': false,
'DTRACE_HTTP_CLIENT_RESPONSE': false,
'LTTNG_HTTP_CLIENT_RESPONSE': false,
'COUNTER_HTTP_CLIENT_RESPONSE': false,
'DTRACE_HTTP_SERVER_REQUEST': false,
'LTTNG_HTTP_SERVER_REQUEST': false,
'COUNTER_HTTP_SERVER_REQUEST': false,
'DTRACE_HTTP_SERVER_RESPONSE': false,
'LTTNG_HTTP_SERVER_RESPONSE': false,
'COUNTER_HTTP_SERVER_RESPONSE': false,
'DTRACE_NET_STREAM_END': false,
'LTTNG_NET_STREAM_END': false,
'COUNTER_NET_SERVER_CONNECTION_CLOSE': false,
'DTRACE_NET_SERVER_CONNECTION': false,
'LTTNG_NET_SERVER_CONNECTION': false,
'COUNTER_NET_SERVER_CONNECTION': false
},
},
'plugins': {
promise
},
'rules': {
'promise/always-return': 'error',
'promise/no-return-wrap': 'error',
'promise/param-names': 'error',
'promise/catch-or-return': 'error',
'promise/no-native': 'off',
'promise/no-nesting': 'warn',
'promise/no-promise-in-callback': 'warn',
'promise/no-callback-in-promise': 'warn',
'promise/no-return-in-finally': 'warn',
// Possible Errors
// http://eslint.org/docs/rules/#possible-errors
'comma-dangle': [2, 'only-multiline'],
'no-control-regex': 2,
'no-debugger': 2,
'no-dupe-args': 2,
'no-dupe-keys': 2,
'no-duplicate-case': 2,
'no-empty-character-class': 2,
'no-ex-assign': 2,
'no-extra-boolean-cast': 2,
'no-extra-parens': [2, 'functions'],
'no-extra-semi': 2,
'no-func-assign': 2,
'no-invalid-regexp': 2,
'no-irregular-whitespace': 2,
'no-negated-in-lhs': 2,
'no-obj-calls': 2,
'no-proto': 2,
'no-unexpected-multiline': 2,
'no-unreachable': 2,
'use-isnan': 2,
'valid-typeof': 2,
// Best Practices
// http://eslint.org/docs/rules/#best-practices
'no-fallthrough': 2,
'no-octal': 2,
'no-redeclare': 2,
'no-self-assign': 2,
'no-unused-labels': 2,
// Strict Mode
// http://eslint.org/docs/rules/#strict-mode
'strict': [2, 'never'],
// Variables
// http://eslint.org/docs/rules/#variables
'no-delete-var': 2,
'no-undef': 2,
'no-unused-vars': [2, { 'args': 'none' }],
// Node.js and CommonJS
// http://eslint.org/docs/rules/#nodejs-and-commonjs
'no-mixed-requires': 2,
'no-new-require': 2,
'no-path-concat': 2,
'no-restricted-modules': [2, 'sys', '_linklist'],
// Stylistic Issues
// http://eslint.org/docs/rules/#stylistic-issues
'comma-spacing': 2,
'eol-last': 2,
'indent': [2, 2, { 'SwitchCase': 1 }],
'keyword-spacing': 2,
'max-len': [2, 120, 2],
'new-parens': 2,
'no-mixed-spaces-and-tabs': 2,
'no-multiple-empty-lines': [2, { 'max': 2 }],
'no-trailing-spaces': [2, { 'skipBlankLines': false }],
'quotes': [2, 'single', 'avoid-escape'],
'semi': 2,
'space-before-blocks': [2, 'always'],
'space-before-function-paren': [2, 'never'],
'space-in-parens': [2, 'never'],
'space-infix-ops': 2,
'space-unary-ops': 2,
// ECMAScript 6
// http://eslint.org/docs/rules/#ecmascript-6
'arrow-parens': [2, 'always'],
'arrow-spacing': [2, { 'before': true, 'after': true }],
'constructor-super': 2,
'no-class-assign': 2,
'no-confusing-arrow': 2,
'no-const-assign': 2,
'no-dupe-class-members': 2,
'no-new-symbol': 2,
'no-this-before-super': 2,
'prefer-const': 2
},
'ignores': []
}
];
+1 -3
View File
@@ -14,9 +14,7 @@ module.exports = (opts, logger) => {
purgeTtsCache: require('./lib/purge-tts-cache').bind(null, client, logger),
addFileToCache: require('./lib/add-file-to-cache').bind(null, client, logger),
synthAudio: require('./lib/synth-audio').bind(null, client, createHash, retrieveHash, logger),
getVerbioAccessToken: require('./lib/get-verbio-token').bind(null, client, logger),
getNuanceAccessToken: require('./lib/get-nuance-access-token').bind(null, client, logger),
getIbmAccessToken: require('./lib/get-ibm-access-token').bind(null, client, logger),
getAwsAuthToken: require('./lib/get-aws-sts-token').bind(null, logger, createHash, retrieveHash),
getTtsVoices: require('./lib/get-tts-voices').bind(null, client, createHash, retrieveHash, logger),
};
+34 -3
View File
@@ -1,9 +1,31 @@
const fs = require('fs/promises');
const {noopLogger, makeSynthKey} = require('./utils');
const EXPIRES = (process.env.JAMBONES_TTS_CACHE_DURATION_MINS || 4 * 60) * 60; // cache tts for 4 hours
const {JAMBONES_TTS_CACHE_DURATION_MINS} = require('./config');
const EXPIRES = JAMBONES_TTS_CACHE_DURATION_MINS;
function getExtensionAndSampleRate(path) {
const match = path.match(/\.([^.]*)$/);
if (!match) {
//default should be wav file.
return ['wav', 8000];
}
const extension = match[1];
const sampleRateMap = {
r8: 8000,
r16: 16000,
r24: 24000,
r44: 44100,
r48: 48000,
r96: 96000,
};
const sampleRate = sampleRateMap[extension] || 8000;
return [extension, sampleRate];
}
async function addFileToCache(client, logger, path,
{account_sid, vendor, language, voice, deploymentId, engine, text}) {
{account_sid, vendor, language, voice, deploymentId, engine, model, text, instructions}) {
let key;
logger = logger || noopLogger;
@@ -14,10 +36,19 @@ async function addFileToCache(client, logger, path,
language: language || '',
voice: voice || deploymentId,
engine,
model,
text,
instructions
});
const [extension, sampleRate] = getExtensionAndSampleRate(path);
const audioBuffer = await fs.readFile(path);
await client.setex(key, EXPIRES, audioBuffer.toString('base64'));
await client.setex(key, EXPIRES, JSON.stringify(
{
audioContent: audioBuffer.toString('base64'),
extension,
sampleRate
}
));
} catch (err) {
logger.error(err, 'addFileToCache: Error');
return;
+27
View File
@@ -0,0 +1,27 @@
const JAMBONES_TTS_TRIM_SILENCE = process.env.JAMBONES_TTS_TRIM_SILENCE;
const JAMBONES_DISABLE_TTS_STREAMING = process.env.JAMBONES_DISABLE_TTS_STREAMING;
const JAMBONES_DISABLE_AZURE_TTS_STREAMING = process.env.JAMBONES_DISABLE_AZURE_TTS_STREAMING;
const JAMBONES_EAGERLY_PRE_CACHE_AUDIO = process.env.JAMBONES_EAGERLY_PRE_CACHE_AUDIO;
const JAMBONES_AZURE_ENABLE_SSML = process.env.JAMBONES_AZURE_ENABLE_SSML;
const JAMBONES_HTTP_PROXY_IP = process.env.JAMBONES_HTTP_PROXY_IP;
const JAMBONES_HTTP_PROXY_PORT = process.env.JAMBONES_HTTP_PROXY_PORT;
const JAMBONES_TTS_CACHE_DURATION_MINS =
(parseInt(process.env.JAMBONES_TTS_CACHE_DURATION_MINS) || 4 * 60) * 60; // cache tts for 4 hours
const TMP_FOLDER = process.env.JAMBONES_TMP_FOLDER || '/tmp';
const HTTP_TIMEOUT = 5000;
module.exports = {
JAMBONES_TTS_TRIM_SILENCE,
JAMBONES_DISABLE_TTS_STREAMING,
JAMBONES_DISABLE_AZURE_TTS_STREAMING,
JAMBONES_HTTP_PROXY_IP,
JAMBONES_HTTP_PROXY_PORT,
JAMBONES_TTS_CACHE_DURATION_MINS,
JAMBONES_EAGERLY_PRE_CACHE_AUDIO,
TMP_FOLDER,
HTTP_TIMEOUT,
JAMBONES_AZURE_ENABLE_SSML
};
-3
View File
@@ -1,3 +0,0 @@
module.exports = {
HTTP_TIMEOUT: 5000
};
+38 -14
View File
@@ -1,35 +1,58 @@
const { STSClient, GetSessionTokenCommand, AssumeRoleCommand } = require('@aws-sdk/client-sts');
const {makeAwsKey, noopLogger} = require('./utils');
const debug = require('debug')('jambonz:speech-utils');
const EXPIRY = 3600;
const EXPIRY = process.env.AWS_STS_SESSION_DURATION || 3600;
// by default reset aws session before expiry time 10 mins
const CACHE_EXPIRY = process.env.AWS_STS_SESSION_RESET_EXPIRY || (EXPIRY - 600);
async function getAwsAuthToken(
logger, createHash, retrieveHash,
awsAccessKeyId, awsSecretAccessKey, awsRegion, roleArn = null) {
{speech_credential_sid, accessKeyId, secretAccessKey, region, roleArn}) {
logger = logger || noopLogger;
try {
const key = makeAwsKey(roleArn || awsAccessKeyId);
// if incase instance profile is used, speech_credential_sid will be used as key to lookup cache
const key = makeAwsKey(roleArn || accessKeyId || speech_credential_sid);
const obj = await retrieveHash(key);
if (obj) return {...obj, servedFromCache: true};
/* access token not found in cache, so generate it using STS */
let data;
let expiry = CACHE_EXPIRY;
if (roleArn) {
const stsClient = new STSClient({ region: awsRegion});
const stsClient = new STSClient({ region });
const roleToAssume = { RoleArn: roleArn, RoleSessionName: 'Jambonz_Speech', DurationSeconds: EXPIRY};
const command = new AssumeRoleCommand(roleToAssume);
data = await stsClient.send(command);
} else {
/* access token not found in cache, so generate it using STS */
} else if (accessKeyId) {
const stsClient = new STSClient({
region: awsRegion,
region,
credentials: {
accessKeyId: awsAccessKeyId,
secretAccessKey: awsSecretAccessKey,
accessKeyId,
secretAccessKey,
}
});
const command = new GetSessionTokenCommand({DurationSeconds: EXPIRY});
data = await stsClient.send(command);
} else {
// instance profile is used.
const stsClient = new STSClient({ region });
const cred = await stsClient.config.credentials();
// method in the AWS SDK automatically fetches credentials using the default credential
// provider chain. If the credentials come from an instance profile or an environment
// variable, their expiration is controlled by AWS and not explicitly by our code.
if (cred && cred.expiration) {
const currentTime = new Date();
const expiryTime = new Date(cred.expiration);
const remainingTimeInSeconds = Math.round((expiryTime - currentTime) / 1000);
expiry = remainingTimeInSeconds;
}
data = {
Credentials: {
AccessKeyId: cred.accessKeyId,
SecretAccessKey: cred.secretAccessKey,
SessionToken: cred.sessionToken
}
};
}
const credentials = {
@@ -38,10 +61,11 @@ async function getAwsAuthToken(
sessionToken: data.Credentials.SessionToken,
securityToken: data.Credentials.SessionToken
};
/* expire 10 minutes before the hour, so we don't lose the use of it during a call */
createHash(key, credentials, EXPIRY - 600)
.catch((err) => logger.error(err, `Error saving hash for key ${key}`));
// Only cache if expiry is good
if (expiry > 0) {
createHash(key, credentials, expiry)
.catch((err) => logger.error(err, `Error saving hash for key ${key}`));
}
return {...credentials, servedFromCache: false};
} catch (err) {
-48
View File
@@ -1,48 +0,0 @@
const formurlencoded = require('form-urlencoded');
const {Pool} = require('undici');
const pool = new Pool('https://iam.cloud.ibm.com');
const {makeIbmKey, noopLogger} = require('./utils');
const { HTTP_TIMEOUT } = require('./constants');
const debug = require('debug')('jambonz:realtimedb-helpers');
async function getIbmAccessToken(client, logger, apiKey) {
logger = logger || noopLogger;
try {
const key = makeIbmKey(apiKey);
const access_token = await client.get(key);
if (access_token) return {access_token, servedFromCache: true};
/* access token not found in cache, so fetch it from Ibm */
const payload = {
grant_type: 'urn:ibm:params:oauth:grant-type:apikey',
apikey: apiKey
};
const {statusCode, headers, body} = await pool.request({
path: '/identity/token',
method: 'POST',
headers: {
'Content-Type': 'application/x-www-form-urlencoded'
},
body: formurlencoded(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
const json = await body.json();
logger.debug({statusCode, headers, body: json}, 'error fetching access token from Ibm');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
await client.set(key, json.access_token, 'EX', json.expires_in - 30);
return {...json, servedFromCache: false};
} catch (err) {
debug(err, 'getIbmAccessToken: Error retrieving Ibm access token');
logger.error(err, 'getIbmAccessToken: Error retrieving Ibm access token for client_id ${clientId}');
throw err;
}
}
module.exports = getIbmAccessToken;
-49
View File
@@ -1,49 +0,0 @@
const formurlencoded = require('form-urlencoded');
const {Pool} = require('undici');
const pool = new Pool('https://auth.crt.nuance.com');
const {makeNuanceKey, makeBasicAuthHeader, noopLogger} = require('./utils');
const { HTTP_TIMEOUT } = require('./constants');
const debug = require('debug')('jambonz:realtimedb-helpers');
async function getNuanceAccessToken(client, logger, clientId, secret, scope) {
logger = logger || noopLogger;
try {
const key = makeNuanceKey(clientId, secret, scope);
const access_token = await client.get(key);
if (access_token) return {access_token, servedFromCache: true};
/* access token not found in cache, so fetch it from Nuance */
const payload = {
grant_type: 'client_credentials',
scope
};
const auth = makeBasicAuthHeader(clientId, secret);
const {statusCode, headers, body} = await pool.request({
path: '/oauth2/token',
method: 'POST',
headers: {
...auth,
'Content-Type': 'application/x-www-form-urlencoded'
},
body: formurlencoded(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
logger.debug({statusCode, headers, body: body.text()}, 'error fetching access token from Nuance');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
await client.set(key, json.access_token, 'EX', json.expires_in - 30);
return {...json, servedFromCache: false};
} catch (err) {
debug(err, `getNuanceAccessToken: Error retrieving Nuance access token for client_id ${clientId}`);
logger.error(err, `getNuanceAccessToken: Error retrieving Nuance access token for client_id ${clientId}`);
throw err;
}
}
module.exports = getNuanceAccessToken;
+8 -112
View File
@@ -1,91 +1,8 @@
const assert = require('assert');
const {noopLogger, createNuanceClient, createKryptonClient} = require('./utils');
const getNuanceAccessToken = require('./get-nuance-access-token');
const getVerbioAccessToken = require('./get-verbio-token');
const {GetVoicesRequest, Voice} = require('../stubs/nuance/synthesizer_pb');
const TextToSpeechV1 = require('ibm-watson/text-to-speech/v1');
const { IamAuthenticator } = require('ibm-watson/auth');
const {noopLogger} = require('./utils');
const ttsGoogle = require('@google-cloud/text-to-speech');
const { PollyClient, DescribeVoicesCommand } = require('@aws-sdk/client-polly');
const getAwsAuthToken = require('./get-aws-sts-token');
const {Pool} = require('undici');
const { HTTP_TIMEOUT } = require('./constants');
const verbioVoicePool = new Pool('https://us.rest.speechcenter.verbio.com');
const getIbmVoices = async(client, logger, credentials) => {
const {tts_region, tts_api_key} = credentials;
console.log(`region: ${tts_region}, api_key: ${tts_api_key}`);
const textToSpeech = new TextToSpeechV1({
authenticator: new IamAuthenticator({
apikey: tts_api_key,
}),
serviceUrl: `https://api.${tts_region}.text-to-speech.watson.cloud.ibm.com`
});
const voices = await textToSpeech.listVoices();
return voices;
};
const getNuanceVoices = async(client, logger, credentials) => {
const {client_id: clientId, secret: secret, nuance_tts_uri} = credentials;
return new Promise(async(resolve, reject) => {
/* get a nuance access token */
let token, nuanceClient;
try {
if (nuance_tts_uri) {
nuanceClient = await createKryptonClient(nuance_tts_uri);
}
else {
const access_token = await getNuanceAccessToken(client, logger, clientId, secret, 'tts');
token = access_token.access_token;
nuanceClient = await createNuanceClient(token);
}
} catch (err) {
logger.error({err}, 'getTtsVoices: error retrieving access token');
return reject(err);
}
/* retrieve all voices */
const v = new Voice();
const request = new GetVoicesRequest();
request.setVoice(v);
nuanceClient.getVoices(request, (err, response) => {
if (err) {
logger.error({err, clientId, secret, token}, 'getTtsVoices: error retrieving voices');
return reject(err);
}
/* return all the voices that are not restricted and eliminate duplicates */
const voices = response.getVoicesList()
.map((v) => {
return {
language: v.getLanguage(),
name: v.getName(),
model: v.getModel(),
gender: v.getGender() === 1 ? 'male' : 'female',
restricted: v.getRestricted()
};
});
const v = voices
.filter((v) => v.restricted === false)
.map((v) => {
delete v.restricted;
return v;
})
.sort((a, b) => {
if (a.language < b.language) return -1;
if (a.language > b.language) return 1;
if (a.name < b.name) return -1;
return 1;
});
const arr = [...new Set(v.map((v) => JSON.stringify(v)))]
.map((v) => JSON.parse(v));
resolve(arr);
});
});
};
const getGoogleVoices = async(_client, logger, credentials) => {
const client = new ttsGoogle.TextToSpeechClient({credentials});
@@ -107,7 +24,12 @@ const getAwsVoices = async(_client, createHash, retrieveHash, logger, credential
} else if (roleArn) {
client = new PollyClient({
region,
credentials: await getAwsAuthToken(logger, createHash, retrieveHash, null, null, region, roleArn),
credentials: await getAwsAuthToken(
logger, createHash, retrieveHash,
{
region,
roleArn
}),
});
} else {
client = new PollyClient({region});
@@ -121,26 +43,6 @@ const getAwsVoices = async(_client, createHash, retrieveHash, logger, credential
}
};
const getVerbioVoices = async(client, logger, credentials) => {
try {
const access_token = await getVerbioAccessToken(client, logger, credentials);
const { body} = await verbioVoicePool.request({
path: '/api/v1/voices',
method: 'GET',
headers: {
'Authorization': `Bearer ${access_token.access_token}`,
'User-Agent': 'jambonz'
},
timeout: HTTP_TIMEOUT,
followRedirects: false
});
return await body.json();
} catch (err) {
logger.info({err}, 'getVerbioVoices - failed to list voices for Verbio');
throw err;
}
};
/**
* Synthesize speech to an mp3 file, and also cache the generated speech
* in redis (base64 format) for 24 hours so as to avoid unnecessarily paying
@@ -160,21 +62,15 @@ const getVerbioVoices = async(client, logger, credentials) => {
async function getTtsVoices(client, createHash, retrieveHash, logger, {vendor, credentials}) {
logger = logger || noopLogger;
assert.ok(['nuance', 'ibm', 'google', 'aws', 'polly', 'verbio'].includes(vendor),
assert.ok(['google', 'aws', 'polly'].includes(vendor),
`getTtsVoices not supported for vendor ${vendor}`);
switch (vendor) {
case 'nuance':
return getNuanceVoices(client, logger, credentials);
case 'ibm':
return getIbmVoices(client, logger, credentials);
case 'google':
return getGoogleVoices(client, logger, credentials);
case 'aws':
case 'polly':
return getAwsVoices(client, createHash, retrieveHash, logger, credentials);
case 'verbio':
return getVerbioVoices(client, logger, credentials);
default:
break;
}
-51
View File
@@ -1,51 +0,0 @@
const {Pool} = require('undici');
const { noopLogger, makeVerbioKey } = require('./utils');
const { HTTP_TIMEOUT } = require('./constants');
const pool = new Pool('https://auth.speechcenter.verbio.com:444');
const debug = require('debug')('jambonz:realtimedb-helpers');
async function getVerbioAccessToken(client, logger, credentials) {
logger = logger || noopLogger;
const { client_id, client_secret } = credentials;
try {
const key = makeVerbioKey(client_id);
const access_token = await client.get(key);
if (access_token) {
return {access_token, servedFromCache: true};
}
const payload = {
client_id,
client_secret
};
const {statusCode, headers, body} = await pool.request({
path: '/api/v1/token',
method: 'POST',
headers: {
'Content-Type': 'application/json',
'User-Agent': 'jambonz'
},
body: JSON.stringify(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
logger.debug({statusCode, headers, body: await body.text()}, 'error fetching access token from Verbio');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
const expiry = Math.floor(json.expiration_time - Date.now() / 1000 - 30);
await client.set(key, json.access_token, 'EX', expiry);
return {...json, servedFromCache: false};
} catch (err) {
debug(err, `getVerbioAccessToken: Error retrieving Verbio access token for client_id ${client_id}`);
logger.error(err, `getVerbioAccessToken: Error retrieving Verbio access token for client_id ${client_id}`);
throw err;
}
}
module.exports = getVerbioAccessToken;
+3 -1
View File
@@ -12,7 +12,7 @@ const debug = require('debug')('jambonz:realtimedb-helpers');
* @returns {object} result - {error, purgedCount}
*/
async function purgeTtsCache(client, logger, {all, account_sid, vendor,
language, voice, deploymentId, engine, text} = {all: true}) {
language, voice, deploymentId, engine, model, text, instructions} = {all: true}) {
logger = logger || noopLogger;
let purgedCount = 0, error;
@@ -33,7 +33,9 @@ async function purgeTtsCache(client, logger, {all, account_sid, vendor,
language: language || '',
voice: voice || deploymentId,
engine,
model,
text,
instructions
});
purgedCount = await client.del(key);
if (purgedCount === 0) error = 'Specified item not found';
+848 -349
View File
File diff suppressed because it is too large Load Diff
+23 -105
View File
@@ -1,118 +1,43 @@
const crypto = require('crypto');
const {SynthesizerClient} = require('../stubs/nuance/synthesizer_grpc_pb');
const {RivaSpeechSynthesisClient} = require('../stubs/riva/proto/riva_tts_grpc_pb');
const {Pool} = require('undici');
const pool = new Pool('https://auth.crt.nuance.com');
const NUANCE_AUTH_ENDPOINT = 'tts.api.nuance.com:443';
const grpc = require('@grpc/grpc-js');
const formurlencoded = require('form-urlencoded');
const { HTTP_TIMEOUT } = require('./constants');
const { TMP_FOLDER } = require('./config');
const debug = require('debug')('jambonz:realtimedb-helpers');
/**
* Future TODO: cache recently used connections to providers
* to avoid connection overhead during a call.
* Will need to periodically age them out to avoid memory leaks.
*/
//const nuanceClientMap = new Map();
function makeSynthKey({account_sid = '', vendor, language, voice, engine = '', text}) {
function makeSynthKey({
account_sid = '',
vendor,
language,
voice,
engine = '',
model = '',
text,
instructions = '',
}) {
const hash = crypto.createHash('sha1');
hash.update(`${language}:${vendor}:${voice}:${engine}:${text}`);
return `tts${account_sid ? (':' + account_sid) : ''}:${hash.digest('hex')}`;
hash.update(`${language}:${vendor}:${voice}:${engine}:${model}:${text}:${instructions}`);
const hexHashKey = hash.digest('hex');
const accountKey = account_sid ? `:${account_sid}` : '';
const key = `tts${accountKey}:${hexHashKey}`;
return key;
}
function makeFilePath({key, salt = '', extension}) {
return `${TMP_FOLDER}/${key.replace('tts:', `tts-${salt}`)}.${extension}`;
}
const noopLogger = {
info: () => {},
debug: () => {},
error: () => {}
};
const toBase64 = (str) => Buffer.from(str || '', 'utf8').toString('base64');
function makeBasicAuthHeader(username, password) {
if (!username || !password) return {};
const creds = `${encodeURIComponent(username)}:${password || ''}`;
const header = `Basic ${toBase64(creds)}`;
return {Authorization: header};
}
function makeIbmKey(apiKey) {
const hash = crypto.createHash('sha1');
hash.update(apiKey);
return `ibm:${hash.digest('hex')}`;
}
function makeAwsKey(awsAccessKeyId) {
const hash = crypto.createHash('sha1');
hash.update(awsAccessKeyId);
return `aws:${hash.digest('hex')}`;
}
function makeVerbioKey(client_id) {
const hash = crypto.createHash('sha1');
hash.update(client_id);
return `verbio:${hash.digest('hex')}`;
}
function makeNuanceKey(clientId, secret, scope) {
const hash = crypto.createHash('sha1');
hash.update(`${clientId}:${secret}:${scope}`);
return `nuance:${hash.digest('hex')}`;
}
const getNuanceAccessToken = async(clientId, secret, scope = 'asr tts') => {
const payload = {
grant_type: 'client_credentials',
scope
};
const auth = makeBasicAuthHeader(clientId, secret);
const {statusCode, headers, body} = await pool.request({
path: '/oauth2/token',
method: 'POST',
headers: {
...auth,
'Content-Type': 'application/x-www-form-urlencoded'
},
body: formurlencoded(payload),
timeout: HTTP_TIMEOUT,
followRedirects: false
});
if (200 !== statusCode) {
debug({statusCode, headers, body: body.text()}, 'error fetching access token from Nuance');
const err = new Error();
err.statusCode = statusCode;
throw err;
}
const json = await body.json();
return json.access_token;
};
const createKryptonClient = async(uri) => {
const client = new SynthesizerClient(uri, grpc.credentials.createInsecure());
return client;
};
const createNuanceClient = async(access_token) => {
//if (nuanceClientMap.has(access_token)) return nuanceClientMap.get(access_token);
const generateMetadata = (params, callback) => {
var metadata = new grpc.Metadata();
metadata.add('authorization', `Bearer ${access_token}`);
callback(null, metadata);
};
const sslCreds = grpc.credentials.createSsl();
const authCreds = grpc.credentials.createFromMetadataGenerator(generateMetadata);
const combined_creds = grpc.credentials.combineChannelCredentials(sslCreds, authCreds);
const client = new SynthesizerClient(NUANCE_AUTH_ENDPOINT, combined_creds);
//if (process.env.NUANCE_CACHE_TTS_CONNECTIONS) nuanceClientMap.set(access_token, client);
return client;
};
const createRivaClient = async(rivaUri) => {
const client = new RivaSpeechSynthesisClient(rivaUri, grpc.credentials.createInsecure());
return client;
@@ -120,15 +45,8 @@ const createRivaClient = async(rivaUri) => {
module.exports = {
makeSynthKey,
makeNuanceKey,
makeIbmKey,
makeAwsKey,
makeVerbioKey,
getNuanceAccessToken,
createNuanceClient,
createKryptonClient,
createRivaClient,
makeBasicAuthHeader,
NUANCE_AUTH_ENDPOINT,
noopLogger
noopLogger,
makeFilePath
};
+4044 -4176
View File
File diff suppressed because it is too large Load Diff
+15 -13
View File
@@ -1,6 +1,6 @@
{
"name": "@jambonz/speech-utils",
"version": "0.1.1",
"version": "1.0.9",
"description": "TTS-related speech utilities for jambonz",
"main": "index.js",
"author": "Dave Horton",
@@ -12,7 +12,8 @@
"test": "NODE_ENV=test node test/ ",
"coverage": "nyc --reporter html --report-dir ./coverage npm run test",
"jslint": "eslint index.js lib",
"build": "./build_stubs.sh"
"jslint:fix": "npm run jslint --fix",
"prepare": "husky"
},
"repository": {
"type": "git",
@@ -24,26 +25,27 @@
},
"homepage": "https://github.com/jambonz/speech-utils#readme",
"dependencies": {
"23": "^0.0.0",
"@aws-sdk/client-polly": "^3.496.0",
"@aws-sdk/client-sts": "^3.496.0",
"@google-cloud/text-to-speech": "^5.0.2",
"@cartesia/cartesia-js": "^2.2.7",
"@google-cloud/text-to-speech": "^6.4.0",
"@grpc/grpc-js": "^1.9.14",
"@jambonz/realtimedb-helpers": "^0.8.7",
"bent": "^7.3.12",
"debug": "^4.3.4",
"form-urlencoded": "^6.1.4",
"google-protobuf": "^3.21.2",
"ibm-watson": "^8.0.0",
"microsoft-cognitiveservices-speech-sdk": "1.36.0",
"openai": "^4.25.0",
"undici": "^6.4.0"
"microsoft-cognitiveservices-speech-sdk": "1.38.0",
"openai": "^4.98.0",
"undici": "^7.5.0"
},
"devDependencies": {
"config": "^3.3.10",
"eslint": "^8.56.0",
"eslint-plugin-promise": "^6.1.1",
"config": "^4.2.0",
"eslint": "^9.3.0",
"eslint-plugin-promise": "^6.2.0",
"husky": "^9.0.11",
"nyc": "^15.1.0",
"pino": "^8.17.0",
"tape": "^5.7.3"
"pino": "^9.1.0",
"tape": "^5.7.5"
}
}
-332
View File
@@ -1,332 +0,0 @@
syntax = "proto3";
package nuance.tts.v1;
/*
The Synthesizer service offers these functionalities:
* - GetVoices: Queries the list of available voices, with filters to reduce the search space.
* - Synthesize: Synthesizes audio from input text and parameters, and returns an audio stream.
* - UnarySynthesize: Synthesizes audio from input text and parameters, and returns a single audio response.
*/
service Synthesizer {
rpc GetVoices(GetVoicesRequest) returns (GetVoicesResponse) {}
rpc Synthesize(SynthesisRequest) returns (stream SynthesisResponse) {}
rpc UnarySynthesize(SynthesisRequest) returns (UnarySynthesisResponse) {}
}
/*
Input message for [Synthesizer](#synthesizer) - GetVoices, to query voices available to the client.
*/
message GetVoicesRequest {
Voice voice = 1; // Optionally filter the voices to retrieve, e.g. set language to en-US to return only American English voices.
}
/*
Output message for [Synthesizer](#synthesizer) - GetVoices. Includes a list of voices that matched the input criteria, if any.
*/
message GetVoicesResponse {
repeated Voice voices = 1; // Repeated. Voices and characteristics returned.
}
/*
Input message for [Synthesizer](#synthesizer) - Synthesize. Specifies input text, audio parameters, and events to subscribe to, in exchange for synthesized audio.
*/
message SynthesisRequest {
Voice voice = 1; // The voice to use for audio synthesis. Mandatory.
AudioParameters audio_params = 2; // Output audio parameters, such as encoding and volume.
Input input = 3; // Input text to synthesize, tuning data, etc. Mandatory.
EventParameters event_params = 4; // Markers and other info to include in server events returned during synthesis.
map<string, string> client_data = 5; // Repeated. Client-supplied key-value pairs to inject into the call log.
string user_id = 6; // Identifies a particular user within an application.
}
/*
Input or output message for voices. When sent as input:
* - In [GetVoicesRequest](#getvoicesrequest), it filters the list of available voices.
* - In [SynthesisRequest](#synthesisrequest), it specifies the voice to use for synthesis.
When received as output in [GetVoicesResponse](#getvoicesresponse), it returns the list of available voices.
*/
message Voice {
string name = 1; // The voice's name, e.g. 'Evan'. Mandatory for SynthesisRequest.
string model = 2; // The voice's quality model, e.g. 'enhanced' or 'standard'. Mandatory for SynthesisRequest.
string language = 3; // IETF language code, e.g. 'en-US'. Used only for GetVoicesRequest and GetVoicesResponse, to search for voices with a certain mother tongue. Ignored otherwise.
EnumAgeGroup age_group = 4; // Used only for GetVoicesRequest and GetVoicesResponse, to search for adult or child voices. Ignored otherwise.
EnumGender gender = 5; // Used only for GetVoicesRequest and GetVoicesResponse, to search for voices with a certain gender. Ignored otherwise.
uint32 sample_rate_hz = 6; // Used only for GetVoicesRequest and GetVoicesResponse, to search for a certain native sample rate. Ignored otherwise.
string language_tlw = 7; // Used only for GetVoicesRequest and GetVoicesResponse. Three-letter language code (e.g. 'enu' for American English), for configuring language identification in [Input](#input).
bool restricted = 8; // Used only in GetVoicesResponse, to identify restricted voices. These are custom voices available only to specific customers. Ignored otherwise.
string version = 9; // Used only in GetVoicesResponse, to return the voice's version. Ignored otherwise.
}
/*
Input or output field specifying whether the voice uses its adult or child version, if available. Included in [Voice](#voice).
*/
enum EnumAgeGroup {
ADULT = 0; // Adult voice. Default for GetVoicesRequest.
CHILD = 1; // Child voice.
}
/*
Input or output field, specifying gender for voices that support multiple genders. Included in [Voice](#voice).
*/
enum EnumGender {
ANY = 0; // Any gender voice. Default for GetVoicesRequest.
MALE = 1; // Male voice.
FEMALE = 2; // Female voice.
NEUTRAL = 3; // Neutral gender voice.
}
/*
Input message for audio-related parameters during synthesis, including encoding, volume, and audio length. Included in [SynthesisRequest](#synthesisrequest).
*/
message AudioParameters {
AudioFormat audio_format = 1; // Audio encoding. Default PCM 22.5kHz.
uint32 volume_percentage = 2; // Volume amplitude, from 0 to 100. Default 80.
float speaking_rate_factor = 3; // Speaking rate, from 0 to 2.0. Default 1.0.
uint32 audio_chunk_duration_ms = 4; // Maximum duration, in ms, of an audio chunk delivered to the client, from 1 to 60000. Default is 20000 (20 seconds). When this parameter is large enough (for example, 20 or 30 seconds), each audio chunk contains an audible segment surrounded by silence.
uint32 target_audio_length_ms = 5; // Maximum duration, in ms, of synthesized audio. When greater than 0, the server stops ongoing synthesis at the first sentence end, or silence, closest to the value.
bool disable_early_emission = 6; // By default, audio segments are emitted as soon as possible, even if they are not audible. This behavior may be disabled.
}
/*
Input message for audio encoding of synthesized text. Included in [AudioParameters](#audioparameters).
*/
message AudioFormat {
oneof audio_format {
PCM pcm = 1; // Signed 16-bit little endian PCM.
ALaw alaw = 2; // G.711 A-law, 8kHz.
ULaw ulaw = 3; // G.711 Mu-law, 8kHz.
OggOpus ogg_opus = 4; // Ogg Opus, 8kHz, 16kHz or 24kHz.
Opus opus = 5; // Opus, 8kHz, 16kHZ or 24 kHz. The audio will be sent one Opus packet at a time.
}
}
/*
Input message defining PCM sample rate. Included in [Audioformat](#audioformat).
*/
message PCM {
uint32 sample_rate_hz = 1; // Output sample rate: 8000, 16000, 22050 (default), 24000.
}
/*
Input message defining A-law audio format. G.711 audio formats are set to 8kHz. Included in [Audioformat](#audioformat).
*/
message ALaw {}
/*
Input message defining Mu-law audio format. G.711 audio formats are set to 8kHz. Included in [Audioformat](#audioformat).
*/
message ULaw {}
/*
Input message defining Opus output rate. Included in [Audioformat](#audioformat).
*/
message Opus {
uint32 sample_rate_hz = 1; // Output sample rate. Supported values: 8000, 16000, 24000 Hz.
uint32 bit_rate_bps = 2; // Valid range is 500 to 256000 bps. Default 28000 bps.
float max_frame_duration_ms = 3; // Opus frame size, in ms: 2.5, 5, 10, 20, 40, 60. Default 20.
uint32 complexity = 4; // Computational complexity. A complexity of 0 means the codec default.
EnumVariableBitrate vbr = 5; // Variable bitrate. On by default.
}
/*
Input message defining Ogg Opus output rate. Included in [Audioformat](#audioformat).
*/
message OggOpus {
uint32 sample_rate_hz = 1; // Output sample rate. Supported values: 8000, 16000, 24000 Hz.
uint32 bit_rate_bps = 2; // Valid range is 500 to 256000 bps. Default 28000 bps.
float max_frame_duration_ms = 3; // Opus frame size, in ms: 2.5, 5, 10, 20, 40, 60. Default 20.
uint32 complexity = 4; // Computational complexity. A complexity of 0 means the codec default.
EnumVariableBitrate vbr = 5; // Variable bitrate. On by default.
}
/*
Settings for variable bitrate. Included in [OggOpus](#oggopus). Turned on by default.
*/
enum EnumVariableBitrate {
VARIABLE_BITRATE_ON = 0; // Use variable bitrate. Default.
VARIABLE_BITRATE_OFF = 1; // Do not use variable bitrate.
VARIABLE_BITRATE_CONSTRAINED = 2; // Use constrained variable bitrate.
}
/*
Input message containing text to synthesize and synthesis parameters, including tuning data, etc. Included in [SynthesisRequest](#synthesisrequest). The type of input may be:
* - Plain text.
* - An SSML document.
* - An alternating sequence of plain text and Nuance control codes.
*/
message Input {
oneof input_data {
Text text = 1; // Text input.
SSML ssml = 2; // SSML input.
TokenizedSequence tokenized_sequence = 3; // Sequence of text and Nuance control codes.
}
repeated SynthesisResource resources = 4; // Repeated. Synthesis resources (user dictionaries, rulesets, etc.) to tune synthesized audio.
LanguageIdentificationParameters lid_params = 5; // LID parameters.
DownloadParameters download_params = 6; // Remote file download parameters.
}
/*
Input message for synthesizing plain text. The encoding must be in UTF-8.
*/
message Text {
oneof text_data {
string text = 1; // Plain input text in UTF-8 encoding.
string uri = 2; // Remote URI to the plain input text. Disabled for Nuance-hosted TTS.
}
}
/*
Input message for synthesizing an SSML document.
*/
message SSML {
oneof ssml_data {
string text = 1; // SSML input text.
string uri = 2; // Remote URI to the SSML input text. Disabled for Nuance-hosted TTS.
}
EnumSSMLValidationMode ssml_validation_mode = 3; // SSML validation mode. Default STRICT.
}
/*
Input message for synthesizing a sequence of plain text and Nuance control codes.
*/
message TokenizedSequence {
repeated Token tokens = 1;
}
/*
The unit when using TokenizedSequence for input. Each token can either be plain text or a Nuance control code.
*/
message Token {
oneof token_data {
string text = 1; // Plain input text.
ControlCode control_code = 2; // Nuance control code.
}
}
/*
A Nuance control code allows the user to control how text is spoken, similarly to SSML.
*/
message ControlCode {
string key = 1; // Name of the control code, e.g. "pause".
string value = 2; // Value of the control code.
}
/*
Input message specifying the type of file to tune the synthesized output and its location or contents. Included in [Input](#input).
*/
message SynthesisResource {
EnumResourceType type = 1; // Resource type, e.g. user dictionary, etc. Default USER_DICTIONARY.
oneof resource_data {
string uri = 2; // URI to the remote resource, or
bytes body = 3; // For EnumResourceType USER_DICTIONARY, the contents of the file.
}
}
/*
The type of synthesis resource to tune the output. Included in [SynthesisResource](#synthesisresource). User dictionaries provide custom pronunciations, rulesets apply search-and-replace rules to input text and ActivePrompt databases help tune synthesized audio under certain conditions, using Nuance Vocalizer Studio.
*/
enum EnumResourceType {
USER_DICTIONARY = 0; // User dictionary (application/edct-bin-dictionary). Default.
TEXT_USER_RULESET = 1; // Text user ruleset (application/x-vocalizer-rettt+text).
BINARY_USER_RULESET = 2; // Binary user ruleset (application/x-vocalizer-rettt+bin).
ACTIVEPROMPT_DB = 3; // ActivePrompt database (application/x-vocalizer/activeprompt-db).
ACTIVEPROMPT_DB_AUTO = 4; // ActivePrompt database with automatic insertion (application/x-vocalizer/activeprompt-db;mode=automatic).
SYSTEM_DICTIONARY = 5; // Nuance system dictionary (application/sdct-bin-dictionary).
}
/*
SSML validation mode when using SSML input. Included in [Input](#input). Strict by default but can be relaxed.
*/
enum EnumSSMLValidationMode {
STRICT = 0; // Strict SSL validation. Default.
WARN = 1; // Give warning only.
NONE = 2; // Do not validate.
}
/*
Input message controlling the language identifier. Included in [Input](#input). The language identifier runs on input blocks labeled with the <ESC>\lang=unknown\ control sequence or SSML xml:lang="unknown". The language identifier automatically restricts the matched languages to the installed voices. This limits the permissible languages, and also sets the order of precedence (first to last) when they have equal confidence scores.
*/
message LanguageIdentificationParameters {
bool disable = 1; // Whether to disable language identification. Turned on by default.
repeated string languages = 2; // Repeated. List of three-letter language codes (e.g. enu, frc, spm) to restrict language identification results, in order of precedence. Use GetVoicesRequest - Voice - language_tlw to obtain the three-letter codes. Default blank.
bool always_use_highest_confidence = 3; // If enabled, language identification always chooses the language with the highest confidence score, even if the score is low. Default false, meaning use language with any confidence.
}
/*
Input message containing parameters for remote file download, whether for input text (Input.uri) or a SynthesisResource (SynthesisResource.uri). Included in [Input](#input).
*/
message DownloadParameters {
map <string,string> headers = 1; // HTTP headers to include in outgoing requests. Only whitelisted headers will actually be sent.
bool refuse_cookies = 2; // Whether to disable cookies. By default, HTTP requests accept cookies.
oneof optional_download_parameter_request_timeout_ms {
uint32 request_timeout_ms = 3; // Request timeout in ms. Default (0) means server default, usually 30000 (30 seconds).
}
}
/*
Input message that defines event subscription parameters. Included in [SynthesisRequest](#synthesisrequest). Events that are requested are sent throughout the SynthesisResponse stream, when generated. Marker events can send events as certain parts of the synthesized audio are reached, for example, at the end of a word, sentence, or user-defined bookmark.
* Log events are produced throughout a synthesis request for events such as a voice loaded by the server or an audio chunk being ready to send.
*/
message EventParameters {
bool send_sentence_marker_events = 1; // Sentence marker. Default: do not send.
bool send_word_marker_events = 2; // Word marker. Default: do not send.
bool send_phoneme_marker_events = 3; // Phoneme marker. Default: do not send.
bool send_bookmark_marker_events = 4; // Bookmark marker. Default: do not send.
bool send_paragraph_marker_events = 5; // Paragraph marker. Default: do not send.
bool send_visemes = 6; // Lipsync information. Default: do not send.
bool send_log_events = 7; // Whether to log events during synthesis. By default, logging is turned off.
bool suppress_input = 8; // Whether to omit input text and URIs from log events. By default, these items are included.
}
/*
The [Synthesizer](#synthesizer) - Synthesize RPC call returns a stream of SynthesisResponse messages. The response contains one of:
* - A status response, indicating completion or failure of the request. This is received only once and signifies the end of a Synthesize call.
* - A list of events the client has requested. This can be received many times. See EventParameters for details.
* - An audio buffer. This may be received many times.
*/
message SynthesisResponse {
oneof response {
Status status = 1; // A status response, indicating completion or failure of the request.
Events events = 2; // A list of events. See EventParameters for details.
bytes audio = 3; // The latest audio buffer.
}
}
/*
The [Synthesizer](#synthesizer) - UnarySynthesize RPC call returns a single UnarySynthesisResponse message. It is similar to a SynthesisResponse message but includes all the information instead of a single type of response. The response contains:
* - A status response, indicating completion or failure of the request.
* - A list of events the client has requested. See EventParameters for details.
* - The complete audio buffer of the synthesized text.
*/
message UnarySynthesisResponse {
Status status = 1; // A status response, indicating completion or failure of the request.
Events events = 2; // A list of events. See EventParameters for details.
bytes audio = 3; // Audio buffer of the synthesized text.
}
/*
Output message containing a status response, indicating completion or failure of a SynthesisRequest. Included in [SynthesisResponse](#synthesisresponse) and [UnarySynthesisResponse](#unarysynthesisresponse).
*/
message Status {
uint32 code = 1; // HTTP-style return code: 200, 4xx, or 5xx as appropriate.
string message = 2; // Brief description of the status.
string details = 3; // Longer description if available.
}
/*
Output message defining a container for a list of events. This container is needed because oneof does not allow repeated parameters in Protobuf. Included in [SynthesisResponse](#synthesisresponse) and [UnarySynthesisResponse](#unarysynthesisresponse).
*/
message Events {
repeated Event events = 1; // Repeated. One or more events.
}
/*
Output message defining an event message. Included in [Events](#events). See EventParameters for details.
*/
message Event {
string name = 1; // Either "Markers" or the name of the event in the case of a Log Event.
map<string, string> values = 2; // Repeated. Key-value data relevant to the current event.
}
-104
View File
@@ -1,104 +0,0 @@
// GENERATED CODE -- DO NOT EDIT!
'use strict';
var grpc = require('@grpc/grpc-js');
var synthesizer_pb = require('./synthesizer_pb.js');
function serialize_nuance_tts_v1_GetVoicesRequest(arg) {
if (!(arg instanceof synthesizer_pb.GetVoicesRequest)) {
throw new Error('Expected argument of type nuance.tts.v1.GetVoicesRequest');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_GetVoicesRequest(buffer_arg) {
return synthesizer_pb.GetVoicesRequest.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_GetVoicesResponse(arg) {
if (!(arg instanceof synthesizer_pb.GetVoicesResponse)) {
throw new Error('Expected argument of type nuance.tts.v1.GetVoicesResponse');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_GetVoicesResponse(buffer_arg) {
return synthesizer_pb.GetVoicesResponse.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_SynthesisRequest(arg) {
if (!(arg instanceof synthesizer_pb.SynthesisRequest)) {
throw new Error('Expected argument of type nuance.tts.v1.SynthesisRequest');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_SynthesisRequest(buffer_arg) {
return synthesizer_pb.SynthesisRequest.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_SynthesisResponse(arg) {
if (!(arg instanceof synthesizer_pb.SynthesisResponse)) {
throw new Error('Expected argument of type nuance.tts.v1.SynthesisResponse');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_SynthesisResponse(buffer_arg) {
return synthesizer_pb.SynthesisResponse.deserializeBinary(new Uint8Array(buffer_arg));
}
function serialize_nuance_tts_v1_UnarySynthesisResponse(arg) {
if (!(arg instanceof synthesizer_pb.UnarySynthesisResponse)) {
throw new Error('Expected argument of type nuance.tts.v1.UnarySynthesisResponse');
}
return Buffer.from(arg.serializeBinary());
}
function deserialize_nuance_tts_v1_UnarySynthesisResponse(buffer_arg) {
return synthesizer_pb.UnarySynthesisResponse.deserializeBinary(new Uint8Array(buffer_arg));
}
//
// The Synthesizer service offers these functionalities:
// - GetVoices: Queries the list of available voices, with filters to reduce the search space.
// - Synthesize: Synthesizes audio from input text and parameters, and returns an audio stream.
// - UnarySynthesize: Synthesizes audio from input text and parameters, and returns a single audio response.
var SynthesizerService = exports.SynthesizerService = {
getVoices: {
path: '/nuance.tts.v1.Synthesizer/GetVoices',
requestStream: false,
responseStream: false,
requestType: synthesizer_pb.GetVoicesRequest,
responseType: synthesizer_pb.GetVoicesResponse,
requestSerialize: serialize_nuance_tts_v1_GetVoicesRequest,
requestDeserialize: deserialize_nuance_tts_v1_GetVoicesRequest,
responseSerialize: serialize_nuance_tts_v1_GetVoicesResponse,
responseDeserialize: deserialize_nuance_tts_v1_GetVoicesResponse,
},
synthesize: {
path: '/nuance.tts.v1.Synthesizer/Synthesize',
requestStream: false,
responseStream: true,
requestType: synthesizer_pb.SynthesisRequest,
responseType: synthesizer_pb.SynthesisResponse,
requestSerialize: serialize_nuance_tts_v1_SynthesisRequest,
requestDeserialize: deserialize_nuance_tts_v1_SynthesisRequest,
responseSerialize: serialize_nuance_tts_v1_SynthesisResponse,
responseDeserialize: deserialize_nuance_tts_v1_SynthesisResponse,
},
unarySynthesize: {
path: '/nuance.tts.v1.Synthesizer/UnarySynthesize',
requestStream: false,
responseStream: false,
requestType: synthesizer_pb.SynthesisRequest,
responseType: synthesizer_pb.UnarySynthesisResponse,
requestSerialize: serialize_nuance_tts_v1_SynthesisRequest,
requestDeserialize: deserialize_nuance_tts_v1_SynthesisRequest,
responseSerialize: serialize_nuance_tts_v1_UnarySynthesisResponse,
responseDeserialize: deserialize_nuance_tts_v1_UnarySynthesisResponse,
},
};
exports.SynthesizerClient = grpc.makeGenericClientConstructor(SynthesizerService);
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1 +1 @@
// GENERATED CODE -- NO SERVICES IN PROTO
// GENERATED CODE -- NO SERVICES IN PROTO
+5 -5
View File
@@ -57,7 +57,7 @@ function deserialize_nvidia_riva_tts_SynthesizeSpeechResponse(buffer_arg) {
var RivaSpeechSynthesisService = exports.RivaSpeechSynthesisService = {
// Used to request text-to-speech from the service. Submit a request containing the
// desired text and configuration, and receive audio bytes in the requested format.
synthesize: {
synthesize: {
path: '/nvidia.riva.tts.RivaSpeechSynthesis/Synthesize',
requestStream: false,
responseStream: false,
@@ -69,9 +69,9 @@ synthesize: {
responseDeserialize: deserialize_nvidia_riva_tts_SynthesizeSpeechResponse,
},
// Used to request text-to-speech returned via stream as it becomes available.
// Submit a SynthesizeSpeechRequest with desired text and configuration,
// and receive stream of bytes in the requested format.
synthesizeOnline: {
// Submit a SynthesizeSpeechRequest with desired text and configuration,
// and receive stream of bytes in the requested format.
synthesizeOnline: {
path: '/nvidia.riva.tts.RivaSpeechSynthesis/SynthesizeOnline',
requestStream: false,
responseStream: true,
@@ -83,7 +83,7 @@ synthesizeOnline: {
responseDeserialize: deserialize_nvidia_riva_tts_SynthesizeSpeechResponse,
},
// Enables clients to request the configuration of the current Synthesize service, or a specific model within the service.
getRivaSynthesisConfig: {
getRivaSynthesisConfig: {
path: '/nvidia.riva.tts.RivaSpeechSynthesis/GetRivaSynthesisConfig',
requestStream: false,
responseStream: false,
+10 -2
View File
@@ -19,12 +19,20 @@ test('AWS - create and cache auth token', async(t) => {
return;
}
try {
let obj = await getAwsAuthToken(process.env.AWS_ACCESS_KEY_ID, process.env.AWS_SECRET_ACCESS_KEY, process.env.AWS_REGION);
let obj = await getAwsAuthToken({
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
//console.log({obj}, 'received auth token from AWS');
t.ok(obj.securityToken && !obj.servedFromCache, 'successfullY generated auth token from AWS');
await sleep(250);
obj = await getAwsAuthToken(process.env.AWS_ACCESS_KEY_ID, process.env.AWS_SECRET_ACCESS_KEY, process.env.AWS_REGION);
obj = await getAwsAuthToken({
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
region: process.env.AWS_REGION
});
//console.log({obj}, 'received auth token from AWS - second request');
t.ok(obj.securityToken && obj.servedFromCache, 'successfully received access token from cache');
+1 -1
View File
@@ -2,7 +2,7 @@ const test = require('tape').test ;
const exec = require('child_process').exec ;
test('starting docker network..', (t) => {
exec(`docker-compose -f ${__dirname}/docker-compose-testbed.yaml up -d`, (err, stdout, stderr) => {
exec(`docker compose -f ${__dirname}/docker-compose-testbed.yaml up -d`, (err, stdout, stderr) => {
setTimeout(() => {
t.end(err);
}, 2000);
+1 -1
View File
@@ -3,7 +3,7 @@ const exec = require('child_process').exec ;
test('stopping docker network..', (t) => {
t.timeoutAfter(10000);
exec(`docker-compose -f ${__dirname}/docker-compose-testbed.yaml down`, (err, stdout, stderr) => {
exec(`docker compose -f ${__dirname}/docker-compose-testbed.yaml down`, (err, stdout, stderr) => {
//console.log(`stderr: ${stderr}`);
process.exit(0);
});
-78
View File
@@ -1,78 +0,0 @@
const test = require('tape').test ;
const config = require('config');
const opts = config.get('redis');
const fs = require('fs');
const logger = require('pino')({level: 'error'});
process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
});
const stats = {
increment: () => {},
histogram: () => {}
};
test('IBM - create access key', async(t) => {
const fn = require('..');
const {client, getIbmAccessToken} = fn(opts, logger);
if (!process.env.IBM_API_KEY ) {
t.pass('skipping IBM test since no IBM api_key provided');
t.end();
client.quit();
return;
}
try {
let obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM');
t.ok(obj.access_token && !obj.servedFromCache, 'successfull received access token from IBM');
obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM - second request');
t.ok(obj.access_token && obj.servedFromCache, 'successfully received access token from cache');
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('IBM - retrieve tts voices test', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
t.pass('skipping IBM test since no IBM api_key and/or region provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'ibm',
credentials: {
tts_api_key: process.env.IBM_TTS_API_KEY,
tts_region: process.env.IBM_TTS_REGION
}
};
const obj = await getTtsVoices(opts);
const {voices} = obj.result;
//console.log(JSON.stringify(voices));
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from IBM`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
+1 -2
View File
@@ -2,6 +2,5 @@ require('./docker_start');
require('./synth');
require('./list-voices');
require('./aws');
require('./ibm');
require('./nuance');
require('./docker_stop');
-158
View File
@@ -12,164 +12,6 @@ const stats = {
histogram: () => {}
};
test('Verbio - get Access key and voices', async(t) => {
const fn = require('..');
const {client, getTtsVoices, getVerbioAccessToken} = fn(opts, logger);
if (!process.env.VERBIO_CLIENT_ID || !process.env.VERBIO_CLIENT_SECRET) {
t.pass('skipping Verbio test since no Verbio Keys provided');
t.end();
client.quit();
return;
}
try {
const credentials = {
client_id: process.env.VERBIO_CLIENT_ID,
client_secret: process.env.VERBIO_CLIENT_SECRET
};
let obj = await getVerbioAccessToken(credentials);
t.ok(obj.access_token , 'successfully received access token not from cache');
const voices = await getTtsVoices({vendor: 'verbio', credentials});
t.ok(voices && voices.length != 0, 'successfully received verbio voices');
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('IBM - create access key', async(t) => {
const fn = require('..');
const {client, getIbmAccessToken} = fn(opts, logger);
if (!process.env.IBM_API_KEY ) {
t.pass('skipping IBM test since no IBM api_key provided');
t.end();
client.quit();
return;
}
try {
let obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM');
t.ok(obj.access_token && !obj.servedFromCache, 'successfull received access token from IBM');
obj = await getIbmAccessToken(process.env.IBM_API_KEY);
//console.log({obj}, 'received access token from IBM - second request');
t.ok(obj.access_token && obj.servedFromCache, 'successfully received access token from cache');
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('IBM - retrieve tts voices test', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
t.pass('skipping IBM test since no IBM api_key and/or region provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'ibm',
credentials: {
tts_api_key: process.env.IBM_TTS_API_KEY,
tts_region: process.env.IBM_TTS_REGION
}
};
const obj = await getTtsVoices(opts);
const {voices} = obj.result;
//console.log(JSON.stringify(voices));
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from IBM`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance hosted tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET ) {
t.pass('skipping Nuance hosted test since no Nuance client_id and secret provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
client_id: process.env.NUANCE_CLIENT_ID,
secret: process.env.NUANCE_SECRET
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance on-prem tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_TTS_URI ) {
t.pass('skipping Nuance on-prem test since no Nuance uri provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
nuance_tts_uri: process.env.NUANCE_TTS_URI
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Google tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
-81
View File
@@ -1,81 +0,0 @@
const test = require('tape').test ;
const config = require('config');
const opts = config.get('redis');
const fs = require('fs');
const logger = require('pino')({level: 'error'});
process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
});
const stats = {
increment: () => {},
histogram: () => {}
};
test('Nuance hosted tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET ) {
t.pass('skipping Nuance hosted test since no Nuance client_id and secret provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
client_id: process.env.NUANCE_CLIENT_ID,
secret: process.env.NUANCE_SECRET
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance on-prem tests', async(t) => {
const fn = require('..');
const {client, getTtsVoices} = fn(opts, logger);
if (!process.env.NUANCE_TTS_URI ) {
t.pass('skipping Nuance on-prem test since no Nuance uri provided');
t.end();
client.quit();
return;
}
try {
const opts = {
vendor: 'nuance',
credentials: {
nuance_tts_uri: process.env.NUANCE_TTS_URI
}
};
let voices = await getTtsVoices(opts);
t.ok(voices.length > 0 && voices[0].language,
`GetVoices: successfully retrieved ${voices.length} voices from Nuance`);
await client.flushall();
t.end();
}
catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
+629 -187
View File
@@ -5,7 +5,7 @@ const fs = require('fs');
const {makeSynthKey} = require('../lib/utils');
const logger = require('pino')();
const bent = require('bent');
const getJSON = bent('json')
const getJSON = bent('json');
process.on('unhandledRejection', (reason, p) => {
console.log('Unhandled Rejection at: Promise', p, 'reason:', reason);
@@ -41,6 +41,7 @@ test('Google speech synth tests', async(t) => {
gender: 'FEMALE',
text: 'This is a test. This is only a test',
salt: 'foo.bar',
renderForCaching: true,
});
t.ok(!opts.servedFromCache, `successfully synthesized google audio to ${opts.filePath}`);
@@ -55,6 +56,7 @@ test('Google speech synth tests', async(t) => {
language: 'en-GB',
gender: 'FEMALE',
text: 'This is a test. This is only a test',
renderForCaching: true,
});
t.ok(opts.servedFromCache, `successfully retrieved cached google audio from ${opts.filePath}`);
@@ -78,6 +80,7 @@ test('Google speech synth tests', async(t) => {
language: 'en-GB',
gender: 'FEMALE',
text: 'This is a test. This is only a test',
renderForCaching: true,
});
t.ok(!opts.servedFromCache, `successfully synthesized google audio regardless of current cache to ${opts.filePath}`);
} catch (err) {
@@ -91,14 +94,17 @@ test('Google speech Custom voice synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GCP_CUSTOM_VOICE_FILE && !process.env.GCP_CUSTOM_VOICE_JSON_KEY || !process.env.GCP_CUSTOM_VOICE_MODEL) {
t.pass('skipping google speech synth tests since neither GCP_CUSTOM_VOICE_FILE nor GCP_CUSTOM_VOICE_JSON_KEY provided, GCP_CUSTOM_VOICE_MODEL is not provided');
if (!process.env.GCP_CUSTOM_VOICE_FILE &&
!process.env.GCP_CUSTOM_VOICE_JSON_KEY ||
!process.env.GCP_CUSTOM_VOICE_MODEL) {
t.pass(`skipping google speech synth tests since neither
GCP_CUSTOM_VOICE_FILE nor GCP_CUSTOM_VOICE_JSON_KEY provided, GCP_CUSTOM_VOICE_MODEL is not provided`);
return t.end();
}
try {
const str = process.env.GCP_CUSTOM_VOICE_JSON_KEY || fs.readFileSync(process.env.GCP_CUSTOM_VOICE_FILE);
const creds = JSON.parse(str);
let opts = await synthAudio(stats, {
const opts = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
@@ -109,9 +115,10 @@ test('Google speech Custom voice synth tests', async(t) => {
language: 'en-AU',
text: 'This is a test. This is only a test',
voice: {
reportedUsage:"REALTIME",
reportedUsage: 'REALTIME',
model: process.env.GCP_CUSTOM_VOICE_MODEL
}
},
renderForCaching: true,
});
t.ok(!opts.servedFromCache, `successfully synthesized google custom voice audio to ${opts.filePath}`);
} catch (err) {
@@ -121,6 +128,417 @@ test('Google speech Custom voice synth tests', async(t) => {
client.quit();
});
test('Google speech voice cloning synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GCP_CUSTOM_VOICE_FILE &&
!process.env.GCP_CUSTOM_VOICE_JSON_KEY ||
!process.env.GCP_VOICE_CLONING_FILE &&
!process.env.GCP_VOICE_CLONING_JSON_KEY) {
t.pass(`skipping google speech synth tests since neither
GCP_CUSTOM_VOICE_FILE nor GCP_CUSTOM_VOICE_JSON_KEY provided,
GCP_VOICE_CLONING_FILE nor GCP_VOICE_CLONING_JSON_KEY is not provided`);
return t.end();
}
try {
const googleKey = process.env.GCP_CUSTOM_VOICE_JSON_KEY ||
fs.readFileSync(process.env.GCP_CUSTOM_VOICE_FILE);
const voice_cloning_key = process.env.GCP_VOICE_CLONING_JSON_KEY ||
fs.readFileSync(process.env.GCP_VOICE_CLONING_FILE).toString();
const creds = JSON.parse(googleKey);
const opts = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
project_id: creds.project_id
},
},
language: 'en-US',
text: 'This is a test. This is only a test. This is a test. This is only a test. This is a test. This is only a test',
voice: {
voice_cloning_key
},
renderForCaching: true,
});
t.ok(!opts.servedFromCache, `successfully synthesized google voice cloning audio to ${opts.filePath}`);
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Google Gemini TTS synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GCP_FILE && !process.env.GCP_JSON_KEY) {
t.pass('skipping Google Gemini TTS synth tests since neither GCP_FILE nor GCP_JSON_KEY provided');
return t.end();
}
try {
const str = process.env.GCP_JSON_KEY || fs.readFileSync(process.env.GCP_FILE);
const creds = JSON.parse(str);
const geminiModel = process.env.GCP_GEMINI_TTS_MODEL || 'gemini-2.5-flash-tts';
// Test Gemini TTS with model and instructions (both required for Gemini)
let result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Kore',
model: geminiModel,
text: 'Hello, this is a test of Google Gemini text to speech.',
instructions: 'Speak clearly and naturally.',
renderForCaching: true,
});
t.ok(!result.servedFromCache, `successfully synthesized Google Gemini TTS audio to ${result.filePath}`);
t.ok(result.filePath.endsWith('.mp3'), 'Gemini TTS audio file has correct extension');
// Test Gemini TTS with different voice and instructions
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Charon',
model: geminiModel,
text: 'Welcome to our service. How can I help you today?',
instructions: 'Speak in a warm, friendly and professional tone.',
renderForCaching: true,
});
t.ok(!result.servedFromCache, `successfully synthesized Gemini TTS with instructions to ${result.filePath}`);
// Test cache retrieval
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Kore',
model: geminiModel,
text: 'Hello, this is a test of Google Gemini text to speech.',
instructions: 'Speak clearly and naturally.',
renderForCaching: true,
});
t.ok(result.servedFromCache, `successfully retrieved Gemini TTS audio from cache ${result.filePath}`);
// Test SSML stripping (Gemini doesn't support SSML)
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Leda',
model: geminiModel,
text: '<speak>This SSML should be stripped for Gemini TTS.</speak>',
instructions: 'Speak naturally.',
disableTtsCache: true,
renderForCaching: true,
});
t.ok(!result.servedFromCache, `successfully synthesized Gemini TTS with SSML stripped to ${result.filePath}`);
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Google TTS streaming tests (!JAMBONES_DISABLE_TTS_STREAMING)', async(t) => {
// Ensure streaming is enabled (default behavior)
delete process.env.JAMBONES_DISABLE_TTS_STREAMING;
// Clear require cache to reload config with new env var
delete require.cache[require.resolve('../lib/config')];
delete require.cache[require.resolve('../lib/synth-audio')];
delete require.cache[require.resolve('..')];
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GCP_FILE && !process.env.GCP_JSON_KEY) {
t.pass('skipping Google TTS streaming tests since neither GCP_FILE nor GCP_JSON_KEY provided');
return t.end();
}
try {
const str = process.env.GCP_JSON_KEY || fs.readFileSync(process.env.GCP_FILE);
const creds = JSON.parse(str);
const geminiModel = process.env.GCP_GEMINI_TTS_MODEL || 'gemini-2.5-flash-tts';
// Test 1: Standard voice streaming (use_live_api=0)
let result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'en-US-Wavenet-D',
gender: 'MALE',
text: 'This is a test of standard voice streaming.',
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'Standard voice returns streaming say: path');
t.ok(result.filePath.includes('vendor=google'), 'Standard voice streaming path contains vendor=google');
t.ok(result.filePath.includes('api_mode=tts'), 'Standard voice uses api_mode=tts');
t.ok(result.filePath.includes('voice=en-US-Wavenet-D'), 'Standard voice streaming path contains voice');
// Verify credentials are base64 encoded (no raw JSON braces that would break FreeSWitch parsing)
t.ok(result.filePath.includes('credentials='), 'Standard voice streaming path contains credentials');
t.ok(!result.filePath.includes('credentials={'), 'Credentials are not raw JSON (base64 encoded)');
// Test 2: HD voice streaming (api_mode=live)
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'en-US-Chirp3-HD-Charon',
text: 'This is a test of HD voice streaming.',
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'HD voice returns streaming say: path');
t.ok(result.filePath.includes('vendor=google'), 'HD voice streaming path contains vendor=google');
t.ok(result.filePath.includes('api_mode=live'), 'HD voice uses api_mode=live');
t.ok(result.filePath.includes('voice=en-US-Chirp3-HD-Charon'), 'HD voice streaming path contains voice');
// Test 3: Gemini TTS streaming (api_mode=gemini)
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Kore',
model: geminiModel,
text: 'This is a test of Gemini TTS streaming.',
instructions: 'Speak naturally.',
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'Gemini TTS returns streaming say: path');
t.ok(result.filePath.includes('vendor=google'), 'Gemini TTS streaming path contains vendor=google');
t.ok(result.filePath.includes('api_mode=gemini'), 'Gemini TTS uses api_mode=gemini');
t.ok(result.filePath.includes(`model_name=${geminiModel}`), 'Gemini TTS streaming path contains model_name');
t.ok(result.filePath.includes('prompt=Speak naturally.'), 'Gemini TTS streaming path contains prompt');
// Test 4: Gemini TTS with SSML stripping in streaming mode
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Leda',
model: geminiModel,
text: '<speak>This SSML should be stripped.</speak>',
instructions: 'Speak naturally.',
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'Gemini TTS with SSML returns streaming say: path');
t.ok(!result.filePath.includes('<speak>'), 'SSML tags are stripped from streaming path');
t.ok(result.filePath.includes('This SSML should be stripped.'), 'Text content is preserved after SSML stripping');
// Test 5: Gemini TTS with prompt containing special characters
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Kore',
model: geminiModel,
text: 'Testing special characters in prompt.',
options: { prompt: 'Speak in a warm, friendly tone' },
disableTtsCache: true
});
t.ok(result.filePath.startsWith('say:'), 'Gemini TTS with special chars returns streaming say: path');
// Commas in prompt should be replaced with semicolons
t.ok(result.filePath.includes('prompt=Speak in a warm; friendly tone'), 'Commas in prompt are escaped to semicolons');
// Test 6: options.apiMode override (force live on standard voice)
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'en-US-Wavenet-D',
text: 'Testing apiMode option override to live.',
options: { apiMode: 'live' },
disableTtsCache: true
});
t.ok(result.filePath.includes('api_mode=live'), 'options.apiMode=live overrides default for standard voice');
// Test 7: options.apiMode override (force gemini without model)
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Kore',
text: 'Testing apiMode option override to gemini.',
options: { apiMode: 'gemini' },
disableTtsCache: true
});
t.ok(result.filePath.includes('api_mode=gemini'), 'options.apiMode=gemini overrides default');
// Test 8: options.apiMode override (force tts on HD voice)
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'en-US-Chirp3-HD-Charon',
text: 'Testing apiMode option override to tts on HD voice.',
options: { apiMode: 'tts' },
disableTtsCache: true
});
t.ok(result.filePath.includes('api_mode=tts'), 'options.apiMode=tts overrides HD voice default');
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Google TTS non-streaming tests (JAMBONES_DISABLE_TTS_STREAMING=true)', async(t) => {
// Enable streaming disable flag
process.env.JAMBONES_DISABLE_TTS_STREAMING = 'true';
// Clear require cache to reload config with new env var
delete require.cache[require.resolve('../lib/config')];
delete require.cache[require.resolve('../lib/synth-audio')];
delete require.cache[require.resolve('..')];
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.GCP_FILE && !process.env.GCP_JSON_KEY) {
t.pass('skipping Google TTS non-streaming tests since neither GCP_FILE nor GCP_JSON_KEY provided');
delete process.env.JAMBONES_DISABLE_TTS_STREAMING;
return t.end();
}
try {
const str = process.env.GCP_JSON_KEY || fs.readFileSync(process.env.GCP_FILE);
const creds = JSON.parse(str);
const geminiModel = process.env.GCP_GEMINI_TTS_MODEL || 'gemini-2.5-flash-tts';
// Test 1: Standard voice falls back to non-streaming API
let result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'en-US-Wavenet-D',
gender: 'MALE',
text: 'This is a test with streaming disabled.',
disableTtsCache: true
});
t.ok(!result.filePath.startsWith('say:'), 'Standard voice does NOT return streaming say: path when disabled');
t.ok(result.filePath.endsWith('.mp3'), 'Standard voice returns mp3 file path');
// Test 2: HD voice falls back to non-streaming API
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'en-US-Chirp3-HD-Charon',
text: 'This is a test of HD voice with streaming disabled.',
disableTtsCache: true
});
t.ok(!result.filePath.startsWith('say:'), 'HD voice does NOT return streaming say: path when disabled');
t.ok(result.filePath.endsWith('.mp3'), 'HD voice returns mp3 file path');
// Test 3: Gemini TTS falls back to non-streaming API
result = await synthAudio(stats, {
vendor: 'google',
credentials: {
credentials: {
client_email: creds.client_email,
private_key: creds.private_key,
},
},
language: 'en-US',
voice: 'Kore',
model: geminiModel,
text: 'This is a test of Gemini TTS with streaming disabled.',
instructions: 'Speak naturally.',
disableTtsCache: true
});
t.ok(!result.filePath.startsWith('say:'), 'Gemini TTS does NOT return streaming say: path when disabled');
t.ok(result.filePath.endsWith('.mp3'), 'Gemini TTS returns mp3 file path');
} catch (err) {
console.error(err);
t.end(err);
} finally {
// Clean up: restore default behavior
delete process.env.JAMBONES_DISABLE_TTS_STREAMING;
delete require.cache[require.resolve('../lib/config')];
delete require.cache[require.resolve('../lib/synth-audio')];
delete require.cache[require.resolve('..')];
}
client.quit();
});
test('AWS speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -140,6 +558,7 @@ test('AWS speech synth tests', async(t) => {
language: 'en-US',
voice: 'Joey',
text: 'This is a test. This is only a test',
renderForCaching: true,
});
t.ok(!opts.servedFromCache, `successfully synthesized aws audio to ${opts.filePath}`);
@@ -153,6 +572,7 @@ test('AWS speech synth tests', async(t) => {
language: 'en-US',
voice: 'Joey',
text: 'This is a test. This is only a test',
renderForCaching: true,
});
t.ok(opts.servedFromCache, `successfully retrieved aws audio from cache ${opts.filePath}`);
} catch (err) {
@@ -339,82 +759,6 @@ test('Azure custom voice speech synth tests', async(t) => {
client.quit();
});
test('Nuance hosted speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.NUANCE_CLIENT_ID || !process.env.NUANCE_SECRET) {
t.pass('skipping Nuance speech synth tests since NUANCE_CLIENT_ID or NUANCE_SECRET not provided');
return t.end();
}
try {
let opts = await synthAudio(stats, {
vendor: 'nuance',
credentials: {
client_id: process.env.NUANCE_CLIENT_ID,
secret: process.env.NUANCE_SECRET,
},
language: 'en-US',
voice: 'Evan',
text: 'This is a test. This is only a test',
});
t.ok(!opts.servedFromCache, `successfully synthesized nuance audio to ${opts.filePath}`);
opts = await synthAudio(stats, {
vendor: 'nuance',
credentials: {
client_id: process.env.NUANCE_CLIENT_ID,
secret: process.env.NUANCE_SECRET,
},
language: 'en-US',
voice: 'Evan',
text: 'This is a test. This is only a test',
});
t.ok(opts.servedFromCache, `successfully retrieved nuance audio from cache ${opts.filePath}`);
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nuance on-prem speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.NUANCE_TTS_URI) {
t.pass('skipping Nuance on prem speech synth tests since NUANCE_TTS_URI not provided');
return t.end();
}
try {
let opts = await synthAudio(stats, {
vendor: 'nuance',
credentials: {
nuance_tts_uri: process.env.NUANCE_TTS_URI
},
language: 'en-US',
voice: 'Evan',
text: 'This is a test of on-prem. This is only a test',
});
t.ok(!opts.servedFromCache, `successfully synthesized nuance audio to ${opts.filePath}`);
opts = await synthAudio(stats, {
vendor: 'nuance',
credentials: {
nuance_tts_uri: process.env.NUANCE_TTS_URI
},
language: 'en-US',
voice: 'Evan',
text: 'This is a test of on-prem. This is only a test',
});
t.ok(opts.servedFromCache, `successfully retrieved nuance audio from cache ${opts.filePath}`);
} catch (err) {
console.error(err);
t.end(err);
}
client.quit();
});
test('Nvidia speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -452,46 +796,6 @@ test('Nvidia speech synth tests', async(t) => {
client.quit();
});
test('IBM watson speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.IBM_TTS_API_KEY || !process.env.IBM_TTS_REGION) {
t.pass('skipping IBM Watson speech synth tests since IBM_TTS_API_KEY or IBM_TTS_API_KEY not provided');
return t.end();
}
const text = `<speak> Hi there and welcome to jambones! jambones is the <sub alias="seapass">CPaaS</sub> designed with the needs of communication service providers in mind. This is an example of simple text-to-speech, but there is so much more you can do. Try us out!</speak>`;
try {
let opts = await synthAudio(stats, {
vendor: 'ibm',
credentials: {
tts_api_key: process.env.IBM_TTS_API_KEY,
tts_region: process.env.IBM_TTS_REGION,
},
language: 'en-US',
voice: 'en-US_AllisonV2Voice',
text,
});
t.ok(!opts.servedFromCache, `successfully synthesized ibm audio to ${opts.filePath}`);
opts = await synthAudio(stats, {
vendor: 'ibm',
credentials: {
tts_api_key: process.env.IBM_TTS_API_KEY,
tts_region: process.env.IBM_TTS_REGION,
},
language: 'en-US',
voice: 'en-US_AllisonV2Voice',
text,
});
t.ok(opts.servedFromCache, `successfully retrieved ibm audio from cache ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('Custom Vendor speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -507,8 +811,10 @@ test('Custom Vendor speech synth tests', async(t) => {
language: 'en-US',
voice: 'English-US.Female-1',
text: 'This is a test. This is only a test',
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized custom vendor audio to ${opts.filePath}`);
t.ok(opts.filePath.endsWith('wav'), 'audio is cached as wav file');
let obj = await getJSON(`http://127.0.0.1:3100/lastRequest/somethingnew`);
t.ok(obj.headers.Authorization == 'Bearer some_jwt_token', 'Custom Vendor Authentication Header is correct');
t.ok(obj.body.language == 'en-US', 'Custom Vendor Language is correct');
@@ -516,6 +822,22 @@ test('Custom Vendor speech synth tests', async(t) => {
t.ok(obj.body.type == 'text', 'Custom Vendor type is correct');
t.ok(obj.body.text == 'This is a test. This is only a test', 'Custom Vendor text is correct');
// Checking if cache is stored with wav format
opts = await synthAudio(stats, {
vendor: 'custom:somethingnew',
credentials: {
use_for_tts: 1,
custom_tts_url: "http://127.0.0.1:3100/somethingnew",
auth_token: 'some_jwt_token'
},
language: 'en-US',
voice: 'English-US.Female-1',
text: 'This is a test. This is only a test',
renderForCaching: true
});
t.ok(opts.servedFromCache, `successfully get custom vendor cached audio to ${opts.filePath}`);
t.ok(opts.filePath.endsWith('wav'), 'audio is cached as wav file');
opts = await synthAudio(stats, {
vendor: 'custom:somethingnew2',
credentials: {
@@ -526,6 +848,7 @@ test('Custom Vendor speech synth tests', async(t) => {
language: 'en-US',
voice: 'English-US.Female-1',
text: '<speak>This is a test. This is only a test</speak>',
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized Custom Vendor audio to ${opts.filePath}`);
obj = await getJSON(`http://127.0.0.1:3100/lastRequest/somethingnew2`);
@@ -574,41 +897,34 @@ test('Elevenlabs speech synth tests', async(t) => {
t.end(err);
}
client.quit();
})
});
test('PlayHT speech synth tests', async(t) => {
test('Cartesia speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.PLAYHT_API_KEY || !process.env.PLAYHT_USER_ID) {
t.pass('skipping PlayHT speech synth tests since PLAYHT_API_KEY or PLAYHT_USER_ID is/are not provided');
if (!process.env.CARTESIA_API_KEY) {
t.pass('skipping Cartesia speech synth tests since CARTESIA_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones!';
const text = 'Hi there and welcome to jambones! ' + Date.now();
try {
let opts = await synthAudio(stats, {
vendor: 'playht',
const opts = await synthAudio(stats, {
vendor: 'cartesia',
credentials: {
api_key: process.env.PLAYHT_API_KEY,
user_id: process.env.PLAYHT_USER_ID,
voice_engine: 'PlayHT2.0-turbo',
api_key: process.env.CARTESIA_API_KEY,
model_id: 'sonic-english',
options: JSON.stringify({
quality: "medium",
speed: 1,
seed: 1,
temperature: 1,
emotion: "female_happy",
voice_guidance: 3,
style_guidance: 20,
text_guidance: 1,
emotion: 'female_happy',
})
},
language: 'en-US',
voice: 's3://voice-cloning-zero-shot/d9ff78ba-d016-47f6-b0ef-dd630f59414e/female-cs/manifest.json',
language: 'en',
voice: '694f9389-aac1-45b6-b726-9d9369183238',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully playht eleven audio to ${opts.filePath}`);
t.ok(!opts.servedFromCache, `successfully cartesia eleven audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
@@ -617,7 +933,66 @@ test('PlayHT speech synth tests', async(t) => {
client.quit();
});
test('rimelabs speech synth tests', async(t) => {
test('inworld speech synth', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.INWORLD_API_KEY) {
t.pass('skipping inworld speech synth tests since INWORLD_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones!';
try {
const opts = await synthAudio(stats, {
vendor: 'inworld',
credentials: {
api_key: process.env.INWORLD_API_KEY,
model_id: 'inworld-tts-1'
},
language: 'en',
voice: 'Ashley',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized inworld audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('resemble speech synth', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.RESEMBLE_API_KEY) {
t.pass('skipping resemble speech synth tests since RESEMBLE_API_KEY is not provided');
return t.end();
}
const text = '<speak prompt="Speak in an excited, upbeat tone">Hello from Resemble!</speak>';
try {
const opts = await synthAudio(stats, {
vendor: 'resemble',
credentials: {
api_key: process.env.RESEMBLE_API_KEY,
},
language: 'en',
voice: '3f5fb9f1',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized resemble audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('rimelabs speech synth tests mist', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -627,7 +1002,7 @@ test('rimelabs speech synth tests', async(t) => {
}
const text = 'Hi there and welcome to jambones!';
try {
let opts = await synthAudio(stats, {
const opts = await synthAudio(stats, {
vendor: 'rimelabs',
credentials: {
api_key: process.env.RIMELABS_API_KEY,
@@ -637,7 +1012,7 @@ test('rimelabs speech synth tests', async(t) => {
reduceLatency: false
})
},
language: 'en-US',
language: 'eng',
voice: 'amber',
text,
renderForCaching: true
@@ -651,6 +1026,40 @@ test('rimelabs speech synth tests', async(t) => {
client.quit();
});
test('rimelabs speech synth tests mistv2', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.RIMELABS_API_KEY) {
t.pass('skipping rimelabs speech synth tests since RIMELABS_API_KEY is not provided');
return t.end();
}
const text = 'Hi there and welcome to jambones!';
try {
const opts = await synthAudio(stats, {
vendor: 'rimelabs',
credentials: {
api_key: process.env.RIMELABS_API_KEY,
model_id: 'mistv2',
options: JSON.stringify({
speedAlpha: 1.0,
reduceLatency: false
})
},
language: 'spa',
voice: 'pablo',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized rimelabs mistv2 audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('whisper speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -681,39 +1090,6 @@ test('whisper speech synth tests', async(t) => {
client.quit();
});
test('Verbio speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.VERBIO_CLIENT_ID || !process.env.VERBIO_CLIENT_SECRET) {
t.pass('skipping Verbio Synthesize test since no Verbio Keys provided');
t.end();
client.quit();
return;
}
const text = 'Hi there and welcome to jambones!';
try {
let opts = await synthAudio(stats, {
vendor: 'verbio',
credentials: {
client_id: process.env.VERBIO_CLIENT_ID,
client_secret: process.env.VERBIO_CLIENT_SECRET
},
language: 'en-US',
voice: 'tommy_en-us',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized whisper audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
})
test('Deepgram speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
@@ -742,6 +1118,69 @@ test('Deepgram speech synth tests', async(t) => {
client.quit();
})
test('Deepgram Flux speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.DEEPGRAM_API_KEY) {
t.pass('skipping Deepgram Flux speech synth tests since DEEPGRAM_API_KEY');
return t.end();
}
const text = 'Hi there and welcome to jambones!';
try {
const opts = await synthAudio(stats, {
vendor: 'deepgramflux',
credentials: {
api_key: process.env.DEEPGRAM_API_KEY
},
model: process.env.DEEPGRAM_FLUX_MODEL || 'flux-alexis-en',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache, `successfully synthesized deepgramflux audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('xai speech synth tests', async(t) => {
const fn = require('..');
const {synthAudio, client} = fn(opts, logger);
if (!process.env.XAI_API_KEY) {
t.pass('skipping xai speech synth tests - no XAI_API_KEY');
return t.end();
}
const text = 'Hi there and welcome to jambones!';
try {
const opts = await synthAudio(stats, {
vendor: 'xai',
credentials: {
api_key: process.env.XAI_API_KEY,
options: JSON.stringify({
voice: process.env.XAI_VOICE || 'eve',
speed: 1.0,
optimize_streaming_latency: 1,
text_normalization: true
})
},
language: 'en',
voice: process.env.XAI_VOICE || 'eve',
text,
renderForCaching: true
});
t.ok(!opts.servedFromCache && opts.filePath, `successfully synthesized xai audio to ${opts.filePath}`);
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
}
client.quit();
});
test('TTS Cache tests', async(t) => {
const fn = require('..');
const {purgeTtsCache, getTtsSize, client} = fn(opts, logger);
@@ -750,11 +1189,12 @@ test('TTS Cache tests', async(t) => {
// save some random tts keys to cache
const minRecords = 8;
for (const i in Array(minRecords).fill(0)) {
await client.set(makeSynthKey({vendor: i, language: i, voice: i, engine: i, text: i}), i);
await client.set(makeSynthKey({vendor: i, language: i, voice: i, engine: i, model: i, text: i,
instructions: i}), i);
}
const count = await getTtsSize();
t.ok(count >= minRecords, 'getTtsSize worked.');
const {purgedCount} = await purgeTtsCache();
t.ok(purgedCount >= minRecords, `successfully purged at least ${minRecords} tts records from cache`);
@@ -769,7 +1209,7 @@ test('TTS Cache tests', async(t) => {
try {
// save some random tts keys to cache
for (const i in Array(10).fill(0)) {
await client.set(makeSynthKey({vendor: i, language: i, voice: i, engine: i, text: i}), i);
await client.set(makeSynthKey({vendor: i, language: i, voice: i, engine: i, text: i, instructions: i}), i);
}
// save a specific key to tts cache
const opts = {vendor: 'aws', language: 'en-US', voice: 'MALE', engine: 'Engine', text: 'Hello World!'};
@@ -785,13 +1225,15 @@ test('TTS Cache tests', async(t) => {
language: 'non-existing',
voice: 'non-existing',
});
t.ok(purgedCountWhenErrored === 0, `purged no records when specified key was not found`);
t.ok(error, `error returned when specified key was not found`);
t.ok(purgedCountWhenErrored === 0, 'purged no records when specified key was not found');
t.ok(error, 'error returned when specified key was not found');
// make sure other tts keys are still there
const cached = (await client.keys('tts:*')).length;
t.ok(cached >= 1, `successfully kept all non-specified tts records in cache`);
const cached = await client.keys('tts:*');
t.ok(cached.length >= 1, 'successfully kept all non-specified tts records in cache');
process.env.VG_TRIM_TTS_SILENCE = 'true';
await client.set(makeSynthKey({ vendor: 'azure' }), 'value');
} catch (err) {
console.error(JSON.stringify(err));
t.end(err);
@@ -805,10 +1247,10 @@ test('TTS Cache tests', async(t) => {
const account_sid = "12412512_cabc_5aff"
const account_sid2 = "22412512_cabc_5aff"
for (const i in Array(minRecords).fill(0)) {
await client.set(makeSynthKey({account_sid, vendor: i, language: i, voice: i, engine: i, text: i}), i);
await client.set(makeSynthKey({account_sid, vendor: i, language: i, voice: i, engine: i, text: i, instructions: i}), i);
}
for (const i in Array(minRecords).fill(0)) {
await client.set(makeSynthKey({account_sid: account_sid2, vendor: i, language: i, voice: i, engine: i, text: i}), i);
await client.set(makeSynthKey({account_sid: account_sid2, vendor: i, language: i, voice: i, engine: i, text: i, instructions: i}), i);
}
const {purgedCount} = await purgeTtsCache({account_sid});
t.equal(purgedCount, minRecords, `successfully purged at least ${minRecords} tts records from cache for account_sid:${account_sid}`);