v0.7.3 release notes (#29)

This commit is contained in:
Dave Horton
2022-02-11 10:49:47 -05:00
committed by GitHub
parent c449927c96
commit fd945d5f04
4 changed files with 74 additions and 19 deletions
+3
View File
@@ -109,6 +109,9 @@ navi:
path: release-notes
title: Release Notes
pages:
-
path: v0.7.3
title: v0.7.3
-
path: v0.7.2
title: v0.7.2
+22
View File
@@ -0,0 +1,22 @@
# Release v0.7.3
> Release Date: Feb 11, 2022
#### New Features
- Adds support for redis authentication (via JAMBONES_REDIS_USERNAME / JAMBONES_REDIS_PASSWORD env vars).
- Adds support for voice activity detection to `gather` and `transcribe` verbs: if activated, the connection to the speech recognizer will be delayed until speech has been detected. (In some deployments this can save on costs).
- When using the `dial` verb with multiple targets, you can specify different custom headers to be used on a per-target basis.
- Adds `overrideTo` property to `dial` verb to allow an outbound call to be placed to a registered user, but provide a different user or extension on the outbound invite. This can be useful for PBXs that register with jambonz under a certain user but can accept calls to a number of different phone numbers or extensions.
- Adds fs_sip_address and api_base_url in webhook payloads; this aids in troubleshooting and can simplfy the coding of webhook apps that also want to use the REST api.
- Adds fs_public_ip to webhook payload when running in ec2 autoscale group.
#### Bug fixes
- Fixes memory leaks resulting from not closing speech synthesizer connections properly.
- Fixes race condition on hangup that caused duplicate call-status webhooks.
- Fixes race condition on hangup that sometimes resulted in outbound call attempt even though caller had hung up.
- Adds missing Azure regions to jambonz portal for speech credentials.
#### Availability
- Available shortly on <a href="https://aws.amazon.com/marketplace/pp/prodview-55wp45fowbovo" target="_blank" >AWS Marketplace</a>
- Deploy to Kubernetes using [this Helm chart](https://github.com/jambonz/helm-charts)
**Questions?** Contact us at <a href="mailto:support@jambonz.org">support@jambonz.org</a>
+22 -11
View File
@@ -28,17 +28,28 @@ You can use the following options in the `gather` command:
| option | description | required |
| ------------- |-------------| -----|
| actionHook | webhook POST to invoke with the collected digits or speech. The payload will include a 'speech' or 'dtmf' property along with the standard attributes. See below for more detail.| yes |
| finishOnKey | dmtf key that signals the end of input | no |
| input | array, specifying allowed types of input: ['digits'], ['speech'], or ['digits', 'speech']. Default: ['digits'] | no |
| numDigits | number of dtmf digits expected to gather | no |
| partialResultHook | webhook to send interim transcription results to. Partial transcriptions are only generated if this property is set. | no |
| play | nested [play](#play) command that can be used to prompt the user | no |
| recognizer.hints | array of words or phrases to assist speech detection | no |
| recognizer.language | language code to use for speech detection. Defaults to the application level setting, or 'en-US' if not set | no |
| recognizer.profanityFilter | if true, filter profanity from speech transcription. Default: no| no |
| recognizer.vendor | speech vendor to use (currently only Google supported) | no |
| say | nested [say](#say) command that can be used to prompt the user | no |
| actionHook | Webhook POST to invoke with the collected digits or speech. The payload will include a 'speech' or 'dtmf' property along with the standard attributes. See below for more detail.| yes |
| finishOnKey | Dmtf key that signals the end of input | no |
| input |Array, specifying allowed types of input: ['digits'], ['speech'], or ['digits', 'speech']. Default: ['digits'] | no |
| numDigits | Number of dtmf digits expected to gather | no |
| partialResultHook | Webhook to send interim transcription results to. Partial transcriptions are only generated if this property is set. | no |
| play | nested [play](#play) Command that can be used to prompt the user | no |
| recognizer.vendor | Speech vendor to use (google, aws, or microsoft) | no |
| recognizer.language | Language code to use for speech detection. Defaults to the application level setting| no |
| recognizer.vad.enable|If true, delay connecting to cloud recognizer until speech is detected|no|
| recognizer.vad.voiceMs|If vad is enabled, the number of milliseconds of speech required before connecting to cloud recognizer|no|
| recognizer.vad.mode|If vad is enabled, this setting governs the sensitivity of the voice activity detector; value must be between 0 to 3 inclusive (lower numbers mean more sensitivity, i.e. more likely to return a false positive). Default: 2|no|
| recognizer.hints | (google and microsoft only) Array of words or phrases to assist speech detection | no |
| recognizer.altLanguages |(google only) An array of alternative languages that the speaker may be using | no |
| recognizer.profanityFilter | (google only) If true, filter profanity from speech transcription . Default: no| no |
| recognizer.vocabularyName | (aws only) The name of a vocabulary to use when processing the speech.| no |
| recognizer.vocabularyFilterName | (aws only) The name of a vocabulary filter to use when processing the speech.| no |
| recognizer.filterMethod | (aws only) The method to use when filtering the speech: remove, mask, or tag.| no |
| recognizer.profanityOption | (microsoft only) masked, removed, or raw. Default: raw| no |
| recognizer.outputFormat | (microsoft only) simple or detailed. Default: simple| no |
| recognizer.requestSnr | (microsoft only) Request signal to noise information| no |
| recognizer.initialSpeechTimeoutMs | (microsoft only) Initial speech timeout in milliseconds| no |
| say | nested [say](#say) Command that can be used to prompt the user | no |
| timeout | The number of seconds of silence or inaction that denote the end of caller input. The timeout timer will begin after any nested play or say command completes. Defaults to 5 | no |
In the case of speech input, the actionHook payload will include a `speech` object with the response from Google speech:
+27 -8
View File
@@ -20,14 +20,33 @@ You can use the following options in the `transcribe` command:
| option | description | required |
| ------------- |-------------| -----|
| recognizer.dualChannel | if true, transcribe the parent call as well as the child call | no |
| recognizer.interim | if true interim transcriptions are sent | no (default: false) |
| recognizer.language | language to use for speech transcription | yes |
| recognizer.profanityFilter | if true, filter profanity from speech transcription. Default: no| no |
| recognizer.vendor | speech vendor to use (currently only Google supported) | no |
| transcriptionHook | webhook to call when a transcription is received. Due to the richness of information in the transcription an HTTP POST will always be sent. | yes |
> **Note**: the `dualChannel` property is not currently implemented.
| recognizer.vendor | Speech vendor to use (google, aws, or microsoft) | no |
| recognizer.language | Language code to use for speech detection. Defaults to the application level setting, or 'en-US' if not set | no |
| recognizer.interim | If true interim transcriptions are sent | no (default: false) |
| recognizer.vad.enable|If true, delay connecting to cloud recognizer until speech is detected|no|
| recognizer.vad.voiceMs|If vad is enabled, the number of milliseconds of speech required before connecting to cloud recognizer|no|
| recognizer.vad.mode|If vad is enabled, this setting governs the sensitivity of the voice activity detector; value must be between 0 to 3 inclusive, lower numbers mean more sensitive|no|
| recognizer.separateRecognitionPerChannel | If true, recognize both caller and called party speech | no |
| recognizer.altLanguages |(google only) An array of alternative languages that the speaker may be using | no |
| recognizer.punctuation |(google only) Enable automatic punctuation | no |
| recognizer.enhancedModel |(google only) Use enhanced model | no |
| recognizer.words |(google only) Enable word offsets | no |
| recognizer.diarization |(google only) Enable speaker diarization | no |
| recognizer.diarizationMinSpeakers |(google only) Set the minimum speaker count | no |
| recognizer.diarizationMaxSpeakers |(google only) Set the maximum speaker count | no |
| recognizer.interactionType |(google only) Set the interaction type: discussion, presentation, phone_call, voicemail, professionally_produced, voice_search, voice_command, dictation | no |
| recognizer.naicsCode |(google only) set an industry [NAICS](https://www.census.gov/naics/?58967?yearbck=2022) code that is relevant to the speech | no |
| recognizer.hints | (google and microsoft only) Array of words or phrases to assist speech detection | no |
| recognizer.profanityFilter | (google only) If true, filter profanity from speech transcription . Default: no| no |
| recognizer.vocabularyName | (aws only) The name of a vocabulary to use when processing the speech.| no |
| recognizer.vocabularyFilterName | (aws only) The name of a vocabulary filter to use when processing the speech.| no |
| recognizer.filterMethod | (aws only) The method to use when filtering the speech: remove, mask, or tag.| no |
| recognizer.identifyChannels | (aws only) Enable channel identification. | no |
| recognizer.profanityOption | (microsoft only) masked, removed, or raw. Default: raw| no |
| recognizer.outputFormat | (microsoft only) simple or detailed. Default: simple| no |
| recognizer.requestSnr | (microsoft only) Request signal to noise information| no |
| recognizer.initialSpeechTimeoutMs | (microsoft only) Initial speech timeout in milliseconds| no |
| transcriptionHook | Webhook to receive an HTPP POST when an interim or final transcription is received. | yes |
<p class="flex">
<a href="/docs/webhooks/tag">Prev: tag</a>