* fix(gather): clear a finalized deepgram interim word so UtteranceEnd is not deferred forever
The #1088 check defers UtteranceEnd while an interim word newer than
last_word_end is pending. The marker was only cleared by a non-empty
is_final, so an interim word that Deepgram later dropped (e.g. line
noise, followed by an empty final) left it set: the gather waited for
the caller to speak again and merged both utterances.
Clear the marker on any Deepgram is_final result whose window
(start + duration) covers it. A word that Deepgram carries into the
next segment still defers UtteranceEnd.
Port of jambonz/feature-server#223.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(gather): return the buffer when a deferred deepgram UtteranceEnd is satisfied
Clearing the unprocessed-word marker only helps if the covering final
arrives before UtteranceEnd. If it arrives after, UtteranceEnd has
already been deferred and Deepgram sends no further UtteranceEnd until
new speech, so the next utterance is still merged in.
Remember a deferred UtteranceEnd, and when a later final clears the
marker, return the buffered transcript once that final has been
handled (setImmediate, since _onTranscription has early returns).
Port of jambonz/feature-server#225.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(gather): only an empty final ends a gather early after a deferred UtteranceEnd
A final with words means the caller resumed talking after the pause;
Deepgram will send a new UtteranceEnd after those words, so keep the
existing flow instead of cutting the caller off mid-sentence (or
skipping the asrTimeout wait). Only a final that drops the pending word
(empty) leaves nothing else to end the gather.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This repo is public and held a long-lived AWS access key as repository secrets
(set 2023-11-22). It is replaced with short-lived credentials from the GitHub OIDC
provider; no AWS key and no account id remain in the repo.
create-test-db.js wrote {access_key_id, secret_access_key, aws_region} into the test
database as the aws speech credential, which sends speech-utils' getAwsAuthToken down
its access-key branch and calls GetSessionToken -- rejected by AWS for session
credentials. The role_arn branch calls AssumeRole instead, which accepts them, and is
already plumbed through db-utils.js, call-session.js and stt-task.js. The pinned
speech-utils 0.2.30 already supports it, so no dependency change is needed.
Fork pull requests receive neither secrets nor an OIDC token, so the credentials step
is guarded by a condition; the AWS tests then skip for forks exactly as they do today.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The optionalDependencies pinned utf-8-validate ^6.0.3, but ws@8 only
accepts ^5.0.2 as its optional peer and ibm-watson's websocket dep
hard-requires ^5.0.2. That 5.x-vs-6.x split forced a dual-version tree
that npm 10 and npm 11 lay out differently, so a lockfile generated with
npm 11 (local) failed 'npm ci' under npm 10 (CI, node 20) with
'Missing: utf-8-validate@5.0.10 from lock file'. The 6.x pin never
accelerated ws anyway. Removing it collapses to a single utf-8-validate
5.0.10 that both npm majors resolve identically.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix race condition where gather resolves with speech transcript but timeout timer gets set after the resolve and is left running after gather completes
* remove unneeded line of code
* for deepgram flux include the turn taking events in the transcription payload
* for deepgram flux, including turn_taking_event in the speech payload
* fix prev commit which used wrong field
* update to speech-utils that generates playback id
* modify tts and say task to track current playback id and match against start and stop events
* bump speech utils
* wip
* wip
* fix race condition where say with playbackId gets stop event from previous play from cache file
* logging
* wip
* fix comparison when playing cached files
* logging
* fix for race condition when play-stop event from earlier command received
* wip
* say verb should not cache if disableTtsCache = true
* logging
* modify race condition logic to validate playback id in playback-stopped matches that from playback-start
* logging
---------
Co-authored-by: Quan HL <quan.luuhoang8@gmail.com>
* Revert "Update dial.js (#1243)"
This reverts commit 259dedcded.
* add to .gitignore
* when we receive a REFER on the parent leg, after adulting the child the dial task in the parent session should end
* better handling of flush commands
* rework buffering of tokens
* gather: when returning low confidence also provide the transcript
* better error handling in tts:tokens
* special handling of asr timeout for speechmatics
* remove some logs that were excessively wordy
* add speechmatics options
* wip
* speechmatics does not do endpointing for us so we need to flip on continuousAsr
* speechmatics: continousAsr should be at least equal to max_delay, if set
* wip
* add TtsStreamingBuffer class to abstract handling of streaming tokens
* wip
* add throttling support
* support background ttsStream (#995)
* wip
* add TtsStreamingBuffer class to abstract handling of streaming tokens
* wip
* support background ttsStream
* wip
---------
Co-authored-by: Dave Horton <daveh@beachdognet.com>
* wip
* dont send if we have nothing to send
* initial testing with cartesia
* wip
---------
Co-authored-by: Hoan Luu Huu <110280845+xquanluu@users.noreply.github.com>
* fix transcribe fixes for speechmatics
* update to verb-specs with fixes for speechmatics
* add support for speechmatics translation
* add handlers for receiving translations
* call translation hookd
* gather: no need to restart speechmatics after a final transcript during continuous asr
* graceful shutdown
* wip
* wip
* wip
* wip
* wip
* add support for aws language model name when transcribing
* wip - from prev branch
* wip
* support both aws grpc and ws api - detect based on transcription payload
* update to drachtio-fsmrf@4.0.0
* fix for grpc compatibility, requires JAMBONES_AWS_TRANSCRIBE_USE_GRPC env
* back out major update to drachtio-srf and fsmrf; that should come in a separate PR
* update drachtio-srf