* feat: surface the SIP Reason header on call status events
Carriers fronting ISDN/E1 PRI trunks put the authoritative disconnect cause in
an RFC 3326 Reason header rather than in the SIP status line, e.g.
SIP/2.0 408 Request Timeout
Reason: Q.850 ;cause=18
Q.850 cause 18 is "no user responding" - nobody answered, not a platform fault.
Different causes also arrive under the same SIP status (503 with cause=38
network out of order, or cause=41 temporary failure), so the status code alone
cannot classify the outcome of the call.
drachtio relays the header intact and it is present on the response object, but
the feature server only read the status line from it, so the cause was lost at
the application boundary and never reached the call status webhook.
Rather than extract the header at each emit site, carry the SIP message that
caused the status change on the callStatusChange event and derive from it in one
place, so provisional responses, the 200, final failures, BYE and CANCEL are all
covered by the same code and future headers cost one line.
Note this partly overlaps the existing _extractCustomHeaders/sip_headers
passthrough: that already exposes a Reason header arriving on a BYE, but it only
runs on the hangup path, so nothing covered the outbound INVITE failure
responses where a Q.850 cause matters most. sip_reason_header is a dedicated,
documented field that behaves the same on every status event.
sip_reason keeps meaning the status line phrase, and no key is added when there
is no Reason header, so existing consumers see an unchanged payload.
Adds a test:unit script so the smoke test runs without the docker testbed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: note that the Reason header arrives re-serialized, not verbatim
Verified live end-to-end against a cluster, capturing both the external leg and
the leg into the feature server:
14.226.234.142 -> 10.0.197.31:5060 Reason: Q.850 ;cause=31 (as sent)
10.0.197.31:5060 -> :5070 Reason: Q.850;cause=31 (to fs)
Proxying re-serializes the header, normalizing the optional whitespace RFC 3326
permits around ';'. The same normalization appears in customer captures from an
unrelated deployment, so this is drachtio behaviour, not cluster-specific.
sip_reason_header is therefore the header as this process received it, not the
carrier's exact bytes - worth stating outright, since the spacing inconsistency
is exactly what consumers ask about, and a consumer who string-matches on the
spaced form would silently never match.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: clear a stale Reason header in redis, pass it on the alloc path, run the tests in CI
Three problems found reviewing the earlier commits.
1. The redis call record kept a stale header forever. updateCallStatus assigns
unconditionally, but toJSON only emits truthy values, and the same object is
written to the redis hash with hmset - a MERGE. An absent key therefore left
the previous status change's header in place, readable via GET /Calls/:sid:
a leg that got 183 + "Reason: Q.850;cause=31", then answered on a clean 200
and completed on a plain BYE, reported a Q.850 temporary-failure cause for a
call that ended normally. The webhooks were right; only the call record was
wrong. Storing absence as '' overwrites it. The existing "does not linger"
test passed throughout because it only exercised toJSON in memory, so this
adds one that pins what survives the redis filter (verified: it fails
without the fix).
2. The endpoint-allocation failure path already extracted the Reason header and
relayed it to the SBC, but called _notifyCallStatusChange without it - so a
FreeSWITCH 488 with "Reason: Q.850;cause=88 INCOMPATIBLE_DESTINATION" told
the SBC the cause and the application nothing, which is exactly the case
this feature exists to expose. That error is an fsmrf object rather than a
SipMessage, so it cannot go through msg (the guard in
reasonHeaderFromSipMessage would silently return undefined); the event now
takes an explicit sipReasonHeader for callers holding the value already.
3. The test:unit script added with these tests was never wired into CI - the
workflow runs jslint and npm test, and npm test enumerates its files
explicitly and skips test/unit entirely, so the suite would have rotted
unnoticed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: clear the stale header on the CallSession path too, and keep the CANCEL
Both from PR review.
1. The previous fix only worked for SingleDialer. Storing '' on the instance is
enough for callers that write the instance itself, but CallSession writes
toJSON(), and toJSON() drops falsy values on purpose so the key stays out of
the webhook payload - so '' never reached hmset and the stale header
survived. Reproduced:
after 183, toJSON has: Q.850;cause=31
raw instance value: ""
instance -> redis has key: true (SingleDialer, clears)
toJSON() -> redis has key: false (CallSession, stale persists)
The split is by session class, not call direction: place-outdial is required
only by dial.js, so SingleDialer covers dial-verb child legs while inbound,
REST-created and adulting all inherit CallSession._notifyCallStatusChange -
including the REST outdial path this feature was written for.
Adds CallInfo.toRedisJSON(), a named projection for the merge-semantics
store, so the divergence from toJSON() lives in one documented place rather
than being rediscovered at each call site. The unit test asserted against the
raw instance, which is why it passed while the real path was broken; it now
goes through both writers and fails if either regresses.
2. The caller-abandoned race dropped the CANCEL. middleware.js had it in hand
and discarded it, so the constructor's _onCancel() passed nothing whenever
the CANCEL beat the application fetch - the header landed or not depending on
timing. Worth closing because the 487 and its 'Request Terminated' phrase are
both ours, so every abandoned inbound call looks identical: a
Reason: SIP;cause=200;text="Call completed elsewhere" is what separates a
forked branch losing the race from a caller who gave up, and today both are
just no-answer.
Confirmed on a deployed srf (5.0.27) that this is not a no-op:
copyUASHeaderToUACForOnlyCancel forwards a hardcoded ['Reason', 'X-Reason']
when proxying a CANCEL, so the header does reach us.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: keep the Reason header out of the redis call record
Reversing the earlier approach after a closer look at who reads what.
The header was reaching the redis call record simply because CallInfo feeds two
sinks with opposite semantics: the status webhook is an EVENT (each POST is
independent, an absent key means "not in this event") while the redis record is
STATE written with hmset, a MERGE (an absent key means "keep what was there").
sipReasonHeader is the first field here that can legitimately go from set back
to unset - sipStatus and sipReason are only ever overwritten - so it was the
first to expose the mismatch, and two rounds of fixes had to chase it because
the two writers project from different bases.
Rather than keep managing that, exclude it: redis holds calls that are still
live, where there is usually no interesting cause yet, and by the time there is
one the call is over. Call history is served from RecentCalls (the CDR), which
is what the webapp reads - not GET /Calls. So the field earned very little there
while costing an invariant that has to be remembered forever.
Worth being explicit that this is NOT simply a revert: dropping the '' would
only have stopped the empty value being written. When the header is PRESENT it
still reached redis through both writers - toJSON includes it, and SingleDialer
writes the instance's own properties - and was then never cleared. Keeping it out
takes the same machinery as clearing it, just inverted, so this is a choice about
where the field belongs rather than a saving.
WEBHOOK_ONLY_FIELDS names that intent in one place, with the reasoning, so the
next field with the same property has somewhere obvious to go. The test asserts
both writers, since they project from different bases and checking one would
pass while the other still wrote the field (verified: it fails if the exclusion
is removed).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: apply the redis projection at the boundary, and drop dead adulting plumbing
From PR review.
1. There was a THIRD call-record writer. Asking each call site to remember the
projection was the wrong shape, and review found the proof: this class writes
the record from the status change, the recording flag AND (private only)
_persistConferenceState, and the last one still passed toJSON() straight
through. A conferenced caller whose BYE carried Reason: Q.850;cause=16 would
have that written into the call hash on the next conference-state persist,
where hmset merges and nothing can ever clear it.
Fixed by wrapping updateCallStatus once where it is bound, so every write is
projected and a newly added writer cannot bypass it by forgetting to ask. The
call sites go back to passing plain shapes. This is a class of miss that a
unit test cannot catch - the contract test passed the whole time - so the fix
is structural rather than another assertion.
2. The msg/byeReq parameters added to AdultingCallSession were dead code. The
only path that reacts to the far-end BYE there is the inline
sd.dlg.on('destroy') handler, which calls _callReleased() and discards the
request; _hangup is reachable only from CallSession.hangup() (LCC), which
passes nothing. The Reason header on that leg does reach the webhook, via
SingleDialer's own destroy handler - so the plumbing was not just unused but
misleading about which path carries it. Removed, with a comment at the handler
recording where the event actually comes from so it does not get re-added.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: correct the projection comment in SingleDialer
Copied verbatim from CallSession, where the list of writers (status change,
recording flag, conference state) is accurate. SingleDialer has one writer, so
the comment described code that is not there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add support for sending 'amd' property in createCall REST API and also added support for using any of the speech vendors for STT
---------
Co-authored-by: Dave Horton <daveh@beachdognet.com>
* initial adds for otel tracing
* initial basic testing
* basic tracing for incoming calls
* linting
* add traceId to the webhook params
* trace webhook calls
* tracing: add new commands as tags when receiving async commands over websocket
* tracing new commands
* add summary for config verb
* trace async commands
* bugfix: undefined ref
* tracing: give time for final webhooks before closing root span
* tracing bugfix: span for background gather was not ended
* tracing - minor tag changes
* tracing - add span atttribute for reason call ended
* trace call status webhooks, add app version to trace output
* config: add support for automatically re-enabling
* env var to customize service name in tracing UI
* config: change to use 'sticky' attribute to re-enable bargein automatically
* fix warnings
* when adulting create a new root span
* when background gather triggers bargein via vad clear queue of tasks
* additional trace attributes for dial and refer
* fix dial tracing
* add better summary for dial
* fix prev commit
* add exponential backoff to WsRequestor reconnection logic
* add calling number to log metadata, as this will be frequently the key data given for troubleshooting
* add accountSid to log metadata
* make handshake timeout for ws connections configurable with default 1.5 secs
* rename env var
* fix bug prev checkin
* logging fixes
* consistent env naming
* Dial: handle incoming REFER on either leg by calling referHook, if configured
* lint
* modify payload of referHook
* support target.trunk on rest createCall api
* bugfix: gather partial result hook was not working
* lint
* handling of incoming REFER