Files
jambonz-feature-server/test
Hoan Luu HuuandClaude Opus 5 db18b558ce add sip_reason_header to status Callback- #178 (#1579)
* feat: surface the SIP Reason header on call status events

Carriers fronting ISDN/E1 PRI trunks put the authoritative disconnect cause in
an RFC 3326 Reason header rather than in the SIP status line, e.g.

    SIP/2.0 408 Request Timeout
    Reason: Q.850 ;cause=18

Q.850 cause 18 is "no user responding" - nobody answered, not a platform fault.
Different causes also arrive under the same SIP status (503 with cause=38
network out of order, or cause=41 temporary failure), so the status code alone
cannot classify the outcome of the call.

drachtio relays the header intact and it is present on the response object, but
the feature server only read the status line from it, so the cause was lost at
the application boundary and never reached the call status webhook.

Rather than extract the header at each emit site, carry the SIP message that
caused the status change on the callStatusChange event and derive from it in one
place, so provisional responses, the 200, final failures, BYE and CANCEL are all
covered by the same code and future headers cost one line.

Note this partly overlaps the existing _extractCustomHeaders/sip_headers
passthrough: that already exposes a Reason header arriving on a BYE, but it only
runs on the hangup path, so nothing covered the outbound INVITE failure
responses where a Q.850 cause matters most.  sip_reason_header is a dedicated,
documented field that behaves the same on every status event.

sip_reason keeps meaning the status line phrase, and no key is added when there
is no Reason header, so existing consumers see an unchanged payload.

Adds a test:unit script so the smoke test runs without the docker testbed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: note that the Reason header arrives re-serialized, not verbatim

Verified live end-to-end against a cluster, capturing both the external leg and
the leg into the feature server:

    14.226.234.142 -> 10.0.197.31:5060   Reason: Q.850 ;cause=31   (as sent)
    10.0.197.31:5060 -> :5070            Reason: Q.850;cause=31    (to fs)

Proxying re-serializes the header, normalizing the optional whitespace RFC 3326
permits around ';'. The same normalization appears in customer captures from an
unrelated deployment, so this is drachtio behaviour, not cluster-specific.

sip_reason_header is therefore the header as this process received it, not the
carrier's exact bytes - worth stating outright, since the spacing inconsistency
is exactly what consumers ask about, and a consumer who string-matches on the
spaced form would silently never match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: clear a stale Reason header in redis, pass it on the alloc path, run the tests in CI

Three problems found reviewing the earlier commits.

1. The redis call record kept a stale header forever. updateCallStatus assigns
   unconditionally, but toJSON only emits truthy values, and the same object is
   written to the redis hash with hmset - a MERGE. An absent key therefore left
   the previous status change's header in place, readable via GET /Calls/:sid:
   a leg that got 183 + "Reason: Q.850;cause=31", then answered on a clean 200
   and completed on a plain BYE, reported a Q.850 temporary-failure cause for a
   call that ended normally. The webhooks were right; only the call record was
   wrong. Storing absence as '' overwrites it. The existing "does not linger"
   test passed throughout because it only exercised toJSON in memory, so this
   adds one that pins what survives the redis filter (verified: it fails
   without the fix).

2. The endpoint-allocation failure path already extracted the Reason header and
   relayed it to the SBC, but called _notifyCallStatusChange without it - so a
   FreeSWITCH 488 with "Reason: Q.850;cause=88 INCOMPATIBLE_DESTINATION" told
   the SBC the cause and the application nothing, which is exactly the case
   this feature exists to expose. That error is an fsmrf object rather than a
   SipMessage, so it cannot go through msg (the guard in
   reasonHeaderFromSipMessage would silently return undefined); the event now
   takes an explicit sipReasonHeader for callers holding the value already.

3. The test:unit script added with these tests was never wired into CI - the
   workflow runs jslint and npm test, and npm test enumerates its files
   explicitly and skips test/unit entirely, so the suite would have rotted
   unnoticed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: clear the stale header on the CallSession path too, and keep the CANCEL

Both from PR review.

1. The previous fix only worked for SingleDialer. Storing '' on the instance is
   enough for callers that write the instance itself, but CallSession writes
   toJSON(), and toJSON() drops falsy values on purpose so the key stays out of
   the webhook payload - so '' never reached hmset and the stale header
   survived. Reproduced:

       after 183, toJSON has: Q.850;cause=31
       raw instance value:    ""
       instance   -> redis has key: true    (SingleDialer, clears)
       toJSON()   -> redis has key: false   (CallSession, stale persists)

   The split is by session class, not call direction: place-outdial is required
   only by dial.js, so SingleDialer covers dial-verb child legs while inbound,
   REST-created and adulting all inherit CallSession._notifyCallStatusChange -
   including the REST outdial path this feature was written for.

   Adds CallInfo.toRedisJSON(), a named projection for the merge-semantics
   store, so the divergence from toJSON() lives in one documented place rather
   than being rediscovered at each call site. The unit test asserted against the
   raw instance, which is why it passed while the real path was broken; it now
   goes through both writers and fails if either regresses.

2. The caller-abandoned race dropped the CANCEL. middleware.js had it in hand
   and discarded it, so the constructor's _onCancel() passed nothing whenever
   the CANCEL beat the application fetch - the header landed or not depending on
   timing. Worth closing because the 487 and its 'Request Terminated' phrase are
   both ours, so every abandoned inbound call looks identical: a
   Reason: SIP;cause=200;text="Call completed elsewhere" is what separates a
   forked branch losing the race from a caller who gave up, and today both are
   just no-answer.

   Confirmed on a deployed srf (5.0.27) that this is not a no-op:
   copyUASHeaderToUACForOnlyCancel forwards a hardcoded ['Reason', 'X-Reason']
   when proxying a CANCEL, so the header does reach us.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: keep the Reason header out of the redis call record

Reversing the earlier approach after a closer look at who reads what.

The header was reaching the redis call record simply because CallInfo feeds two
sinks with opposite semantics: the status webhook is an EVENT (each POST is
independent, an absent key means "not in this event") while the redis record is
STATE written with hmset, a MERGE (an absent key means "keep what was there").
sipReasonHeader is the first field here that can legitimately go from set back
to unset - sipStatus and sipReason are only ever overwritten - so it was the
first to expose the mismatch, and two rounds of fixes had to chase it because
the two writers project from different bases.

Rather than keep managing that, exclude it: redis holds calls that are still
live, where there is usually no interesting cause yet, and by the time there is
one the call is over. Call history is served from RecentCalls (the CDR), which
is what the webapp reads - not GET /Calls. So the field earned very little there
while costing an invariant that has to be remembered forever.

Worth being explicit that this is NOT simply a revert: dropping the '' would
only have stopped the empty value being written. When the header is PRESENT it
still reached redis through both writers - toJSON includes it, and SingleDialer
writes the instance's own properties - and was then never cleared. Keeping it out
takes the same machinery as clearing it, just inverted, so this is a choice about
where the field belongs rather than a saving.

WEBHOOK_ONLY_FIELDS names that intent in one place, with the reasoning, so the
next field with the same property has somewhere obvious to go. The test asserts
both writers, since they project from different bases and checking one would
pass while the other still wrote the field (verified: it fails if the exclusion
is removed).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: apply the redis projection at the boundary, and drop dead adulting plumbing

From PR review.

1. There was a THIRD call-record writer. Asking each call site to remember the
   projection was the wrong shape, and review found the proof: this class writes
   the record from the status change, the recording flag AND (private only)
   _persistConferenceState, and the last one still passed toJSON() straight
   through. A conferenced caller whose BYE carried Reason: Q.850;cause=16 would
   have that written into the call hash on the next conference-state persist,
   where hmset merges and nothing can ever clear it.

   Fixed by wrapping updateCallStatus once where it is bound, so every write is
   projected and a newly added writer cannot bypass it by forgetting to ask. The
   call sites go back to passing plain shapes. This is a class of miss that a
   unit test cannot catch - the contract test passed the whole time - so the fix
   is structural rather than another assertion.

2. The msg/byeReq parameters added to AdultingCallSession were dead code. The
   only path that reacts to the far-end BYE there is the inline
   sd.dlg.on('destroy') handler, which calls _callReleased() and discards the
   request; _hangup is reachable only from CallSession.hangup() (LCC), which
   passes nothing. The Reason header on that leg does reach the webhook, via
   SingleDialer's own destroy handler - so the plumbing was not just unused but
   misleading about which path carries it. Removed, with a comment at the handler
   recording where the event actually comes from so it does not get re-added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: correct the projection comment in SingleDialer

Copied verbatim from CallSession, where the list of writers (status change,
recording flag, conference state) is accurate. SingleDialer has one writer, so
the comment described code that is not there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 07:35:38 -04:00
..
2025-06-28 15:01:09 -04:00
2025-06-28 15:01:09 -04:00
2023-06-03 08:16:05 -04:00
2020-01-07 10:34:03 -05:00
2023-06-03 08:16:05 -04:00
2023-12-28 14:59:59 -05:00
2023-06-03 09:09:49 -04:00
2023-06-09 14:54:53 -04:00
2023-06-03 08:16:05 -04:00
2023-02-06 08:06:41 -05:00