Hoan Luu HuuandClaude Opus 5 db18b558ce add sip_reason_header to status Callback- #178 (#1579)
* feat: surface the SIP Reason header on call status events

Carriers fronting ISDN/E1 PRI trunks put the authoritative disconnect cause in
an RFC 3326 Reason header rather than in the SIP status line, e.g.

    SIP/2.0 408 Request Timeout
    Reason: Q.850 ;cause=18

Q.850 cause 18 is "no user responding" - nobody answered, not a platform fault.
Different causes also arrive under the same SIP status (503 with cause=38
network out of order, or cause=41 temporary failure), so the status code alone
cannot classify the outcome of the call.

drachtio relays the header intact and it is present on the response object, but
the feature server only read the status line from it, so the cause was lost at
the application boundary and never reached the call status webhook.

Rather than extract the header at each emit site, carry the SIP message that
caused the status change on the callStatusChange event and derive from it in one
place, so provisional responses, the 200, final failures, BYE and CANCEL are all
covered by the same code and future headers cost one line.

Note this partly overlaps the existing _extractCustomHeaders/sip_headers
passthrough: that already exposes a Reason header arriving on a BYE, but it only
runs on the hangup path, so nothing covered the outbound INVITE failure
responses where a Q.850 cause matters most.  sip_reason_header is a dedicated,
documented field that behaves the same on every status event.

sip_reason keeps meaning the status line phrase, and no key is added when there
is no Reason header, so existing consumers see an unchanged payload.

Adds a test:unit script so the smoke test runs without the docker testbed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: note that the Reason header arrives re-serialized, not verbatim

Verified live end-to-end against a cluster, capturing both the external leg and
the leg into the feature server:

    14.226.234.142 -> 10.0.197.31:5060   Reason: Q.850 ;cause=31   (as sent)
    10.0.197.31:5060 -> :5070            Reason: Q.850;cause=31    (to fs)

Proxying re-serializes the header, normalizing the optional whitespace RFC 3326
permits around ';'. The same normalization appears in customer captures from an
unrelated deployment, so this is drachtio behaviour, not cluster-specific.

sip_reason_header is therefore the header as this process received it, not the
carrier's exact bytes - worth stating outright, since the spacing inconsistency
is exactly what consumers ask about, and a consumer who string-matches on the
spaced form would silently never match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: clear a stale Reason header in redis, pass it on the alloc path, run the tests in CI

Three problems found reviewing the earlier commits.

1. The redis call record kept a stale header forever. updateCallStatus assigns
   unconditionally, but toJSON only emits truthy values, and the same object is
   written to the redis hash with hmset - a MERGE. An absent key therefore left
   the previous status change's header in place, readable via GET /Calls/:sid:
   a leg that got 183 + "Reason: Q.850;cause=31", then answered on a clean 200
   and completed on a plain BYE, reported a Q.850 temporary-failure cause for a
   call that ended normally. The webhooks were right; only the call record was
   wrong. Storing absence as '' overwrites it. The existing "does not linger"
   test passed throughout because it only exercised toJSON in memory, so this
   adds one that pins what survives the redis filter (verified: it fails
   without the fix).

2. The endpoint-allocation failure path already extracted the Reason header and
   relayed it to the SBC, but called _notifyCallStatusChange without it - so a
   FreeSWITCH 488 with "Reason: Q.850;cause=88 INCOMPATIBLE_DESTINATION" told
   the SBC the cause and the application nothing, which is exactly the case
   this feature exists to expose. That error is an fsmrf object rather than a
   SipMessage, so it cannot go through msg (the guard in
   reasonHeaderFromSipMessage would silently return undefined); the event now
   takes an explicit sipReasonHeader for callers holding the value already.

3. The test:unit script added with these tests was never wired into CI - the
   workflow runs jslint and npm test, and npm test enumerates its files
   explicitly and skips test/unit entirely, so the suite would have rotted
   unnoticed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: clear the stale header on the CallSession path too, and keep the CANCEL

Both from PR review.

1. The previous fix only worked for SingleDialer. Storing '' on the instance is
   enough for callers that write the instance itself, but CallSession writes
   toJSON(), and toJSON() drops falsy values on purpose so the key stays out of
   the webhook payload - so '' never reached hmset and the stale header
   survived. Reproduced:

       after 183, toJSON has: Q.850;cause=31
       raw instance value:    ""
       instance   -> redis has key: true    (SingleDialer, clears)
       toJSON()   -> redis has key: false   (CallSession, stale persists)

   The split is by session class, not call direction: place-outdial is required
   only by dial.js, so SingleDialer covers dial-verb child legs while inbound,
   REST-created and adulting all inherit CallSession._notifyCallStatusChange -
   including the REST outdial path this feature was written for.

   Adds CallInfo.toRedisJSON(), a named projection for the merge-semantics
   store, so the divergence from toJSON() lives in one documented place rather
   than being rediscovered at each call site. The unit test asserted against the
   raw instance, which is why it passed while the real path was broken; it now
   goes through both writers and fails if either regresses.

2. The caller-abandoned race dropped the CANCEL. middleware.js had it in hand
   and discarded it, so the constructor's _onCancel() passed nothing whenever
   the CANCEL beat the application fetch - the header landed or not depending on
   timing. Worth closing because the 487 and its 'Request Terminated' phrase are
   both ours, so every abandoned inbound call looks identical: a
   Reason: SIP;cause=200;text="Call completed elsewhere" is what separates a
   forked branch losing the race from a caller who gave up, and today both are
   just no-answer.

   Confirmed on a deployed srf (5.0.27) that this is not a no-op:
   copyUASHeaderToUACForOnlyCancel forwards a hardcoded ['Reason', 'X-Reason']
   when proxying a CANCEL, so the header does reach us.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: keep the Reason header out of the redis call record

Reversing the earlier approach after a closer look at who reads what.

The header was reaching the redis call record simply because CallInfo feeds two
sinks with opposite semantics: the status webhook is an EVENT (each POST is
independent, an absent key means "not in this event") while the redis record is
STATE written with hmset, a MERGE (an absent key means "keep what was there").
sipReasonHeader is the first field here that can legitimately go from set back
to unset - sipStatus and sipReason are only ever overwritten - so it was the
first to expose the mismatch, and two rounds of fixes had to chase it because
the two writers project from different bases.

Rather than keep managing that, exclude it: redis holds calls that are still
live, where there is usually no interesting cause yet, and by the time there is
one the call is over. Call history is served from RecentCalls (the CDR), which
is what the webapp reads - not GET /Calls. So the field earned very little there
while costing an invariant that has to be remembered forever.

Worth being explicit that this is NOT simply a revert: dropping the '' would
only have stopped the empty value being written. When the header is PRESENT it
still reached redis through both writers - toJSON includes it, and SingleDialer
writes the instance's own properties - and was then never cleared. Keeping it out
takes the same machinery as clearing it, just inverted, so this is a choice about
where the field belongs rather than a saving.

WEBHOOK_ONLY_FIELDS names that intent in one place, with the reasoning, so the
next field with the same property has somewhere obvious to go. The test asserts
both writers, since they project from different bases and checking one would
pass while the other still wrote the field (verified: it fails if the exclusion
is removed).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: apply the redis projection at the boundary, and drop dead adulting plumbing

From PR review.

1. There was a THIRD call-record writer. Asking each call site to remember the
   projection was the wrong shape, and review found the proof: this class writes
   the record from the status change, the recording flag AND (private only)
   _persistConferenceState, and the last one still passed toJSON() straight
   through. A conferenced caller whose BYE carried Reason: Q.850;cause=16 would
   have that written into the call hash on the next conference-state persist,
   where hmset merges and nothing can ever clear it.

   Fixed by wrapping updateCallStatus once where it is bound, so every write is
   projected and a newly added writer cannot bypass it by forgetting to ask. The
   call sites go back to passing plain shapes. This is a class of miss that a
   unit test cannot catch - the contract test passed the whole time - so the fix
   is structural rather than another assertion.

2. The msg/byeReq parameters added to AdultingCallSession were dead code. The
   only path that reacts to the far-end BYE there is the inline
   sd.dlg.on('destroy') handler, which calls _callReleased() and discards the
   request; _hangup is reachable only from CallSession.hangup() (LCC), which
   passes nothing. The Reason header on that leg does reach the webhook, via
   SingleDialer's own destroy handler - so the plumbing was not just unused but
   misleading about which path carries it. Removed, with a comment at the handler
   recording where the event actually comes from so it does not get re-added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: correct the projection comment in SingleDialer

Copied verbatim from CallSession, where the list of writers (status change,
recording flag, conference state) is accurate. SingleDialer has one writer, so
the comment described code that is not there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 07:35:38 -04:00
2023-12-11 08:37:36 -05:00
2020-01-07 10:34:03 -05:00
2021-08-30 17:12:00 -04:00
2025-06-28 15:01:09 -04:00
2025-08-21 14:09:20 -04:00
2026-03-20 08:57:16 -04:00
2024-06-28 09:05:05 -04:00

jambonz-feature-server CI

This application implements the core feature server of the jambones platform.

Note: If you are a developer looking to work on the code please read our how-to for that.

Configuration

Configuration is provided via environment variables:

variable meaning required?
AWS_ACCESS_KEY_ID aws access key id, used for TTS/STT as well SNS notifications no
AWS_REGION aws region no
AWS_SECRET_ACCESS_KEY aws secret access key, used per above no
AWS_SNS_TOPIC_ARN aws sns topic arn that scale-in lifecycle notifications will be published to no
DRACHTIO_HOST ip address of drachtio server (typically '127.0.0.1') yes
DRACHTIO_PORT listening port of drachtio server for control connections (typically 9022) yes
DRACHTIO_SECRET shared secret yes
ENABLE_METRICS if 1, metrics will be generated no
ENCRYPTION_SECRET secret for credential encryption(JWT_SECRET is deprecated) yes
GOOGLE_APPLICATION_CREDENTIALS path to gcp service key file yes
HTTP_PORT tcp port to listen on for API requests from jambonz-api-server yes
HTTP_IP IP Address for API requests from jambonz-api-server no
JAMBONES_GATHER_EARLY_HINTS_MATCH if true and hints are provided, gather will opportunistically review interim transcripts if possible to reduce ASR latency no
JAMBONES_FREESWITCH IP:port:secret for Freeswitch server (e.g. '127.0.0.1:8021:JambonzR0ck$' yes
JAMBONES_HOLD_UNHOLD_EVENTS if set, surface in-dialog hold/un-hold: emits a 'hold'/'unhold' event and, when sipRequestWithinDialogHook is configured, delivers it to the application. Disabled when unset no
JAMBONES_LOGLEVEL log level for application, 'info' or 'debug' no
JAMBONES_MYSQL_HOST mysql host yes
JAMBONES_MYSQL_USER mysql username yes
JAMBONES_MYSQL_PASSWORD mysql password yes
JAMBONES_MYSQL_DATABASE mysql data yes
JAMBONES_MYSQL_CONNECTION_LIMIT mysql connection limit no
JAMBONES_NETWORK_CIDR CIDR of private network that feature server is running in (e.g. '172.31.0.0/16') yes
JAMBONES_REDIS_HOST redis host yes
JAMBONES_REDIS_PORT redis port yes
JAMBONES_SBCS list of IP addresses (on the internal network) of SBCs, comma-separated yes
STATS_HOST ip address of metrics host (usually '127.0.0.1' since telegraf is installed locally no
STATS_PORT listening port for metrics host no
STATS_PROTOCOL 'tcp' or 'udp' no
STATS_TELEGRAF if 1, metrics will be generated in telegraf format no
JAMBONZ_RECORD_WS_BASE_URL recording websocket URL to send the recording audio no
JAMBONZ_RECORD_WS_USERNAME recording websocket username no
JAMBONZ_RECORD_WS_PASSWORD recording websocket password no
ANCHOR_MEDIA_ALWAYS keep media on media server no
JAMBONZ_DISABLE_DIAL_PAI_HEADER control P-Asserted-Identity header on B-Leg no

running under pm2

Typically, this application runs under pm2 using an ecosystem.config.js file similar to this:

module.exports = {
  apps : [
  {
    name: 'jambonz-feature-server',
    cwd: '/home/admin/apps/jambonz-feature-server',
    script: 'app.js',
    instance_var: 'INSTANCE_ID',
    out_file: '/home/admin/.pm2/logs/jambonz-feature-server.log',
    err_file: '/home/admin/.pm2/logs/jambonz-feature-server.log',
    exec_mode: 'fork',
    instances: 1,
    autorestart: true,
    watch: false,
    max_memory_restart: '1G',
    env: {
      NODE_ENV: 'production',
      GOOGLE_APPLICATION_CREDENTIALS: '/home/admin/credentials/gcp.json',
      AWS_ACCESS_KEY_ID: 'XXXXXXXXXXXX',
      AWS_SECRET_ACCESS_KEY: 'YYYYYYYYYYYYYYYYYYYYY',
      AWS_REGION: 'us-west-1',
      ENABLE_METRICS: 1,
      STATS_HOST: '127.0.0.1',
      STATS_PORT: 8125,
      STATS_PROTOCOL: 'tcp',
      STATS_TELEGRAF: 1,
      AWS_SNS_TOPIC_ARN: 'arn:aws:sns:us-west-1:xxxxxxxxxxx:terraform-20201107200347128600000002',
      JAMBONES_NETWORK_CIDR: '172.31.0.0/16',
      JAMBONES_MYSQL_HOST: 'aurora-cluster-jambonz.cluster-yyyyyyyyyyy.us-west-1.rds.amazonaws.com',
      JAMBONES_MYSQL_USER: 'admin',
      JAMBONES_MYSQL_PASSWORD: 'foobarbz',
      JAMBONES_MYSQL_DATABASE: 'jambones',
      JAMBONES_MYSQL_CONNECTION_LIMIT: 10,
      JAMBONES_REDIS_HOST: 'jambonz.zzzzzzz.0001.usw1.cache.amazonaws.com',
      JAMBONES_REDIS_PORT: 6379,
      JAMBONES_LOGLEVEL: 'debug',
      HTTP_PORT: 3000,
      DRACHTIO_HOST: '127.0.0.1',
      DRACHTIO_PORT: 9022,
      DRACHTIO_SECRET: 'sharedsecret',
      JAMBONES_SBCS: '172.31.32.10',
      JAMBONES_FREESWITCH: '127.0.0.1:8021:sharedsecret'
    }
  }]
};

Running the test suite

Please see this.

S
Description
Core telephony feature server for the jambones platform
Readme MIT
10 MiB
Languages
JavaScript 99.9%