The Call-ID of outbound registrations was the sip_gateway_sid, the same on
every SBC. When the regbot role moved to the other SBC, the registrar saw a
refresh of an existing binding from a different source address and Contact.
Some registrars 200 such a refresh without updating their routing, so inbound
calls to the registered trunk fail with 404 until the binding is recreated.
The Call-ID is now sip_gateway_sid@<sending SBC public IP>: stable across
refreshes and restarts of one SBC, new when the role moves, so a move looks
like a new registration. register_status also records the sending SBC as
sbcAddress.
With AWS_LIFECYCLE_DRAIN enabled the sidecar polls IMDS
autoscaling/target-lifecycle-state (the signal inbound drains on; detection
only, inbound completes the lifecycle hook). When the instance is being scaled
in, the regbot holder releases the lease while still running instead of after
the instance is gone; until now the draining SBC kept the registrations, so
carriers kept sending registration-trunk calls to an SBC that answers new
INVITEs with 503. It never claims the role back. Once another SBC has claimed
it, the draining SBC un-REGISTERs (Expires: 0) the bindings whose Contact
carries its own IP. Bindings with an AoR or realm Contact are left alone: the
new SBC sends the same Contact, and its REGISTER has already replaced ours.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix: set alert_type on OPTIONS-ping and registration alerts
The three writeAlerts() calls in sip-trunk-options-ping.js and the two in
regbot.js omitted alert_type, so the alerts were stored as 'undefined'.
Use AlertType.SIP_GATEWAY_OPTIONS_FAILURE / SIP_GATEWAY_REGISTRATION_FAILURE.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: bump @jambonz/time-series to ^0.5.4
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix: expire FS/RTP roster members on shared last-seen, not one SBC's local view
The active-fs, fs-service-url and active-rtp redis sets are shared by every
SBC in the cluster, but the expiry sweep in lib/options.js removed members
based solely on when THIS sidecar last received an OPTIONS ping from them.
An SBC that was taken out of active-sip (so FS/RTP servers stopped pinging
it) therefore deleted every feature server and rtpengine from the shared
rosters 60s later, rejecting all inbound calls until the other SBC re-added
them on its next ping cycle.
Each SBC now records the last ping it received per member in a shared redis
hash (<setName>:lastseen, member -> epoch ms) and the sweep removes a member
only when that shared timestamp is older than EXPIRES_INTERVAL. Members with
no shared timestamp are never expired by the sweep, so a parked or
mixed-version SBC cannot remove members the others are still hearing from.
Adds a unit test that reproduces the failure against the old code.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix: make sbc_addresses keepalive recurring instead of a single setTimeout
addSbcAddress() refreshes the row's last_updated and cleanSbcAddresses()
deletes rows older than DEAD_SBC_IN_SECOND (default 3600s), but app.js only
re-called addSbcAddress once, 15 minutes after connecting. The row then
went stale, and the next sidecar to start anywhere in the cluster deleted
the healthy SBC's row, so new sip realms were provisioned with one SBC IP
instead of two.
Run one recurring timer (SBC_PUBLIC_ADDRESS_KEEP_ALIVE_IN_MILISECOND, default
15 min) that refreshes this SBC's rows and only then reaps stale ones, so the
cleaner never runs ahead of this process's own keepalive. The timer is
unref'd and replaced (not stacked) on drachtio reconnect.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
- jambonzVersion is read from the schema_version table via the existing
db-helpers pool (now exposed on srf.locals); a DB error yields null
rather than failing discovery
- drachtioVersion is captured from the drachtio connect handshake and
stored on srf.locals
- README example updated
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix: add mutex to prevent concurrent regbot rebuilds (#2719)
updateCarrierRegbots was called via .then() (not awaited) from checkStatus,
allowing concurrent invocations to race on the shared regbots array. The
last writer wins, and concurrent cleanup can stop() regbots still in use,
permanently breaking their timer chain.
Add a boolean mutex (rebuildInProgress) with try/finally to skip concurrent
invocations. Promote rebuild completion log from debug to info for production
visibility. Add standalone test demonstrating both the bug and the fix.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* raise default minimum registration expires from 30s to 90s
Prevents overly aggressive re-registration cycles when the registrar
returns a short Expires value, reducing regbot churn.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
The base image was pinned to linux/amd64, so the arm64 build leg pulled the
amd64 image and failed with 'exec format error'. Remove the pin so each
platform builds on its native base (jambonz v11 arm64 support).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a cloud-driver buildx step and platforms: linux/amd64,linux/arm64 so the
published image supports arm64 (jambonz v11 arm64 support). Mirrors the house
pattern (upload-recordings / drachtio-server).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix timing issue of ephemeral gateways update/deletion
* Fix for potential regbot Zombie, and other concerns
* address performance concerns in regbot behavior
* use gw sid as call id and fix bug on updateing regbots
we must not mutate gw directly, as it is cached in the gateways array
and adding properties would cause the JSON.stringify comparison to detect a
false change on every checkStatus cycle, triggering unnecessary re-registrations.
* fix tests
add influx and also the registered now changes to remove previous reg on update
* better test scenario
still checks that we have 1 registered as expires now at 600s so it doesn't expire before the 65sec delay
* support draining feature server manually
* support draining feature server manually
* support CLI to add or remove feature server
* wip
* wip
* wip
* add redis key for feature server and integration test
* wip
* add support for registration trunks which result in a set of ephemeral sip gateways to be stored in redis
* wip
* refactor createEphemeralGateways into realtime dbhelpers
* minor
* update eslint