The Call-ID of outbound registrations was the sip_gateway_sid, the same on
every SBC. When the regbot role moved to the other SBC, the registrar saw a
refresh of an existing binding from a different source address and Contact.
Some registrars 200 such a refresh without updating their routing, so inbound
calls to the registered trunk fail with 404 until the binding is recreated.
The Call-ID is now sip_gateway_sid@<sending SBC public IP>: stable across
refreshes and restarts of one SBC, new when the role moves, so a move looks
like a new registration. register_status also records the sending SBC as
sbcAddress.
With AWS_LIFECYCLE_DRAIN enabled the sidecar polls IMDS
autoscaling/target-lifecycle-state (the signal inbound drains on; detection
only, inbound completes the lifecycle hook). When the instance is being scaled
in, the regbot holder releases the lease while still running instead of after
the instance is gone; until now the draining SBC kept the registrations, so
carriers kept sending registration-trunk calls to an SBC that answers new
INVITEs with 503. It never claims the role back. Once another SBC has claimed
it, the draining SBC un-REGISTERs (Expires: 0) the bindings whose Contact
carries its own IP. Bindings with an AoR or realm Contact are left alone: the
new SBC sends the same Contact, and its REGISTER has already replaced ours.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix timing issue of ephemeral gateways update/deletion
* Fix for potential regbot Zombie, and other concerns
* address performance concerns in regbot behavior