Api-realtime SIP Refer not reaching far end

Realtime SIP: in-dialog REFER is never delivered when the dialog has a Record-Route

POST /v1/realtime/calls/{call_id}/refer returns HTTP 200 and raises no error on the realtime session, but no REFER reaches the far end and no transfer happens. This occurs whenever the SIP dialog carries a Record-Route header, which is the normal case for any call arriving through a carrier that proxies signalling, including Twilio Programmable Voice (the Dial + Sip verbs).

The failure is completely silent. The API reports success, the session reports nothing, and the caller sits in silence until they hang up.

Impact

Every call transfer fails, with no error surfaced anywhere in the API. In our case this ran for two days before we could prove where it was going wrong, because a 200 from /refer is documented to mean the REFER was relayed to the SIP provider.

A directly-connected SIP peer (no proxy, therefore no Record-Route) receives the REFER correctly, so basic integration tests pass while every real carrier call fails.

Reproduction

No carrier account needed. Two runs of one script, changing a single header.

Prerequisite: a webhook handler for realtime.call.incoming that accepts the call and, a few seconds later, calls POST /v1/realtime/calls/{call_id}/refer with a target_uri of tel:+15551234567. Point one project’s webhook at it.

The SIP client: a ~130 line Node script (Node 20+, no dependencies) that opens one TLS connection, sends an INVITE, ACKs the 200, and prints any in-dialog request it receives. It is here: PASTE_YOUR_GIST_LINK_HERE

Run it twice:

SIP_PROJECT=proj_xxxxx node repro-uac.mjs
RESULT  Record-Route=none  ->  REFER RECEIVED
SIP_PROJECT=proj_xxxxx RECORD_ROUTE='<sip:203.0.113.9;r2=on;lr>' node repro-uac.mjs
RESULT  Record-Route=<sip:203.0.113.9;r2=on;lr>  ->  REFER NOT received

203.0.113.9 is TEST-NET-3, deliberately unreachable. Both runs return HTTP 200 from /refer and produce no error on the realtime session.

Results matrix

# Record-Route on the INVITE In-dialog REFER
1 absent received, ~3.5 s after the API call
2 <sip:203.0.113.9;r2=on;lr> never arrives
3 absent received
4 <sip:203.0.113.9;transport=tls;lr> never arrives

Runs 3 and 4 alternate with 1 and 2 to show it is deterministic, not a transient.

Where it goes wrong

Repeating run 2 with the Record-Route pointing at an address we actually controlled (a public TCP endpoint) showed the mechanism.

  1. The REFER is not sent over the established dialog connection. A new outbound connection is opened to the route-set target. We observed it connect, then retry about 8 seconds later.
  2. That connection always uses TLS, even when the route URI says transport=tcp. A plain TCP listener received a TLS ClientHello: 16 03 01 05 e9 01 00 05 e5 03 03 ...
  3. The server certificate is validated. Terminating TLS with a self-signed certificate gets the handshake aborted with sslv3 alert bad certificate (TLS alert 42).
  4. Then it gives up silently. One retry, then nothing. No error on the session, and the 200 was returned several seconds before any of this was attempted.

Why this breaks Twilio Programmable Voice

Twilio’s INVITE to sip.api.openai.com carries these headers:

Record-Route: <sip:203.0.113.2;r2=on;lr>
Contact:      <sip:+1XXXXXXXXXX@10.x.x.x:5060;transport=udp>
Via:          SIP/2.0/UDP 10.x.x.x:5060;rport=5060;branch=...

The Record-Route is a bare sip: URI with no transport parameter and no port, which per RFC 3263 resolves to UDP port 5060. The REFER is then attempted as TLS against port 5060, where Twilio serves plain SIP (TLS is on 5061). The handshake cannot complete, so the REFER is never delivered, Twilio’s referUrl is never invoked, and nothing reports a failure.

This also explains why a direct peer works: with no Record-Route there is no second connection to fail, and the REFER goes down the existing dialog.

Expected behaviour

Three things, in order of how much they would have helped.

  1. /refer should not return 200 before delivery is attempted, or a delivery failure should surface as an error on the realtime session. The silent success is what made this take two days to locate. A 202 Accepted plus a later session event would be enough.
  2. The transport in the route URI should be honoured, rather than TLS being forced.
  3. Ideally the REFER should reuse the dialog’s existing connection, as most SIP stacks do for in-dialog requests where the route set resolves to the peer that established it.

Environment

  • sip.api.openai.com:5061, transport=tls (US endpoint, not sip-eu)
  • Carrier: Twilio Programmable Voice, Dial + Sip verbs, with referUrl set
  • target_uri: both tel:+E.164 and sip:user@host fail identically
  • Reproduced across 6 OpenAI projects, 5 API keys (sk-proj-* and sk-svcacct-*), inbound and outbound calls, and 3 destination numbers
  • Not an SDK issue: the repro script uses no OpenAI SDK, only node:tls plus one HTTP POST
  • The /refer response carries no x-request-id; cf-ray is the only per-request handle available

One possible workaround, still looking at others

If your carrier exposes a call-control API, bypass the REFER: detect Record-Route on the incoming INVITE and redirect the caller’s leg directly instead. On Twilio that is POST /2010-04-01/Accounts/{Sid}/Calls/{CallSid}.json with a Twiml parameter containing a Dial + Number, applied to the caller’s leg. Note that leg is the parent for inbound calls but the child for calls where you dialled the SIP leg first, so resolve it rather than assuming.

this issue seems to be resolved. We are no longer seeing any of the record-route or other issues as of this morning. We were told by openai support that our logs should a defect or misalignment with the document, but no acknowledgment of a fix or intended change.

Thanks for coming back to let us know. The community appreciates it.

Happy coding…