A ghost answer is a call your carrier marks as answered with a 200 OK that carries no SDP, after which your own switch tears the call down the moment it is answered. Asterisk reports a healthy answer rate while your dialer logs a pile of zero-second no-answers. Here is how I tracked one carrier doing this, with the Homer SQL, Asterisk log lines and CDR joins I used to prove it.
Tested on: ViciDial 2.14 (SVN r3988) with Asterisk 18.26 (chan_sip), MariaDB 10.11, and Homer 7 using heplify-server 1.59.6 with the homer7 schema on PostgreSQL 16. In this install hep_proto_1_call.raw is a varchar. If yours stores raw as bytea, wrap it as convert_from(raw, 'UTF8') before applying the text functions below.
What is a ghost answer?
On a normal outbound SIP call the far end answers with 200 OK and an SDP body that says where to send audio and which codec to use. Your side sends ACK, RTP flows, and the call has a conversation.
A ghost answer breaks that. The carrier sends 200 OK with an empty body. No SDP, no media address, no codec. Depending on what happened earlier in the call, Asterisk either falls back to the media details from an earlier 183 Session Progress, or it gives up and sends BYE right after its own ACK. In the second case the gap between ACK and BYE is typically a fraction of a millisecond.
Is an SDP-less 200 OK legal? It depends. With reliable provisional responses (100rel and PRACK, RFC 3262) the answer can arrive in a 183, and the 200 does not have to repeat it. Without them, RFC 3261 section 13.2.1 is blunt: when the offer is in the INVITE, the answer must be in a reliable message, "For this specification, that is only the final 2xx response to that INVITE."
In my capture window, all 23,167 initial INVITEs to my three carriers carried an SDP offer, none advertised 100rel, no 180 or 183 carried RSeq or Require: 100rel, and Homer saw zero PRACKs. So these 200s were missing the SDP answer that this exchange required. Beyond the RFC, what makes it a ghost answer is scale and outcome: one route doing it on most answers while your other routes do it on none, and a large share of those calls being torn down instantly.
How did the problem first show up in my numbers?
I run a press-1 survey campaign that splits outbound traffic randomly across three carriers, one third each, on the same lead lists. One morning the dialer-side answer rate on one carrier, Carrier A, was half of what it had been the day before. The other two had not moved.
The odd part was that the Asterisk-side answer rate had barely changed. ViciDial's vicidial_carrier_log table records the DIALSTATUS Asterisk saw for each call. Joining that to vicidial_log (the disposition the dialer finally recorded for the lead) showed the gap:
| Day | Carrier | Dials | Asterisk ANSWER rate | Dialer answer rate | ANSWER but logged NA |
|---|---|---|---|---|---|
| Day 1 | Carrier A | 88,019 | 32.2% | 31.2% | 83 |
| Day 1 | Carrier B | 87,962 | 38.3% | 36.2% | 56 |
| Day 1 | Carrier C | 87,483 | 40.0% | 39.1% | 25 |
| Day 2 | Carrier A | 84,793 | 31.6% | 15.6% | 13,041 |
| Day 2 | Carrier B | 85,337 | 38.3% | 36.1% | 28 |
| Day 2 | Carrier C | 85,363 | 39.1% | 38.2% | 13 |
"Asterisk ANSWER rate" is the share of dials where Dial() returned ANSWER. It is measured on my side, not taken from the carrier's reports. "Dialer answer rate" is the share of dials whose ViciDial disposition was anything other than a no-answer, busy, disconnect or drop.
Carrier A's ANSWER rate stayed around 32%. But overnight, roughly half of those answers turned into NA records in vicidial_log, every one of them with length_in_sec = 0. Over the next seven dialing days Carrier A produced between 12,963 and 17,459 of these per day. The other two carriers stayed between 10 and 60 per day.
The signal is the disagreement between the two rates; either one alone looks unremarkable.
That table is from an incident a few months ago. My Homer instance only keeps three days of SIP, so I no longer have the signaling for those days. The same carrier was still doing it when I wrote this, and every SIP-level number below comes from a live capture of that current traffic. That the July discrepancy had the same cause is an inference, not something I can still prove from packets.
How do I measure the gap between Asterisk answers and dialer answers?
You need two things: the per-call DIALSTATUS from Asterisk, and a way to know which carrier each call took. ViciDial gives you the first in vicidial_carrier_log. For the second, I log the carrier at dial time into a small attribution table (carrier_route_log here: uniqueid, carrier, call_date) from the dialplan. That table joins to vicidial_carrier_log on uniqueid, because both are written with the dial-time uniqueid.
The join to vicidial_log is the tricky part. On my install, answered calls in vicidial_log do not carry the dial-time uniqueid: on one day, 9,542 answered-status rows had just 1 match in my attribution table, while all 16,744 no-answer/busy/drop rows matched. A uniqueid join therefore loses exactly the calls you care about. I join on lead_id plus a time window instead, and count matches per call so I can see when the join misbehaves:
SELECT d, carrier, COUNT(*) dials,
SUM(ds='ANSWER') ast_answer,
ROUND(100*SUM(ds='ANSWER')/COUNT(*),1) ast_answer_pct,
SUM(nmatch=0) unmatched, SUM(nmatch>1) multi,
SUM(ds='ANSWER' AND st='NA') ans_but_na,
SUM(ds='ANSWER' AND st='NA' AND len=0) ans_but_na_0s,
ROUND(100*SUM(nmatch=1 AND st NOT IN
('NA','AB','ADC','DC','B','CONGST','DROP'))/COUNT(*),1) dialer_answer_pct
FROM (
SELECT DATE(c.call_date) d, r.carrier, c.uniqueid, c.dialstatus ds,
COUNT(vl.uniqueid) nmatch, MAX(vl.status) st, MAX(vl.length_in_sec) len
FROM vicidial_carrier_log c
JOIN carrier_route_log r USING(uniqueid)
-- scope the dial cohort to the campaign BEFORE joining dispositions
JOIN vicidial_list l ON l.lead_id = c.lead_id
JOIN vicidial_lists ls ON ls.list_id = l.list_id
AND ls.campaign_id = 'YOUR_CAMPAIGN'
LEFT JOIN vicidial_log vl
ON vl.lead_id = c.lead_id AND vl.campaign_id = 'YOUR_CAMPAIGN'
AND vl.call_date BETWEEN c.call_date - INTERVAL 5 SECOND
AND c.call_date + INTERVAL 120 SECOND
WHERE c.call_date >= CURDATE() - INTERVAL 1 DAY
AND c.call_date < CURDATE()
GROUP BY c.uniqueid) x
GROUP BY d, carrier;
The campaign filter has to sit on the dial cohort, not just on the vicidial_log side. Otherwise other campaigns' dials stay in the denominator and simply show up as unmatched. I scope by the campaign that owns the lead's list. On my install that removed 98 dials from another campaign on one day out of more than two million, and no other day had any. Before trusting it, check three things for your interval:
- One dial per lead per day. On every day in my table, zero leads were dialed more than once, so a
lead_idjoin cannot pick up the wrong attempt. Check yours with aGROUP BY lead_id, DATE(call_date) HAVING COUNT(*) > 1. multiis near zero. Mine never exceeded 1 per carrier per day.unmatchedis explained. Mine were mostlyBUSYcalls whosevicidial_logrow landed outside the window. Between 1.8% and 5.6% ofANSWERrows were unmatched on the day I checked, so the dialer answer rate is a slight undercount for all three carriers.
The window is wide on purpose. In the sample call I trace below, the vicidial_log row was stamped 43 seconds after the dial started.
Treat ans_but_na as a screening metric, not a ghost counter. An ANSWER that ends up NA with zero length is what a ghost answer looks like in the CDRs, but other faults could produce the same row. On a clean route it is a few dozen per day out of tens of thousands of answers. When one route jumps to thousands, go to the SIP capture to confirm what is happening.
What does a ghost answer look like in Asterisk logs?
chan_sip logs a warning when it receives an SDP-less 200 OK it cannot use:
WARNING[10434][C-01204f61] chan_sip.c: Received response: "200 OK" from 'carrier_a' without SDP
Here is one complete ghost call from the full log, trimmed to the interesting lines:
[09:02:16] app_dial.c: Called SIP/carrier_a/39XXXXXXXXXX
[09:02:19] app_dial.c: SIP/carrier_a-00d92d01 is making progress passing it to Local/...;2
[09:02:19] app_dial.c: Call on SIP/carrier_a-00d92d01 placed on hold
[09:02:19] res_musiconhold.c: Started music on hold, class 'default', on channel 'Local/...;2'
[09:02:28] chan_sip.c: Received response: "200 OK" from 'carrier_a' without SDP
[09:02:28] app_dial.c: SIP/carrier_a-00d92d01 answered Local/...;2
[09:02:28] bridge_channel.c: Channel SIP/carrier_a-00d92d01 joined 'simple_bridge' basic-bridge <...>
[09:02:28] bridge_channel.c: Channel SIP/carrier_a-00d92d01 left 'simple_bridge' basic-bridge <...>
[09:02:28] pbx.c: Executing [h@default:1] AGI("Local/...;2", "agi://127.0.0.1:4577/call_log--HVcauses--PRI-----NODEBUG-----16-----ANSWER-----11-----0-----SIP 200 OK)")
The SIP channel joins the bridge and leaves it in the same second. The hangup handler records DIALSTATUS=ANSWER after 11 seconds of dial time. ViciDial then logged this lead as NA with length_in_sec = 0.
These commands assume the default full log format with timestamps like [Oct 5 09:02:29]. To count warnings per hour per peer:
awk -F"'" '/Received response: "200 OK" from .* without SDP/ {
split($0, a, ":"); print substr(a[1], 2), $2 }' /var/log/asterisk/full | sort | uniq -c
And to compare against total answers on that carrier over the same hour:
awk '/^\[Oct 5 09:/' /var/log/asterisk/full | \
grep -cE 'app_dial.c(:[0-9]+)?: SIP/carrier_a-[0-9a-f]+ answered'
awk '/^\[Oct 5 09:/' /var/log/asterisk/full | \
grep -cE 'Received response: "200 OK" from .carrier_a. without SDP'
If you use ISO timestamps (dateformat in logger.conf), adjust the hour extraction to match.
Does the Asterisk warning catch every SDP-less 200 OK?
No, and this matters. On one dialer, Homer saw 1,304 SDP-less 200 OK answers from Carrier A during my window, but Asterisk logged far fewer warnings. To see which calls the warnings belonged to, I matched them one by one. For every call answered before the cutoff where Asterisk sent BYE within 50 ms of its ACK, I took the dialed number from Homer. Then I matched each warning line to the dialed number in the Called SIP/... line with the same C- call ID. On that dialer, 595 calls appeared in both lists, with none left over on either side. On a second dialer it was 84 and 84.
So each warning is one call that Asterisk killed, and SDP-less answers that survived produce no warning. The reason is in chan_sip's handle_response_invite() (Asterisk 18 branch source; the build I tested was 18.26.4):
} else if (!reinvite) {
struct ast_sockaddr remote_address = {{0,}};
ast_rtp_instance_get_requested_target_address(p->rtp, &remote_address);
if (ast_sockaddr_isnull(&remote_address) || (!ast_strlen_zero(p->theirprovtag) && strcmp(p->theirtag, p->theirprovtag))) {
ast_log(LOG_WARNING, "Received response: \"200 OK\" from '%s' without SDP\n", p->relatedpeer->name);
ast_set_flag(&p->flags[0], SIP_PENDINGBYE);
ast_rtp_instance_activate(p->rtp);
}
}
When a 200 OK arrives without SDP, chan_sip checks whether it already has a remote media address (normally from an earlier 183 with SDP) and whether the To-tag matches the provisional response. If the address is missing, or the tag changed, it logs the warning and sets SIP_PENDINGBYE, which schedules a BYE once the ACK is out. If it already has a usable address, it silently carries on with the early-media details, and the call can survive.
The practical lesson: a grep 'without SDP' count is a count of warning events, which on my system matched the killed calls one for one. Use Homer for the real count of SDP-less answers.
How do I find ghost answers in Homer?
heplify-server writes SIP calls into hep_proto_1_call (registrations and other methods go to separate tables), with the full message in raw, addresses in protocol_header, and parsed cseq, method and callid in data_header. Using the parsed fields is more robust than regexing headers yourself, because heplify has already handled case and compact header forms.
The queries below count answered dialogs, not messages, and must run in one psql session because they build temp tables and use psql variables. They take each initial INVITE transaction (no To-tag) that your dialers sent inside a cohort window, keep the first 200 OK per dialog, and check the body for real SDP instead of trusting Content-Length. A dialog is matched exactly on Call-ID, our From-tag and the carrier's To-tag, the tags compared as parsed values, so a forked answer would show up as its own row. Messages are collected over a longer observation window than the INVITE cohort, so a call dialed at 09:59:58 and answered at 10:00:02 is still counted:
SET TIME ZONE 'UTC';
-- cohort = INVITEs sent in [c_start, c_end); observe their messages until o_end
\set c_start '''2026-10-05 00:00'''
\set c_end '''2026-10-05 07:45'''
\set o_end '''2026-10-05 08:00'''
-- 1. every message of every call your dialers sent to a carrier in the cohort window
CREATE TEMP TABLE m AS
SELECT create_date t,
protocol_header->>'srcIp' src, protocol_header->>'dstIp' dst,
data_header->>'method' meth,
data_header->>'callid' callid,
data_header->>'from_tag' ftag,
split_part(data_header->>'cseq', ' ', 1) cseqn,
split_part(data_header->>'cseq', ' ', 2) cseqm,
substring(split_part(raw, E'\r\n\r\n', 1)
from '(?in)^(?:to|t)[ \t]*:[^\r\n]*;[ \t]*tag=([^;>,[:space:]]+)') ttag,
substr(raw, length(split_part(raw, E'\r\n\r\n', 1)) + 5) body
FROM hep_proto_1_call
WHERE create_date >= :c_start AND create_date < :o_end
AND sid IN (SELECT sid FROM hep_proto_1_call
WHERE create_date >= :c_start AND create_date < :c_end
AND data_header->>'method' = 'INVITE'
AND protocol_header->>'dstIp' IN
('203.0.113.10', '203.0.113.20', '203.0.113.30')); -- carrier IPs
CREATE INDEX ON m (callid);
-- 2. initial INVITE transactions (no To-tag) sent inside the cohort window
CREATE TEMP TABLE inv AS
SELECT callid, cseqn, MIN(ftag) ftag, MIN(dst) carrier, MIN(src) dialer
FROM m
WHERE meth = 'INVITE' AND ttag IS NULL
AND dst IN ('203.0.113.10', '203.0.113.20', '203.0.113.30')
GROUP BY callid, cseqn
HAVING MIN(t) < :c_end;
-- 3. first 200 OK per answered dialog: Call-ID + our From-tag + their To-tag
-- (one row per fork if the carrier answers more than one early dialog)
CREATE TEMP TABLE ok AS
SELECT DISTINCT ON (i.callid, i.cseqn, m.ttag)
i.callid, i.cseqn, i.ftag, i.carrier, i.dialer, m.ttag, m.t t200,
(m.body ~ '(?n)^v=0' AND m.body ~ '(?n)^m=audio') has_sdp
FROM inv i
JOIN m ON m.callid = i.callid AND m.cseqn = i.cseqn AND m.cseqm = 'INVITE'
AND m.meth = '200' AND m.src = i.carrier AND m.ftag = i.ftag
ORDER BY i.callid, i.cseqn, m.ttag, m.t;
-- 4. per-carrier share of answered calls with no SDP, plus sanity counters
SELECT carrier, COUNT(*) answered_dialogs,
SUM((NOT has_sdp)::int) no_sdp,
ROUND(100.0 * SUM((NOT has_sdp)::int) / COUNT(*), 1) pct_no_sdp
FROM ok GROUP BY carrier ORDER BY carrier;
SELECT (SELECT COUNT(*) FROM inv) cohort_invites,
(SELECT COUNT(*) FROM (SELECT callid FROM ok GROUP BY callid
HAVING COUNT(*) > 1) x) calls_with_forks,
(SELECT COUNT(*) FROM inv i WHERE NOT EXISTS
(SELECT 1 FROM m WHERE m.callid = i.callid AND m.cseqn = i.cseqn
AND m.cseqm = 'INVITE' AND m.meth ~ '^[2-6][0-9][0-9]$')) pending;
The window values are an example; they are what I used. The cohort is calls whose INVITE went out between midnight and 07:45 UTC on the day I checked (dialing started around 07:00), observed until 08:00. It returned:
| Carrier | Answered calls | 200 OK without SDP |
|---|---|---|
| Carrier A | 2,366 | 1,475 (62.3%) |
| Carrier B | 2,755 | 0 (0.0%) |
| Carrier C | 3,682 | 0 (0.0%) |
All 1,475 of Carrier A's SDP-less 200s had a completely empty body. heplify-server's Dedup option was off on my install. I checked the window for duplicate captures and retransmitted INVITEs or 200s and found none. Of the 23,167 cohort INVITEs, none had more than one answered fork and none were still pending at the end of the observation window.
Run it as docker exec -i <postgres-container> psql -U homer -d homer_data if Homer runs in Docker. The SET TIME ZONE 'UTC' makes the window and any timestamps you print explicit. On my system Asterisk logs in local time (UTC+2), so the same call shows up two hours apart in the two sources.
Why does my Homer query flag clean carriers too?
Because of one innocent-looking filter. The obvious quick way to restrict raw messages to answers of INVITEs is:
AND raw LIKE '%CSeq:%INVITE%'
That also matches the 200 OK to a BYE or CANCEL, because those responses carry an Allow: header after the CSeq line:
CSeq: 104 BYE
Allow: INVITE, ACK, CANCEL, OPTIONS, BYE, REFER, SUBSCRIBE, NOTIFY, INFO, PUBLISH, MESSAGE
%CSeq:%INVITE% happily spans from CSeq: to the INVITE in Allow:. Counting raw 200 OK messages from my three carriers over the morning, the sloppy filter returned 12,453 messages. Only 8,754 of them were answers to an INVITE.
Those BYE and CANCEL responses normally have no body, so they look exactly like ghost answers. With the sloppy filter, Carrier C had 7,363 matching messages instead of 3,664, and showed up as 50.2% "without SDP" when the real figure was zero. Carriers A and B were unaffected only because none of their BYE or CANCEL responses matched the pattern. You cannot count on that.
If you do filter raw text, anchor the method to the CSeq line itself (raw ~ 'CSeq: [0-9]+ INVITE'). Better still, use the parsed data_header->>'cseq' field as in the query above.
How do I measure how fast the call dies?
For each SDP-less answer, find the ACK for that INVITE CSeq in the same dialog (same Call-ID, From-tag and To-tag), and the first BYE in that dialog in either direction. A BYE we send carries our tag in From; a BYE from the carrier has the tags swapped. Then classify, keeping incomplete and impossible cases visible instead of letting them fall into a bucket:
-- same psql session as the previous query (reuses temp tables m and ok)
CREATE TEMP TABLE g AS
SELECT o.*,
(SELECT MIN(t) FROM m
WHERE m.callid = o.callid AND m.meth = 'ACK' AND m.cseqn = o.cseqn
AND m.ftag = o.ftag AND m.ttag = o.ttag AND m.t >= o.t200) tack,
b.t tbye, b.src bye_src, b.dir bye_dir
FROM ok o
LEFT JOIN LATERAL (
SELECT t, src,
CASE WHEN m.ftag = o.ftag THEN 'ours' ELSE 'theirs' END dir
FROM m
WHERE m.callid = o.callid AND m.meth = 'BYE' AND m.t >= o.t200
AND ((m.ftag = o.ftag AND m.ttag = o.ttag) -- BYE sent by us
OR (m.ftag = o.ttag AND m.ttag = o.ftag)) -- BYE sent by the carrier
ORDER BY t LIMIT 1) b ON true
WHERE NOT o.has_sdp;
SELECT carrier,
CASE WHEN tack IS NULL OR tbye IS NULL THEN 'incomplete'
WHEN tbye < tack THEN 'negative'
WHEN tbye - tack < interval '50 ms' THEN 'instant'
ELSE 'survived' END AS cls,
COUNT(*) calls,
SUM((bye_dir = 'ours')::int) bye_from_us,
ROUND((percentile_cont(0.5) WITHIN GROUP
(ORDER BY EXTRACT(EPOCH FROM tbye - tack)))::numeric * 1000, 3) AS median_ms
FROM g GROUP BY 1, 2 ORDER BY 1, 2;
Of Carrier A's 1,475 SDP-less answers, 681 (46.2%) were instant: the BYE followed the ACK by a median of 0.038 ms, and by at most 18 ms. In all 681 the BYE was ours: our tag in From, sent from the same dialer IP the 200 OK was delivered to. From the 200 OK itself to the BYE, the median was 0.3 ms. The other 794 survived, with a median of 5.4 seconds between ACK and BYE. None were negative, and none were incomplete. If you run it while calls are still up, the last few minutes of the cohort will show incomplete rows, which is what that bucket is for.
These calls had 183s with SDP, so early media could have flowed before the answer, and I have not checked RTP for them. What a BYE less than a millisecond after the ACK does show is that my side tore the call down immediately on answer, leaving no room for a post-answer conversation. If you need to prove what audio was or wasn't present, you need the RTP or a recording. Signaling alone won't tell you.
What does the SIP ladder look like?
Here is the Homer ladder for the same call I showed from the Asterisk log (addresses replaced with documentation ranges, timestamps UTC):
07:02:16.827944 dialer -> carrier INVITE CSeq: 102 INVITE CL: 334
07:02:16.830601 carrier -> dialer 100 Giving it a try CSeq: 102 INVITE CL: 0
07:02:19.240142 carrier -> dialer 183 Session Progress CSeq: 102 INVITE CL: 254 c=198.51.100.7
07:02:19.433423 carrier -> dialer 183 Session Progress CSeq: 102 INVITE CL: 254 c=203.0.113.10
07:02:19.546739 carrier -> dialer 183 Session Progress CSeq: 102 INVITE CL: 254 c=203.0.113.10
07:02:19.661336 carrier -> dialer 183 Session Progress CSeq: 102 INVITE CL: 254 c=203.0.113.10
07:02:28.020534 carrier -> dialer 200 OK CSeq: 102 INVITE CL: 0
07:02:28.042133 dialer -> carrier ACK CSeq: 102 ACK
07:02:28.042431 dialer -> carrier BYE CSeq: 103 BYE
07:02:28.045899 carrier -> dialer 200 OK CSeq: 103 BYE
The ACK and the BYE are 0.3 ms apart, and the BYE went out 22 ms after the 200 OK arrived. If you prefer watching live on the dialer, sngrep -c host 203.0.113.10 shows only INVITE dialogs to and from that carrier, and you can open any call to get the same ladder.
To pull a ladder like this from Homer, grab the sid of the INVITE and list every message in that dialog:
SET TIME ZONE 'UTC';
SELECT to_char(create_date,'HH24:MI:SS.US') t,
protocol_header->>'srcIp' src, protocol_header->>'dstIp' dst,
split_part(raw, E'\r\n', 1) first_line,
data_header->>'cseq' cseq,
substring(raw from '(?i)Content-Length: *[0-9]+') cl,
substring(raw from 'c=IN IP4 [0-9.]+') c_line
FROM hep_proto_1_call
WHERE sid = '<sid>'
ORDER BY create_date;
Why do some ghost answers die and others survive?
Look at the SDP in those 183s. The first one says a=sendrecv from one media address. The second switches to a different address and a=sendonly. That matches the "placed on hold" log line from the same second.
I checked this across the whole sample. For each SDP-less call I took the last 183 with SDP that arrived before the 200 OK in the same dialog. I split its SDP into media sections at each m= line and read the direction attribute from the audio section, falling back to a session-level attribute (before the first m=) and then to the sendrecv default. Every one of those 183 SDPs had exactly one media section, audio. For 680 of the 681 instant kills the direction was sendonly, and the remaining one had no 183 with SDP at all. The two clean carriers sent about 7,000 183 messages with SDP each, and none of them had a sendonly audio section. The killed calls also averaged about 1.7 distinct media addresses across their 183s, versus about 0.7 for the survivors.
I'd call that a strong correlation rather than the full mechanism. Among the 794 survivors, 120 also ended ringing on a sendonly 183 and still lived, and 384 had no 183 with SDP at all. So sendonly alone does not decide it. What I can say for certain is that the calls that die are the ones where chan_sip hits the branch in the code above. For detection you do not need to know more than that.
Can I just fail over to another carrier on ghost answers?
Not with a normal DIALSTATUS failover. My dialplan already reroutes on carrier errors. This is a fragment: FBCNT is initialized earlier, DSTR holds the dial string, and the end and fb labels are defined further down:
exten => _88X.,n,Dial(${DSTR})
exten => _88X.,n,GotoIf($[${FBCNT}>=2]?end)
exten => _88X.,n,GotoIf($["${DIALSTATUS}"="CONGESTION" | "${DIALSTATUS}"="CHANUNAVAIL"]?fb)
A ghost answer returns DIALSTATUS=ANSWER, exactly like a real answer, so it bypasses this fallback completely. What happens to the lead afterwards depends on your campaign. On mine the call is dispositioned NA, so whether it gets dialed again comes down to your recycling rules, not your dialplan.
The response that works is at the route level:
- Reweight the split. With a
RAND()-based split, take the ghosting carrier out of the random pick or give it a small share. On ViciDial, make the change in the carrier's dialplan entry in the admin GUI so it survives config regeneration. - Keep a small share flowing if you need ongoing evidence. A few percent of traffic is enough to keep the Homer query and the
ans_but_nacolumn telling you whether the carrier fixed it. - Re-check daily. In my case it went from a small baseline to thousands of calls a day overnight. It can stop or change just as abruptly, and you want to know when.
What evidence should I send the carrier?
"Your ASR dropped" goes nowhere; their answer rate looks fine. Send them their own signaling:
- The Homer per-carrier table: share of answered calls whose
200 OKhad no SDP, theirs versus your other carriers on the same traffic, same time window. - Two or three full SIP ladders with Call-ID, timestamps and the raw
200 OK, showing the empty body. Include one where the 183 flips tosendonly. - The RFC reference. If the INVITE carried an SDP offer and the 183s were not sent reliably, quote RFC 3261 section 13.2.1. The 200 OK is missing the SDP answer this exchange required. That turns "your ASR looks odd" into a specific protocol fault their engineers can check.
- The daily table of Asterisk ANSWER versus dialer-answered calls, starting a few days before the change, so the step is obvious.
- An invoice question. Ask them to confirm whether calls answered without SDP and released within a second are billed. If they are, ask for a credit on those calls from the start date.
Keep the Call-IDs in the ticket. Those are what their NOC can actually look up.
How do I watch for this going forward?
Put two checks on a schedule:
- The per-call Homer query, run hourly with the previous complete hour as the INVITE cohort and an observation window that runs at least 15 minutes past it, so calls answered after the hour boundary still count. Report the
pendingcounter alongside it. Alert when a carrier has a meaningful number of answered calls (I'd want at least a hundred) and more than a few percent of them have no SDP. Exclude calls where a reliable 183 (withRSeq) already carried the SDP answer, because those 200s are legitimately empty. On my traffic there were none, and clean routes sat at 0.0%. - The
ans_but_nadaily query, run on the previous complete day. Alert on the ratio ofans_but_natoast_answerper carrier rather than the raw count, so volume changes and partial days don't trigger it.
If your Homer retention is short, run the Homer check while the data still exists. Mine keeps three days, so a weekend is enough to lose the SIP evidence for a Friday incident. Save the query output and a handful of ladders as soon as you notice something.
FAQ
Is a 200 OK without SDP always fraud? No. With 100rel and PRACK (RFC 3262) the answer can come in a reliable provisional response, and the 200 does not need to repeat it. I can't tell you whether my carrier did it on purpose or broke something. It becomes your problem when one route does it on most answers, your other routes do it on none, and the calls produce no conversation.
Does this affect billing? That depends on whether your carrier bills calls with a sub-second or zero answered duration. Check your invoice detail for the affected period before you assume either way.