A self-hosted AI phone agent can pass every demo and still fail on real outbound traffic. This is a case study of an Italian-language dental agent on Asterisk, fed by a ViciDial press-1 campaign. It covers what broke once real callers reached it: call identity, latency under load, answering machine detection, an overflowing context window, safety guards, and retrieval. The measured numbers come from the system's own traces, database rows, CDRs or test output. Where something is an observation and not a controlled result, I say so.
Configuration snapshot, September 2026: Asterisk 20.6 on the agent host; ViciDial on Asterisk 18.26.4-vici (ViciDial's patched build) on the dialer; AudioSocket; Silero VAD 6.2; faster-whisper 1.2.1 (large-v3); Ollama 0.30.8 with llama3.1:8b (num_ctx 4096); Qdrant 1.18; one 20 GB workstation GPU. Italian TTS was a hosted streaming service, with local Kokoro/Piper code as a fallback. Earlier experiments ran on earlier spec versions of the same stack.
What was the setup?
ViciDial press-1 campaign -> SIP -> Asterisk (agent host)
-> AGI (stage a session) -> AudioSocket (8 kHz 16-bit PCM over TCP)
-> Python gateway: Silero VAD -> faster-whisper -> LLM (Ollama) -> TTS
-> audio back over the same socket
ViciDial dials the leads. The callee hears a short recorded message and presses 1, and ViciDial sends the call to the agent host. The agent speaks Italian and has one job: find out whether the caller has a dental need, then get them to send two photos of their teeth over WhatsApp so a human can prepare a free quote.
Basic demos sounded good. Targeted tests and real-call reviews did not. The checklist at the end sums up what we learned.
Why did the audit call it "a pilot, not production"?
About two weeks after real press-1 calls started, we ran a read-only production-readiness audit. Its first line said: "operational pilot, not ready for an unattended production service." The GPU, models and health checks were fine. What failed was the integration around them.
The worst finding was a P0 in the first step every call runs. The dialplan called a shell AGI script before AudioSocket, and the script read its AGI environment like this:
while read -r line; do
[ -z "$line" ] && break
export "$line"
done
Asterisk sends AGI variables as agi_request: something, with a colon and a space. That is not a shell assignment, and you can reproduce the failure in one line (123.4 is a made-up uniqueid):
/bin/sh -c 'while read -r line; do [ -z "$line" ] && break; export "$line"; done; echo REACHED' <<'EOF'
agi_request: prepare.sh
agi_uniqueid: 123.4
EOF
# /bin/sh: 1: export: agi_request: prepare.sh: bad variable name
# exit=2
The script died before it staged anything. In the gateway's session_prepared events, agent, direction and provider reference were all null. The dialplan also passed a hard-coded UUID per language to AudioSocket(), so every Italian call reached the gateway with the same 3333... UUID. The gateway used that UUID to pick the language line, not to identify the call. Linking a conversation back to its ViciDial lead took a timestamp join.
A demo call doesn't show this, because the agent still answers and talks.
The fix was a small Python AGI that treats AGI input as a protocol, never as shell code:
import re
def read_environment(stream):
"""Read AGI 'agi_key: value' lines until the blank line."""
result = {}
for line in stream:
line = line.rstrip('\r\n')
if not line:
break
key, sep, value = line.partition(':')
if sep and re.fullmatch(r'agi_[a-z0-9_]+', key):
result[key] = value.lstrip()
return result
After reading the environment, the script sets an 8-second signal.alarm. Inside that deadline it reads its config, checks the gateway readiness URL (when one is configured), logs in, generates uuid.uuid4(), stages the session and returns the IDs with SET VARIABLE. The deadline does not cover the initial environment read or the final SET VARIABLE FLOW_STAGE_OK 0 on the error path. If you copy this design, put those under a timeout too.
The dialplan checks the result. This excerpt is from the deployed context (the channel is answered earlier, at Answer()):
same => n,Set(FLOW_STAGE_OK=0)
same => n,AGI(prepare.sh,${EXTEN})
same => n,GotoIf($["${FLOW_STAGE_OK}" != "1"]?flow_unavailable)
same => n,ExecIf($["${CALLSESSIONID}" != ""]?Set(CDR(userfield)=${CALLSESSIONID}))
same => n,ExecIf($["${FLOW_VICI_CID}" != ""]?Set(CDR(userfield)=${CDR(userfield)}|vici:${FLOW_VICI_CID}))
same => n(gw),AudioSocket(${FLOW_CALL_UUID},${FLOW_AUDIO_GW})
...
same => n(flow_unavailable),Playback(custom/agent-unavailable)
same => n,Hangup(41)
As deployed, the failure path plays an "unavailable" message and hangs up. That fails closed, but no human or callback destination exists yet.
On the dialer side, ViciDial adds its caller code as a SIP header (SIPAddHeader(X-vici-cid: ${CALLERID(name)})), and the agent host reads it with PJSIP_HEADER(read,X-vici-cid). For one real call after the fix, the IDs line up like this (values shortened):
| Where | Field | Value |
|---|---|---|
ViciDial vicidial_log_extended |
caller_code / uniqueid / lead_id |
V917…174 / …793.36 / …174 |
| Agent-host CDR | userfield |
<sessionId>|vici:V917…174 |
| Agent-host CDR | uniqueid |
…818.512 |
| Agent-host CDR | lastdata (AudioSocket UUID) |
d26f…6beb |
Gateway session_prepared |
callSessionId, asteriskUuid, externalProviderRef |
<sessionId>, d26f…6beb, …818.512 |
The two uniqueids differ. The caller code in the header and the CDR links them.
Two more blockers:
- No durable outcome. The application's lead table held 0 records across 229 agent conversations (test calls included). The code synced leads only when inline state parsing succeeded. That path was switched off, so the deferred extractor saved state but never created a lead.
- The release gate published first and tested second. The deploy script switched the live spec before it ran the text gate. The gate also skipped any sentence that contained a negation. It flagged "Le ho inviato il messaggio" ("I've sent you the message", a false claim of action) but missed "Non si preoccupi, le ho inviato il messaggio" ("Don't worry, I've sent you the message").
How did latency hold up with several calls at once?
Badly. Single post-deploy canary calls got their first reply in 1.7 to 2.2 s. We then ran six synthetic AudioSocket conversations at once, with Ollama set to three parallel slots. The time from the caller finishing a sentence to the first agent audio was:
4893, 5011, 5020, 8257, 9051, 9320 ms -> median 6.64 s, worst 9.32 s
Cutting the reply cap from 140 to 90 tokens gave a median of 5.27 s and a worst case of 8.60 s. That is one small run per setting, so treat it as a direction, not a percentile. The gateway advertised 10 calls and the upstream route allowed 6. Nobody had measured whether either limit could hold acceptable latency.
The cache matters too: an uncached five-session run hit 19.3 s, while a warm ten-session run with identical caller text topped out at 4.85 s.
Lesson: benchmark several concurrency levels with varied caller audio and a cold cache. Set admission control from the measured p95. Send overflow to a real destination, not to Congestion.
Should you keep ViciDial AMD in front of an AI agent?
For us, no. The campaign used a ViciDial survey extension that runs Asterisk's AMD() before routing:
exten => 8373,n,Playback(sip-silence)
exten => 8373,n,AMD(2000,2000,1000,5000,120,50,4,256)
exten => 8373,n,AGI(VD_amd.agi,${EXTEN})
VD_amd.agi copies the AMDSTATS channel variable into run_time, in milliseconds. Stock Asterisk AMD doesn't set AMDSTATS. ViciDial's patched -vici build does, and on our dialer every row in the window had it populated. The script writes to vicidial_amd_log only when $AMD_LOG is 2 or 3. Ours was 3. Check both before you rely on this query:
SELECT a.AMDSTATUS, a.AMDRESPONSE, vl.status,
COUNT(*) n, COUNT(DISTINCT a.uniqueid) calls,
ROUND(AVG(CAST(a.run_time AS UNSIGNED))/1000,2) avg_s,
ROUND(MAX(CAST(a.run_time AS UNSIGNED))/1000,2) max_s -- run_time is VARCHAR
FROM (SELECT uniqueid, AMDSTATUS, AMDRESPONSE, run_time FROM vicidial_amd_log
WHERE call_date BETWEEN '2026-09-14 15:00:00' AND '2026-09-14 16:05:50'
UNION ALL
SELECT uniqueid, AMDSTATUS, AMDRESPONSE, run_time FROM vicidial_amd_log_archive
WHERE call_date BETWEEN '2026-09-14 15:00:00' AND '2026-09-14 16:05:50') a
JOIN vicidial_log vl ON vl.uniqueid = a.uniqueid
WHERE vl.campaign_id = 'YOUR_AI_CAMPAIGN'
GROUP BY 1,2,3;
The inner join drops AMD rows with no matching vicidial_log row (25 in this window, most from other campaigns). It produced 168 rows for 168 distinct calls in the hour before the change:
| AMD result | ViciDial status | Calls | Avg AMD run time |
|---|---|---|---|
| HUMAN | sent to call menu (SVYCLM) | 43 | 3.17 s |
| NOTSURE / TOOLONG | sent to call menu (SVYCLM) | 30 | 3.31 s |
| NOTSURE / NOAUDIODATA | sent to call menu (SVYCLM) | 15 | 4.80 s |
| NOTSURE / TOOLONG | dropped (PDROP) | 1 | 2.16 s |
| MACHINE (initial silence, long greeting, max words) | hung up (AA) | 79 | 3.25 s |
The 88 calls sent on to the call menu spent an average of 3.50 s in AMD (maximum 5.02 s) first. 79 of 168 (47%) were hung up as machines. A person who pauses or says more than "Pronto?" can trip any of the three machine causes. We didn't listen to the AA calls, so the false-positive share is unknown.
We switched the campaign to the same survey extension without AMD (8366) and compared the gateway's own traces for each window. The gateway count (84) uses a slightly different window from the SQL above:
| Window | Calls reaching agent | Caller said any words | Ended < 9 s, no speech | Voicemails caught by gateway |
|---|---|---|---|---|
| With AMD (15:10 to 16:05) | 84 | 12% | 74% | 1 |
| Without AMD (16:05 to 16:25) | 54 | 30% | 43% | 9 |
This is a before/after observation, not a randomized test. The windows differ in length and time of day, and the lead mix may differ. In this sample, more callers spoke after AMD came out. The cost is that answering machines now reach the agent. The gateway caught 9 of them in that window with an Italian transcript regex ("dopo il segnale acustico", "non è raggiungibile", mailbox menus). We don't know how many it missed. Any it misses run until the idle timeout and hold a call slot.
We also tested one off-the-shelf audio AMD checkpoint, a whisper-tiny telephony classifier. We ran it offline on 15 calls from that day. The input was the first 3 s of the caller channel after the first audible sample, resampled to 16 kHz, and we took the top class. The reference labels came from the gateway traces, not from listening: a voicemail detection meant "machine", two or more substantive transcripts meant "human".
labelled n= 15
HUMAN recall: 3/4 = 75%
MACHINE recall: 0/11 = 0%
So the checkpoint agreed with 0 of 11 trace-labelled machines, calling all of them human. We didn't validate the labels by ear or establish why it disagreed, so treat this as a reason to test any AMD model on your own labelled calls before trusting it.
What happens when the prompt outgrows the context window?
The always-on responseConstraints block in the agent spec kept growing as we added rules to discipline the 8B model. It went from 6,534 to 14,268 characters between two spec versions, with num_ctx at 4096. After that release, calls felt slower and replies drifted.
To test the prompt in isolation, we ran an A/B probe against the same model. Knowledge, history and question were identical. Only the constraints block changed. Each variant got one warm-up request, then four sequential requests, and we report the medians:
[v34] prompt_tok=4095 prefill_ms=1609 out_tok~140 decode_ms=3436 TOTAL_ms=5335
[v35] prompt_tok=2243 prefill_ms=22 out_tok~71 decode_ms=1474 TOTAL_ms=1660
A prompt count of 4095 against a 4096 context strongly suggests the prompt was being cut. The long variant's median output also hit the 140-token cap. In this probe, compacting the block to 5,834 characters cut the median total from 5.3 s to 1.7 s.
The problem came back. Under a newer spec, long calls pushed prompts past 4096 again. Ollama logs a warning each time it truncates. To count them for one UTC day (ollama is your container name):
docker logs --since 2026-09-17T00:00:00Z --until 2026-09-18T00:00:00Z ollama \
> /tmp/ollama-day.log 2>&1 || echo "log fetch failed"
grep -c "truncating input prompt" /tmp/ollama-day.log
# 78
# msg="truncating input prompt" limit=4096 prompt=4173 keep=4 new=4095
On that same day, callers heard replies such as "Sembra che il cliente abbia risposto al telefono... posso continuare seguendo le istruzioni fornite" ("It seems the customer has answered the phone... I can continue following the instructions provided"). Another offered to "fare una simulazione della conversazione" ("run a simulation of the conversation"). We didn't match each of these replies to a specific truncated request. The pattern fits a model that has lost part of its instructions, but treat that link as a hypothesis.
Lesson: budget tokens in code. Keep the fully rendered prompt plus a reserve for the reply under the context window. Alert on the daily truncation count, because the runtime doesn't fail. It just quietly gets worse.
Are prompt rules enough to keep an 8B model safe?
No. The clearest example was a clinical rule: never suggest full-arch treatment ("All-on-4", "riabilitazione dell'arcata") to someone who is missing only one or two teeth. The rule was in the prompt, and one run of the adversarial eval still produced this:
Caller: "Mi mancano due denti davanti." ("I'm missing two front teeth.") Agent: "...La soluzione per te potrebbe essere l'All-on-4 o All-on-6..." ("...The solution for you could be All-on-4 or All-on-6...")
In a scripted golden conversation, the caller said "A me mancano due denti" ("I'm missing two teeth") and the model answered "visto che mancano tutti i denti" ("since all your teeth are missing").
The fix was deterministic code. A per-turn eligibility scope classifies the caller's own words as "few" or "many" missing teeth. When the answer is "few", full-arch FAQs are dropped from retrieval and a speech guard scrubs full-arch phrases from the output. Here is an excerpt (not a complete program) of the detection patterns and test cases, from the Python mirror of the production C# code:
# excerpt: FEWMISS = "few teeth missing"; SOLO and CNT are small helper patterns
FEWMISS = re.compile(
r"manca(?:no)?\s+" + SOLO + CNT + r"?\s*(?:dent|incisiv|molar|canin|premolar)"
r"|\bmi\s+manca\b|\bme\s+ne\s+manca(?:no)?\b"
# ... more alternatives omitted ...
, re.IGNORECASE)
DETECTION_CASES = [
(["mi mancano due denti"], True),
(["non sono senza denti, me ne mancano due"], True),
(["ho solo due denti"], False),
(["mi sono rimasti due denti"], False),
(["ho perso tutti i denti", "in realtà me ne mancano due"], True),
# ... 20 more ...
]
The full regression corpus has 25 detection cases and 41 output-scrub cases, and all 66 pass. It tests a Python mirror of the production C# patterns, not every streamed chunk boundary, cache hit and fallback in the deployed build. That needs integration tests against the deployed version. A reviewer also found a phrasing the corpus misses: the mirror classifies "non ho solo due denti, me ne mancano due" ("I don't have only two teeth, I'm missing two") as near-total loss. Regex guards are only as good as their adversarial cases. After the guard went in, the same eval answered "impianti singoli o un piccolo ponte" ("single implants or a small bridge").
The same went for formal address ("Lei"), hallucinated phone numbers and meta-replies: the prompt rule came first and failed in testing, then code guards replaced it.
Can a safety guard break the business goal?
Yes, and nothing alerts you when it happens. The goal was to give interested callers the WhatsApp number. A guard stopped the model from giving the number before the caller agreed. At the time, the affirmative branch was a short full-match allowlist ("sì", "va bene", "d'accordo", "certo", "ok"), plus exceptions for explicit requests such as "mi dia il numero". Real callers answered the invite with things like "Va bene, ci penso e poi ve la mando. Di dove siete?" ("OK, I'll think about it and send it. Where are you based?"), which failed the match. The guard then deleted the model's number sentence.
Our proxy was the number of calls in which an agent_reply event contained the spelled-out country-code prefix of the number:
2026-09-08: 2 of 86 2026-09-16: 0 of 1049 (guard stripped it in 17 calls)
2026-09-09: 1 of 380 2026-09-17: 1 of 1001 (stripped in 16)
2026-09-11: 1 of 386 2026-09-21: 1 of 880
2026-09-15: 0 of 778 2026-09-22: 2 of 1300
This is a text proxy, not proof of what the caller heard. In the one Sep 17 call where the number appeared, playback events show it came out garbled and the caller interrupted. For more than two weeks, the main conversion step barely appeared even in the logged text, while call volumes looked healthy.
Lesson: every guard needs a counter. Measure the conversion step from playback events, then spot-check against recordings.
What do real calls reveal that test calls don't?
We reviewed one full day of traces (1,001 calls, 385 with any transcribed caller speech) and found failure modes that no synthetic test had produced.
Stuck VAD on stalled audio. 30 calls ran past 60 s with no transcribed caller speech. Every one had an unmatched vad_start and no idle reprompt, and ended when the AudioSocket peer closed, at about 185 s or about 315 s. Together they took roughly 99 gateway-session minutes. We didn't check CDRs for how much of that was billed. The cause was in the gateway code: the VAD force-end only ran inside the per-frame audio handler. When AudioSocket frames stopped arriving mid-utterance, in_speech stayed true, and that blocked the idle and max-duration timers. Here is a filter for candidate calls, written for our trace schema (one JSONL file per call, rel = seconds since the call started):
import glob, json, sys
TRACE_DIR, DAY = sys.argv[1], sys.argv[2] # e.g. ./traces 20260917
for path in sorted(glob.glob(f"{TRACE_DIR}/{DAY}-*.jsonl")):
heard = vad_open = reprompted = False
duration = end_reason = None
with open(path) as fh:
for line in fh:
try:
e = json.loads(line)
except ValueError:
continue
t = e.get("type")
if t == "stt" and (e.get("text") or "").strip():
heard = True
elif t == "vad_start":
vad_open = True
elif t == "vad_end":
vad_open = False
elif t == "idle_reprompt":
reprompted = True
elif t == "call_end":
duration, end_reason = e.get("rel"), e.get("hangupReason")
if duration and duration > 60 and not heard and vad_open and not reprompted:
print(path, round(duration), end_reason)
The fix belongs in a watchdog outside the audio callback. If in_speech is set and no frame has arrived for a few seconds, close the utterance so the normal idle reprompt and goodbye logic can run. Keep that separate from any decision to hang up, and keep an absolute call deadline as well. Don't hang up just because media went quiet: a legitimate call can be silent, and some endpoints use silence suppression.
IVR loops. The backend's transfer intent matched "prema 1" ("press 1"). So did a restaurant's phone menu, and the agent answered it with the same canned line 24 times. Cap identical consecutive replies.
Losing people at the greeting. In the first two days of real calls (37 calls), 15 callers hung up during the greeting without saying a word. All 5 callers whose first sentence named a concrete need (a quote, a denture, implants, yellow teeth, a full set of teeth) were tagged as leads in our review. That is a small sample, and stating a need naturally predicts a lead tag. We didn't test what caused the early hangups. Our working hypothesis is the opening, and it needs a controlled test.
"Grazie a tutti" was also the first transcribed "utterance" in three of those calls, a known Whisper output on silence. Discard such phrases only when VAD or no-speech evidence agrees, since a person can really say them.
Does RAG help a small model on the phone?
Less than we assumed. We replayed 292 real caller utterances offline: 241 from this agent and 51 from the hosted platform described below. The replay used the live embedder (nomic-embed-text), a read-only search of the live Qdrant collection, and a copy of the production ranking logic. For the 136 utterances that were real questions or objections, we hand-labelled the correct FAQ:
| Measure | Result |
|---|---|
| Top-1 correct, live query (prev. turn + utterance) | 54 / 136 (40%) |
| Top-1 correct, raw utterance only | 68 / 136 (50%) |
| None of the top 3 correct | 57 / 136 (42%) |
| Top-1 correct, by source (live query) | this agent 47/107, hosted 7/29 |
| Lowest top-1 / third-ranked cosine, all 292 | 0.66 / 0.62 |
The live config had contextMinScore: 0.52. Even the third-ranked candidate never scored below 0.62, so that cut-off removed nothing in this replay. Prepending the previous caller turn ("contextual" query) cost about 10 points of top-1 accuracy compared with the raw utterance. Correct and wrong top-1 hits had overlapping scores, so a higher cut-off alone discards correct answers too.
Lesson: replay labelled, real utterances against your index before you trust it. Pick a threshold by measuring accuracy, coverage and false rejections on those labels, and allow retrieval to return nothing.
Was a hosted no-code voice platform the answer?
We ran a hosted no-code voice-agent platform in parallel on cold outbound calls, with a 52-node conversation flow. From its own call export over about ten days:
| Metric | Value |
|---|---|
| Real outbound dials | 12,160 |
| Connected | 4,285 |
| Calls where a human spoke (platform label) | 2,363 |
| Of those, left during the intro (platform label) | 2,156 (91%) |
| Tagged "warm" | 5 |
| Tagged "agreed to send photos" | 0 |
| Total platform cost | about $266 |
The labels come from the platform's own LLM post-call analysis, and we had no separate downstream outcome data. Like our agent, it lost most callers early. Testing a different opening is the next experiment. Nothing here shows whether the opening matters more than the model.
What does a production readiness checklist look like?
Each item comes from a failure above.
Call identity and integration
- AGI input parsed as a
key: valueprotocol, neverexported orevaled. Tested offline with real AGI headers. - A unique AudioSocket UUID per call. A documented mapping from dialer uniqueid to agent-host uniqueid to session to recording, checked on real calls.
- A timeout around the whole staging step, including input and error handling. On failure, the dialplan goes to an approved human or callback destination.
- A consenting, qualified test call creates exactly one lead/task row in the application database. Separately, test that a refusal and an opt-out are stored correctly.
Capacity and latency
- Concurrency measured at several levels with a cold cache and varied caller audio, reporting p50 and p95 from the caller's end of speech to the first agent audio.
- Admission limit set from that measurement. Overflow goes to a human queue or a callback, not
Congestion.
Dialer front end
- AMD run time measured from
vicidial_amd_log.run_time, after confirming thatAMDSTATSis populated and DB logging is on. - If AMD stays on, false positives sampled by listening to AA recordings.
- Voicemail detection validated on labelled calls in your own language.
LLM and prompt
- Worst-case rendered prompt (long call, injected FAQs, full history) plus a reply reserve fits under
num_ctx, enforced in code. - Daily count of "truncating input prompt" warnings, with an alert.
- Safety rules that matter have deterministic guards, a regression corpus, and integration tests on the deployed build, including streamed output.
Guards and funnel
- Every guard logs when it fires, and someone reviews the counts daily.
- The conversion step measured from playback events and spot-checked against recordings.
- When a guard empties a reply, a spoken fallback line plays instead of silence.
Real-call hygiene
- A watchdog outside the audio callback closes stuck utterances. Hangup policy is separate and tested with silent callers.
- Identical consecutive replies capped, to break IVR loops.
- Suspected Whisper silence hallucinations discarded only when VAD/no-speech evidence agrees.
Retrieval
- Hit@1 measured on real, labelled caller utterances.
- Threshold chosen on labelled accuracy and false rejections. Retrieval can return zero results.
Release
- The candidate is gated before it takes traffic, with a manifest (spec, KB, code, config, models) and a rehearsed rollback.
What would I do differently?
I'd build the measurement before the agent: per-call identity, a playback-based funnel, guard counters and a concurrency benchmark. Demos and service dashboards missed nearly every problem here. Targeted tests and real-call reviews exposed them.