← All Tutorials

AI Voice Agent Pilot vs Production: Lessons From Real Outbound Calls

AI & Voice Agents Advanced 21 min read #102 Published

A self-hosted AI phone agent can pass every demo and still fail on real outbound traffic. This is a case study of an Italian-language dental agent on Asterisk, fed by a ViciDial press-1 campaign. It covers what broke once real callers reached it: call identity, latency under load, answering machine detection, an overflowing context window, safety guards, and retrieval. The measured numbers come from the system's own traces, database rows, CDRs or test output. Where something is an observation and not a controlled result, I say so.

Configuration snapshot, September 2026: Asterisk 20.6 on the agent host; ViciDial on Asterisk 18.26.4-vici (ViciDial's patched build) on the dialer; AudioSocket; Silero VAD 6.2; faster-whisper 1.2.1 (large-v3); Ollama 0.30.8 with llama3.1:8b (num_ctx 4096); Qdrant 1.18; one 20 GB workstation GPU. Italian TTS was a hosted streaming service, with local Kokoro/Piper code as a fallback. Earlier experiments ran on earlier spec versions of the same stack.

What was the setup?

ViciDial press-1 campaign -> SIP -> Asterisk (agent host)
  -> AGI (stage a session) -> AudioSocket (8 kHz 16-bit PCM over TCP)
  -> Python gateway: Silero VAD -> faster-whisper -> LLM (Ollama) -> TTS
  -> audio back over the same socket

ViciDial dials the leads. The callee hears a short recorded message and presses 1, and ViciDial sends the call to the agent host. The agent speaks Italian and has one job: find out whether the caller has a dental need, then get them to send two photos of their teeth over WhatsApp so a human can prepare a free quote.

Basic demos sounded good. Targeted tests and real-call reviews did not. The checklist at the end sums up what we learned.

Why did the audit call it "a pilot, not production"?

About two weeks after real press-1 calls started, we ran a read-only production-readiness audit. Its first line said: "operational pilot, not ready for an unattended production service." The GPU, models and health checks were fine. What failed was the integration around them.

The worst finding was a P0 in the first step every call runs. The dialplan called a shell AGI script before AudioSocket, and the script read its AGI environment like this:

while read -r line; do
  [ -z "$line" ] && break
  export "$line"
done

Asterisk sends AGI variables as agi_request: something, with a colon and a space. That is not a shell assignment, and you can reproduce the failure in one line (123.4 is a made-up uniqueid):

/bin/sh -c 'while read -r line; do [ -z "$line" ] && break; export "$line"; done; echo REACHED' <<'EOF'
agi_request: prepare.sh
agi_uniqueid: 123.4

EOF
# /bin/sh: 1: export: agi_request: prepare.sh: bad variable name
# exit=2

The script died before it staged anything. In the gateway's session_prepared events, agent, direction and provider reference were all null. The dialplan also passed a hard-coded UUID per language to AudioSocket(), so every Italian call reached the gateway with the same 3333... UUID. The gateway used that UUID to pick the language line, not to identify the call. Linking a conversation back to its ViciDial lead took a timestamp join.

A demo call doesn't show this, because the agent still answers and talks.

The fix was a small Python AGI that treats AGI input as a protocol, never as shell code:

import re

def read_environment(stream):
    """Read AGI 'agi_key: value' lines until the blank line."""
    result = {}
    for line in stream:
        line = line.rstrip('\r\n')
        if not line:
            break
        key, sep, value = line.partition(':')
        if sep and re.fullmatch(r'agi_[a-z0-9_]+', key):
            result[key] = value.lstrip()
    return result

After reading the environment, the script sets an 8-second signal.alarm. Inside that deadline it reads its config, checks the gateway readiness URL (when one is configured), logs in, generates uuid.uuid4(), stages the session and returns the IDs with SET VARIABLE. The deadline does not cover the initial environment read or the final SET VARIABLE FLOW_STAGE_OK 0 on the error path. If you copy this design, put those under a timeout too.

The dialplan checks the result. This excerpt is from the deployed context (the channel is answered earlier, at Answer()):

 same => n,Set(FLOW_STAGE_OK=0)
 same => n,AGI(prepare.sh,${EXTEN})
 same => n,GotoIf($["${FLOW_STAGE_OK}" != "1"]?flow_unavailable)
 same => n,ExecIf($["${CALLSESSIONID}" != ""]?Set(CDR(userfield)=${CALLSESSIONID}))
 same => n,ExecIf($["${FLOW_VICI_CID}" != ""]?Set(CDR(userfield)=${CDR(userfield)}|vici:${FLOW_VICI_CID}))
 same => n(gw),AudioSocket(${FLOW_CALL_UUID},${FLOW_AUDIO_GW})
 ...
 same => n(flow_unavailable),Playback(custom/agent-unavailable)
 same => n,Hangup(41)

As deployed, the failure path plays an "unavailable" message and hangs up. That fails closed, but no human or callback destination exists yet.

On the dialer side, ViciDial adds its caller code as a SIP header (SIPAddHeader(X-vici-cid: ${CALLERID(name)})), and the agent host reads it with PJSIP_HEADER(read,X-vici-cid). For one real call after the fix, the IDs line up like this (values shortened):

Where Field Value
ViciDial vicidial_log_extended caller_code / uniqueid / lead_id V917…174 / …793.36 / …174
Agent-host CDR userfield <sessionId>|vici:V917…174
Agent-host CDR uniqueid …818.512
Agent-host CDR lastdata (AudioSocket UUID) d26f…6beb
Gateway session_prepared callSessionId, asteriskUuid, externalProviderRef <sessionId>, d26f…6beb, …818.512

The two uniqueids differ. The caller code in the header and the CDR links them.

Two more blockers:

How did latency hold up with several calls at once?

Badly. Single post-deploy canary calls got their first reply in 1.7 to 2.2 s. We then ran six synthetic AudioSocket conversations at once, with Ollama set to three parallel slots. The time from the caller finishing a sentence to the first agent audio was:

4893, 5011, 5020, 8257, 9051, 9320 ms   -> median 6.64 s, worst 9.32 s

Cutting the reply cap from 140 to 90 tokens gave a median of 5.27 s and a worst case of 8.60 s. That is one small run per setting, so treat it as a direction, not a percentile. The gateway advertised 10 calls and the upstream route allowed 6. Nobody had measured whether either limit could hold acceptable latency.

The cache matters too: an uncached five-session run hit 19.3 s, while a warm ten-session run with identical caller text topped out at 4.85 s.

Lesson: benchmark several concurrency levels with varied caller audio and a cold cache. Set admission control from the measured p95. Send overflow to a real destination, not to Congestion.

Should you keep ViciDial AMD in front of an AI agent?

For us, no. The campaign used a ViciDial survey extension that runs Asterisk's AMD() before routing:

exten => 8373,n,Playback(sip-silence)
exten => 8373,n,AMD(2000,2000,1000,5000,120,50,4,256)
exten => 8373,n,AGI(VD_amd.agi,${EXTEN})

VD_amd.agi copies the AMDSTATS channel variable into run_time, in milliseconds. Stock Asterisk AMD doesn't set AMDSTATS. ViciDial's patched -vici build does, and on our dialer every row in the window had it populated. The script writes to vicidial_amd_log only when $AMD_LOG is 2 or 3. Ours was 3. Check both before you rely on this query:

SELECT a.AMDSTATUS, a.AMDRESPONSE, vl.status,
       COUNT(*) n, COUNT(DISTINCT a.uniqueid) calls,
       ROUND(AVG(CAST(a.run_time AS UNSIGNED))/1000,2) avg_s,
       ROUND(MAX(CAST(a.run_time AS UNSIGNED))/1000,2) max_s   -- run_time is VARCHAR
FROM (SELECT uniqueid, AMDSTATUS, AMDRESPONSE, run_time FROM vicidial_amd_log
        WHERE call_date BETWEEN '2026-09-14 15:00:00' AND '2026-09-14 16:05:50'
      UNION ALL
      SELECT uniqueid, AMDSTATUS, AMDRESPONSE, run_time FROM vicidial_amd_log_archive
        WHERE call_date BETWEEN '2026-09-14 15:00:00' AND '2026-09-14 16:05:50') a
JOIN vicidial_log vl ON vl.uniqueid = a.uniqueid
WHERE vl.campaign_id = 'YOUR_AI_CAMPAIGN'
GROUP BY 1,2,3;

The inner join drops AMD rows with no matching vicidial_log row (25 in this window, most from other campaigns). It produced 168 rows for 168 distinct calls in the hour before the change:

AMD result ViciDial status Calls Avg AMD run time
HUMAN sent to call menu (SVYCLM) 43 3.17 s
NOTSURE / TOOLONG sent to call menu (SVYCLM) 30 3.31 s
NOTSURE / NOAUDIODATA sent to call menu (SVYCLM) 15 4.80 s
NOTSURE / TOOLONG dropped (PDROP) 1 2.16 s
MACHINE (initial silence, long greeting, max words) hung up (AA) 79 3.25 s

The 88 calls sent on to the call menu spent an average of 3.50 s in AMD (maximum 5.02 s) first. 79 of 168 (47%) were hung up as machines. A person who pauses or says more than "Pronto?" can trip any of the three machine causes. We didn't listen to the AA calls, so the false-positive share is unknown.

We switched the campaign to the same survey extension without AMD (8366) and compared the gateway's own traces for each window. The gateway count (84) uses a slightly different window from the SQL above:

Window Calls reaching agent Caller said any words Ended < 9 s, no speech Voicemails caught by gateway
With AMD (15:10 to 16:05) 84 12% 74% 1
Without AMD (16:05 to 16:25) 54 30% 43% 9

This is a before/after observation, not a randomized test. The windows differ in length and time of day, and the lead mix may differ. In this sample, more callers spoke after AMD came out. The cost is that answering machines now reach the agent. The gateway caught 9 of them in that window with an Italian transcript regex ("dopo il segnale acustico", "non è raggiungibile", mailbox menus). We don't know how many it missed. Any it misses run until the idle timeout and hold a call slot.

We also tested one off-the-shelf audio AMD checkpoint, a whisper-tiny telephony classifier. We ran it offline on 15 calls from that day. The input was the first 3 s of the caller channel after the first audible sample, resampled to 16 kHz, and we took the top class. The reference labels came from the gateway traces, not from listening: a voicemail detection meant "machine", two or more substantive transcripts meant "human".

labelled n= 15
HUMAN recall: 3/4 = 75%
MACHINE recall: 0/11 = 0%

So the checkpoint agreed with 0 of 11 trace-labelled machines, calling all of them human. We didn't validate the labels by ear or establish why it disagreed, so treat this as a reason to test any AMD model on your own labelled calls before trusting it.

What happens when the prompt outgrows the context window?

The always-on responseConstraints block in the agent spec kept growing as we added rules to discipline the 8B model. It went from 6,534 to 14,268 characters between two spec versions, with num_ctx at 4096. After that release, calls felt slower and replies drifted.

To test the prompt in isolation, we ran an A/B probe against the same model. Knowledge, history and question were identical. Only the constraints block changed. Each variant got one warm-up request, then four sequential requests, and we report the medians:

[v34] prompt_tok=4095  prefill_ms=1609  out_tok~140  decode_ms=3436  TOTAL_ms=5335
[v35] prompt_tok=2243  prefill_ms=22    out_tok~71   decode_ms=1474  TOTAL_ms=1660

A prompt count of 4095 against a 4096 context strongly suggests the prompt was being cut. The long variant's median output also hit the 140-token cap. In this probe, compacting the block to 5,834 characters cut the median total from 5.3 s to 1.7 s.

The problem came back. Under a newer spec, long calls pushed prompts past 4096 again. Ollama logs a warning each time it truncates. To count them for one UTC day (ollama is your container name):

docker logs --since 2026-09-17T00:00:00Z --until 2026-09-18T00:00:00Z ollama \
  > /tmp/ollama-day.log 2>&1 || echo "log fetch failed"
grep -c "truncating input prompt" /tmp/ollama-day.log
# 78
# msg="truncating input prompt" limit=4096 prompt=4173 keep=4 new=4095

On that same day, callers heard replies such as "Sembra che il cliente abbia risposto al telefono... posso continuare seguendo le istruzioni fornite" ("It seems the customer has answered the phone... I can continue following the instructions provided"). Another offered to "fare una simulazione della conversazione" ("run a simulation of the conversation"). We didn't match each of these replies to a specific truncated request. The pattern fits a model that has lost part of its instructions, but treat that link as a hypothesis.

Lesson: budget tokens in code. Keep the fully rendered prompt plus a reserve for the reply under the context window. Alert on the daily truncation count, because the runtime doesn't fail. It just quietly gets worse.

Are prompt rules enough to keep an 8B model safe?

No. The clearest example was a clinical rule: never suggest full-arch treatment ("All-on-4", "riabilitazione dell'arcata") to someone who is missing only one or two teeth. The rule was in the prompt, and one run of the adversarial eval still produced this:

Caller: "Mi mancano due denti davanti." ("I'm missing two front teeth.") Agent: "...La soluzione per te potrebbe essere l'All-on-4 o All-on-6..." ("...The solution for you could be All-on-4 or All-on-6...")

In a scripted golden conversation, the caller said "A me mancano due denti" ("I'm missing two teeth") and the model answered "visto che mancano tutti i denti" ("since all your teeth are missing").

The fix was deterministic code. A per-turn eligibility scope classifies the caller's own words as "few" or "many" missing teeth. When the answer is "few", full-arch FAQs are dropped from retrieval and a speech guard scrubs full-arch phrases from the output. Here is an excerpt (not a complete program) of the detection patterns and test cases, from the Python mirror of the production C# code:

# excerpt: FEWMISS = "few teeth missing"; SOLO and CNT are small helper patterns
FEWMISS = re.compile(
    r"manca(?:no)?\s+" + SOLO + CNT + r"?\s*(?:dent|incisiv|molar|canin|premolar)"
    r"|\bmi\s+manca\b|\bme\s+ne\s+manca(?:no)?\b"
    # ... more alternatives omitted ...
    , re.IGNORECASE)

DETECTION_CASES = [
    (["mi mancano due denti"], True),
    (["non sono senza denti, me ne mancano due"], True),
    (["ho solo due denti"], False),
    (["mi sono rimasti due denti"], False),
    (["ho perso tutti i denti", "in realtà me ne mancano due"], True),
    # ... 20 more ...
]

The full regression corpus has 25 detection cases and 41 output-scrub cases, and all 66 pass. It tests a Python mirror of the production C# patterns, not every streamed chunk boundary, cache hit and fallback in the deployed build. That needs integration tests against the deployed version. A reviewer also found a phrasing the corpus misses: the mirror classifies "non ho solo due denti, me ne mancano due" ("I don't have only two teeth, I'm missing two") as near-total loss. Regex guards are only as good as their adversarial cases. After the guard went in, the same eval answered "impianti singoli o un piccolo ponte" ("single implants or a small bridge").

The same went for formal address ("Lei"), hallucinated phone numbers and meta-replies: the prompt rule came first and failed in testing, then code guards replaced it.

Can a safety guard break the business goal?

Yes, and nothing alerts you when it happens. The goal was to give interested callers the WhatsApp number. A guard stopped the model from giving the number before the caller agreed. At the time, the affirmative branch was a short full-match allowlist ("sì", "va bene", "d'accordo", "certo", "ok"), plus exceptions for explicit requests such as "mi dia il numero". Real callers answered the invite with things like "Va bene, ci penso e poi ve la mando. Di dove siete?" ("OK, I'll think about it and send it. Where are you based?"), which failed the match. The guard then deleted the model's number sentence.

Our proxy was the number of calls in which an agent_reply event contained the spelled-out country-code prefix of the number:

2026-09-08: 2 of 86       2026-09-16: 0 of 1049  (guard stripped it in 17 calls)
2026-09-09: 1 of 380      2026-09-17: 1 of 1001  (stripped in 16)
2026-09-11: 1 of 386      2026-09-21: 1 of 880
2026-09-15: 0 of 778      2026-09-22: 2 of 1300

This is a text proxy, not proof of what the caller heard. In the one Sep 17 call where the number appeared, playback events show it came out garbled and the caller interrupted. For more than two weeks, the main conversion step barely appeared even in the logged text, while call volumes looked healthy.

Lesson: every guard needs a counter. Measure the conversion step from playback events, then spot-check against recordings.

What do real calls reveal that test calls don't?

We reviewed one full day of traces (1,001 calls, 385 with any transcribed caller speech) and found failure modes that no synthetic test had produced.

Stuck VAD on stalled audio. 30 calls ran past 60 s with no transcribed caller speech. Every one had an unmatched vad_start and no idle reprompt, and ended when the AudioSocket peer closed, at about 185 s or about 315 s. Together they took roughly 99 gateway-session minutes. We didn't check CDRs for how much of that was billed. The cause was in the gateway code: the VAD force-end only ran inside the per-frame audio handler. When AudioSocket frames stopped arriving mid-utterance, in_speech stayed true, and that blocked the idle and max-duration timers. Here is a filter for candidate calls, written for our trace schema (one JSONL file per call, rel = seconds since the call started):

import glob, json, sys

TRACE_DIR, DAY = sys.argv[1], sys.argv[2]   # e.g. ./traces 20260917

for path in sorted(glob.glob(f"{TRACE_DIR}/{DAY}-*.jsonl")):
    heard = vad_open = reprompted = False
    duration = end_reason = None
    with open(path) as fh:
        for line in fh:
            try:
                e = json.loads(line)
            except ValueError:
                continue
            t = e.get("type")
            if t == "stt" and (e.get("text") or "").strip():
                heard = True
            elif t == "vad_start":
                vad_open = True
            elif t == "vad_end":
                vad_open = False
            elif t == "idle_reprompt":
                reprompted = True
            elif t == "call_end":
                duration, end_reason = e.get("rel"), e.get("hangupReason")
    if duration and duration > 60 and not heard and vad_open and not reprompted:
        print(path, round(duration), end_reason)

The fix belongs in a watchdog outside the audio callback. If in_speech is set and no frame has arrived for a few seconds, close the utterance so the normal idle reprompt and goodbye logic can run. Keep that separate from any decision to hang up, and keep an absolute call deadline as well. Don't hang up just because media went quiet: a legitimate call can be silent, and some endpoints use silence suppression.

IVR loops. The backend's transfer intent matched "prema 1" ("press 1"). So did a restaurant's phone menu, and the agent answered it with the same canned line 24 times. Cap identical consecutive replies.

Losing people at the greeting. In the first two days of real calls (37 calls), 15 callers hung up during the greeting without saying a word. All 5 callers whose first sentence named a concrete need (a quote, a denture, implants, yellow teeth, a full set of teeth) were tagged as leads in our review. That is a small sample, and stating a need naturally predicts a lead tag. We didn't test what caused the early hangups. Our working hypothesis is the opening, and it needs a controlled test.

"Grazie a tutti" was also the first transcribed "utterance" in three of those calls, a known Whisper output on silence. Discard such phrases only when VAD or no-speech evidence agrees, since a person can really say them.

Does RAG help a small model on the phone?

Less than we assumed. We replayed 292 real caller utterances offline: 241 from this agent and 51 from the hosted platform described below. The replay used the live embedder (nomic-embed-text), a read-only search of the live Qdrant collection, and a copy of the production ranking logic. For the 136 utterances that were real questions or objections, we hand-labelled the correct FAQ:

Measure Result
Top-1 correct, live query (prev. turn + utterance) 54 / 136 (40%)
Top-1 correct, raw utterance only 68 / 136 (50%)
None of the top 3 correct 57 / 136 (42%)
Top-1 correct, by source (live query) this agent 47/107, hosted 7/29
Lowest top-1 / third-ranked cosine, all 292 0.66 / 0.62

The live config had contextMinScore: 0.52. Even the third-ranked candidate never scored below 0.62, so that cut-off removed nothing in this replay. Prepending the previous caller turn ("contextual" query) cost about 10 points of top-1 accuracy compared with the raw utterance. Correct and wrong top-1 hits had overlapping scores, so a higher cut-off alone discards correct answers too.

Lesson: replay labelled, real utterances against your index before you trust it. Pick a threshold by measuring accuracy, coverage and false rejections on those labels, and allow retrieval to return nothing.

Was a hosted no-code voice platform the answer?

We ran a hosted no-code voice-agent platform in parallel on cold outbound calls, with a 52-node conversation flow. From its own call export over about ten days:

Metric Value
Real outbound dials 12,160
Connected 4,285
Calls where a human spoke (platform label) 2,363
Of those, left during the intro (platform label) 2,156 (91%)
Tagged "warm" 5
Tagged "agreed to send photos" 0
Total platform cost about $266

The labels come from the platform's own LLM post-call analysis, and we had no separate downstream outcome data. Like our agent, it lost most callers early. Testing a different opening is the next experiment. Nothing here shows whether the opening matters more than the model.

What does a production readiness checklist look like?

Each item comes from a failure above.

Call identity and integration

Capacity and latency

Dialer front end

LLM and prompt

Guards and funnel

Real-call hygiene

Retrieval

Release

What would I do differently?

I'd build the measurement before the agent: per-call identity, a playback-based funnel, guard counters and a concurrency benchmark. Demos and service dashboards missed nearly every problem here. Targeted tests and real-call reviews exposed them.

Stuck on something specific?

Book a free 30-minute call. I run ViciDial call centers serving callers in 6 countries and can usually unblock your setup in one session — or build it for you.

Book a Free Consultation