Complete guide to ai & voice agents for ViciDial and Asterisk. 10 in-depth tutorials covering everything from basics to advanced production setups.
Score call audio quality automatically with a FastAPI service that combines the NISQA neural model, Silero VAD, and Claude for plain-English verdicts.
Build a real-time AI voice agent that answers Asterisk calls: AudioSocket streaming into Deepgram STT, a Groq LLM, and Cartesia TTS replies.
Connect ElevenLabs conversational AI agents to real phone calls over SIP, with webhook tools and dynamic call context injection.
Batch transcription of ViciDial call recordings using Faster-Whisper (OpenAI Whisper optimized with CTranslate2) for speech-to-text at scale — on CPU, without cloud APIs.
Design voice-agent personalities, conversation flows, and tool-calling workflows using patterns mined from real call transcriptions.
ElevenLabs cloud vs a local Deepgram, Groq and Cartesia stack for voice agents: architecture, latency, cost per call, and migration path.
Build an automated quality assurance pipeline that transcribes every inbound call using Faster-Whisper and scores agent performance with AI — running entirely on your existing ViciDial server with zero impact on live calls.
Build a production voice agent with OpenAI's native speech-to-speech Realtime API wired into Asterisk through AudioSocket.
Build a live dashboard that transcribes active calls and shows per-call sentiment in real time, using faster-whisper and a lightweight web UI.
Build a self-hosted answering machine detection (AMD) system that replaces Asterisk's built-in `AMD()` application with a Whisper-based speech recognition + machine learning classifier pipeline. Traditional AMD relies on energy detection and cadence analysis, achieving only 60-70% accuracy in real-world conditions — misclassifying live humans as machines (killing revenue-generating calls) and letting voicemail greetings through to agents (wasting expensive seat time). This tutorial's AI approach transcribes the first 3-5 seconds of answered audio using OpenAI's Whisper model, then feeds the transcript and audio features into a trained ML classifier that distinguishes human pickups from answering machines with 95%+ accuracy. The entire system runs on your own hardware with no per-call API costs, processes decisions in under 2 seconds, and continuously improves as you feed it new labeled data from your call center's actual traffic.