Skip to content
Back to skills

Ai Voice Bots

ASecurity

Builds production voice bots and IVR with Python STT/TTS pipelines. Use when designing telephony, streaming audio, latency budgets, or voice quality monitoring.

  • 89 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 2, 2026
ai-agentspythonrailsazuretestingapidocumentation

Works with

  • claude code
  • api

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add vasilyu1983/AI-Agents-public --skill ai-voice-bots --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Voice Bots?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Voice Bots
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vasilyu1983-ai-voice-bots/badge)](https://www.skillsdirectory.com/skills/vasilyu1983-ai-voice-bots)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-voice-bots
description: "Builds production voice bots and IVR with Python STT/TTS pipelines. Use when designing telephony, streaming audio, latency budgets, or voice quality monitoring."
compatibility: Portable core. Works on Claude Code and Codex.
version: "1.2"
last_validated: 2026-07-11
---

# AI Voice Bots

Use this skill to build, ship, and tune voice bots — phone IVR, real-time speech agents, and voice-first customer interactions — using pure Python frameworks.

This skill owns the voice-specific pipeline: STT, TTS, telephony platforms, latency engineering, and voice quality. For conversation design, persona, and escalation patterns, use [`../ai-bot-builder/SKILL.md`](../ai-bot-builder/SKILL.md).

Default to a streaming STT→LLM→TTS pipeline when text inspection, redaction, or deterministic call state is required. Choose Pipecat for custom transports and LiveKit Agents for LiveKit room or SIP integration. Test speech-to-speech (S2S) against the same target calls when latency or multimodal behavior may justify a different pipeline; see [references/s2s-and-native-voice-apis.md](references/s2s-and-native-voice-apis.md).

## When to Use This Skill

- Building a voice bot for phone, IVR, or real-time speech
- Choosing a telephony platform (Twilio, Vapi, Bland.ai, Retell, Telnyx, Vonage)
- Choosing a voice pipeline framework (Pipecat, LiveKit Agents, Vocode)
- Engineering latency budgets for voice (TTFB, total turn latency)
- Selecting and configuring STT/TTS providers (Deepgram, ElevenLabs, Cartesia, Azure)
- Monitoring voice quality (MOS, WER, call completion rate)
- Designing IVR flows with DTMF and voice hybrid
- Building outbound dialing campaigns

## When NOT to Use This Skill

| Need | Route to |
|------|----------|
| Bot conversation design, persona, escalation | [`../ai-bot-builder/SKILL.md`](../ai-bot-builder/SKILL.md) |
| Text-only bot architecture | [`../ai-bot-builder/SKILL.md`](../ai-bot-builder/SKILL.md) |
| General agent architecture | [`../ai-agents/SKILL.md`](../ai-agents/SKILL.md) |
| WebSocket/SSE infrastructure (non-voice) | [`../software-realtime/SKILL.md`](../software-realtime/SKILL.md) |
| Voice/multimodal reference material | [`../ai-agents/references/voice-multimodal-agents.md`](../ai-agents/references/voice-multimodal-agents.md) |

## Quick Reference

| Need | Default | Notes |
|------|---------|-------|
| Choose telephony platform | `references/telephony-platform-selection.md` | Twilio, Vapi, Bland.ai, Retell, Telnyx, Vonage |
| Design voice pipeline | `references/voice-pipeline-architecture.md` | STT→LLM→TTS streaming, codec selection |
| Build with Pipecat | `references/pipecat-patterns.md` | Processors, transports, production deployment |
| Build with LiveKit Agents | `references/livekit-agents-patterns.md` | AgentSession, rooms, plugins |
| Optimize latency | `references/latency-engineering.md` | Component budgets, edge deployment, caching |
| Monitor voice quality | `references/voice-quality-metrics.md` | MOS, WER, dashboards, alerting |
| Design IVR flows | `references/ivr-design.md` | DTMF, menu trees, hybrid voice+keypad |
| Voice compliance | `references/voice-safety-compliance.md` | Recording consent, PCI, TCPA, GDPR |
| Deploy voice bot to 24/7 production | [references/production-deployment.md](references/production-deployment.md) | Concurrent-call capacity, SIP/PSTN HA, autoscaling, drain, recording compliance, cost model |
| Pick a hosting platform (LiveKit Cloud + Fly.io, Pipecat Cloud, etc.) | [`../software-paas-hosting/references/agent-hosting-matrix.md`](../software-paas-hosting/references/agent-hosting-matrix.md) | Voice stacks BV1–BV3 + what does NOT host voice |

## Default Workflow

1. **Define call flow** — inbound vs outbound, IVR menu tree, conversation states.
2. **Choose telephony platform** — by volume, region, compliance, and API quality.
3. **Choose voice pipeline framework** — Pipecat (default) or LiveKit Agents.
4. **Set latency budgets** — measure each stage on target call conditions and set a product-specific percentile target.
5. **Select STT/TTS providers** — by language support, latency, quality, and cost.
6. **Integrate conversation logic** — use `ai-bot-builder` patterns for the LLM "brain."
7. **Add voice-specific guardrails** — AI identity disclosure where required, recording consent as a separate step, PII in speech, and barge-in safety; use [voice-safety-compliance.md](references/voice-safety-compliance.md).
8. **Instrument voice quality metrics** — MOS, WER, call completion, latency percentiles.
9. **Load test and tune** — verify latency under concurrent call load.

## Voice Pipeline Architecture

```
Phone/WebRTC → Transport → STT → LLM → TTS → Transport → Phone/WebRTC
                  │          │      │      │         │
                  │          │      │      │         └── Audio codec encoding
                  │          │      │      └── Text-to-speech streaming
                  │          │      └── Conversation logic (ai-bot-builder)
                  │          └── Speech-to-text streaming
                  └── WebSocket / WebRTC / SIP
```

**Pipeline budget:** Measure each component and the end-to-end first-audio time before setting numerical targets. Streaming stages overlap, so summing component durations does not necessarily equal perceived turn latency.

| Component | Measure |
|-----------|---------|
| VAD or turn detector | End-of-speech to end-of-turn decision |
| STT | Audio to stable transcript, including partials |
| LLM | Request to first usable output token |
| TTS | Text availability to first playable audio |
| Transport | Caller network and SIP/WebRTC delivery |

Full depth → [references/voice-pipeline-architecture.md](references/voice-pipeline-architecture.md)

## Telephony Platform Selection

Default to the telephony provider already approved for the target geography and call type. Compare its current number/SIP availability, recording controls, transfer behavior, concurrent-call limits, and pricing against the [official provider documentation](references/telephony-platform-selection.md) before changing platforms.

Full comparison → [references/telephony-platform-selection.md](references/telephony-platform-selection.md)

## Framework Selection

| Framework | Best for | Transport | S2S support | Ecosystem |
|-----------|----------|-----------|-------------|-----------|
| **Pipecat** (default) | Custom voice pipelines, multi-transport | WebSocket, Twilio, Daily, WebRTC | Yes (OpenAI Realtime, Gemini Live) | Deepgram, ElevenLabs, Cartesia, Anthropic, OpenAI |
| **LiveKit Agents** | Room-based voice, recording, multi-party | LiveKit (WebRTC) | Yes (OpenAI Realtime) | LiveKit Cloud, STT/TTS plugins |
| **Vocode** | Simple voice bots, telephony focus | Twilio, Vonage, WebSocket | No | Deepgram, Azure, ElevenLabs |

Default: **Pipecat** — strongest Python ecosystem, composable pipeline processors, multi-transport support, and broadest S2S provider coverage.

Use **LiveKit Agents** when: multi-participant calls, built-in recording, or already using LiveKit infrastructure.

### S2S vs cascading

If an inline text gate must block unsafe output before the caller hears it, use the cascading pipeline. If that gate is not required and a same-call-set test shows S2S improves the chosen latency or experience metric, use S2S. A parallel transcript can help audit after the fact but cannot block already-spoken audio. Check the current model, regional availability and price in the provider docs before selection.

Full S2S reference → [references/s2s-and-native-voice-apis.md](references/s2s-and-native-voice-apis.md)

## Production Defaults

- **Framework:** Pipecat with streaming pipeline
- **STT/TTS:** Select from the providers supported by the installed framework. Compare language/accent accuracy, first-audio latency, interruption recovery and current provider terms on target calls; see [voice-pipeline-architecture.md](references/voice-pipeline-architecture.md).
- **LLM:** Choose by measured task success, tool reliability, latency and cost for the call flow.
- **Transport:** Use the approved SIP/PSTN or WebRTC integration for the target region and codec.
- **Latency target:** Set p50/p95 and interruption targets from the product's call tests; do not import a vendor benchmark as the SLO.
- **Quality monitoring:** MOS tracking, WER sampling, call completion rate
- **Compliance:** Recording consent per jurisdiction, PII redaction from transcripts

## Expert Judgment: Latency Budget and Barge-In

**Decompose the budget before optimizing.** Attribute delay to VAD, STT, LLM first token, TTS first audio, and network. The bottleneck can be turn detection or a cold TTS connection even when the LLM is fast. Full instrumentation → [latency-engineering.md](references/latency-engineering.md); queueing effects → [queueing-theory-applied.md](references/queueing-theory-applied.md).

**Barge-in is a cancellation path.** A user talking over the bot must stop the in-flight audio and pending generation promptly. Measure time to actual audio stop, false interruptions and recovery separately from end-of-turn latency. See [queueing-theory-applied.md](references/queueing-theory-applied.md) (P3, A3).

Recompare S2S and cascading when changing models or providers. Use the same target-region calls to measure first audio, task success, tool safety and transcript needs; see [s2s-and-native-voice-apis.md](references/s2s-and-native-voice-apis.md).

## Real-Call Launch Gate

Test the target languages, accents, codecs, handset networks, background noise, silence, barge-in, DTMF, transfer, provider timeout, and reconnect paths with real or replayed call audio. Report end-of-turn and first-audio latency percentiles, interruption success, task completion, false transfer, hang-up, and consent-capture rates by scenario. Launch only when every blocking scenario has a deterministic fallback and the concurrent-call test meets the same bounds. A provider connection or synthetic clean-audio demo is not launch evidence.

## Known Traps

- proving latency with synthetic lab prompts instead of real barge-in, interruption, packet-loss, and handset-network conditions
- treating telephony acceptance as conversation success when the real failure is post-answer latency, bad turn segmentation, or TTS overlap
- mixing recording, transcript retention, PCI redaction, and consent rules across regions without one explicit policy owner
- optimizing only average latency while ignoring p95 or p99 tails that make production calls feel broken
- shipping one STT or TTS provider path with no fallback, rollback, or degraded-mode behavior for provider incidents
- **S2S session-state loss on model switch**: switching between S2S model versions mid-session (any provider — OpenAI Realtime, Gemini Live) drops all ephemeral session state — voice, tone configuration, conversation history, and tool state are not carried over. Resolution: persist conversation state to an external store (Redis or Postgres) after every turn; reload from the store when resuming or switching models. Do not rely on the S2S session as a state store for anything you cannot afford to lose. See `references/s2s-and-native-voice-apis.md` for the full session management pattern.

## Common Anti-Patterns

- **Batch-style voice pipelines** — waiting for full utterances or full synthesis destroys turn-taking and makes the bot feel laggy
- **LLM-first architecture with no deterministic call state** — IVR routing, transfers, and compliance prompts need explicit state machines, not only prompt logic
- **One-metric quality reporting** — MOS alone or WER alone hides interruption quality, completion failures, and escalation pain
- **Treating outbound voice like chat automation** — dialing, consent, voicemail handling, and retry policy need channel-specific controls
- **Using text-bot guardrails unchanged for speech** — voice bots need barge-in, silence, DTMF, and speaking-over-user protections

## Navigation

**References**
- [references/index.md](references/index.md) — Reference navigation map
- [references/s2s-and-native-voice-apis.md](references/s2s-and-native-voice-apis.md) — S2S vs cascading, OpenAI Realtime API, Gemini Live, session management (Jul 2026)
- [references/telephony-platform-selection.md](references/telephony-platform-selection.md) — Platform comparison
- [references/voice-pipeline-architecture.md](references/voice-pipeline-architecture.md) — Pipeline design
- [references/pipecat-patterns.md](references/pipecat-patterns.md) — Pipecat deep dive
- [references/livekit-agents-patterns.md](references/livekit-agents-patterns.md) — LiveKit Agents deep dive
- [references/latency-engineering.md](references/latency-engineering.md) — Latency optimization
- [references/voice-quality-metrics.md](references/voice-quality-metrics.md) — Quality monitoring
- [references/ivr-design.md](references/ivr-design.md) — IVR flow design
- [references/voice-safety-compliance.md](references/voice-safety-compliance.md) — Voice compliance
- [references/queueing-theory-applied.md](references/queueing-theory-applied.md) — Queueing theory applied to voice: latency budget partitioning, jitter buffer sizing, Erlang-C IVR capacity, barge-in priority, TTS streaming targets

**Assets**
- [assets/voice-bot-spec.md](assets/voice-bot-spec.md) — Voice bot specification template
- [assets/voice-latency-budget.md](assets/voice-latency-budget.md) — Latency budget worksheet
- [assets/voice-quality-checklist.md](assets/voice-quality-checklist.md) — Pre-launch quality gate
- [assets/voice-eval-scenarios.md](assets/voice-eval-scenarios.md) — End-to-end voice-agent eval scenario template

**Scripts**
- `python3 scripts/voice_latency_audit.py --input pipeline_logs.jsonl` — Pipeline latency breakdown
- `python3 scripts/call_quality_scorer.py --input calls.jsonl` — Call quality scoring

**Data**
- [data/sources.json](data/sources.json) — Curated voice-specific sources

## Related Skills

- [../ai-bot-builder/SKILL.md](../ai-bot-builder/SKILL.md) — Conversation design, persona, escalation, LangGraph
- [../ai-context-layer/references/conversational-surfaces-cross-platform.md](../ai-context-layer/references/conversational-surfaces-cross-platform.md) — Cross-platform composition recipe; voice section specifies the RA13 hot/cold memory tier split required for sub-300 ms turn latency
- [../ai-agents/SKILL.md](../ai-agents/SKILL.md) — Agent architecture decisions
- [../ai-agents/references/voice-multimodal-agents.md](../ai-agents/references/voice-multimodal-agents.md) — Voice/multimodal agent reference
- [../software-realtime/SKILL.md](../software-realtime/SKILL.md) — WebSocket/SSE infrastructure
- [../qa-agent-testing/SKILL.md](../qa-agent-testing/SKILL.md) — Agent eval harnesses
- [../qa-observability/SKILL.md](../qa-observability/SKILL.md) — Pipeline telemetry

## Learnings Loop

When prior decisions or pitfalls are relevant, consult `learnings.consolidated.md` if present; use `learnings.md` only for needed history or as the available fallback. Otherwise skip both.

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to `learnings.md` via `agents-skills-feedback-loop/scripts/append_learning.py`. Do not modify `SKILL.md` itself.

Files in this skill

  • SKILL.md18.4 KB
  • agents/openai.yaml410 B
  • assets/voice-bot-spec.md3.1 KB
  • assets/voice-eval-scenarios.md2.9 KB
  • assets/voice-latency-budget.md2.3 KB
  • assets/voice-quality-checklist.md2.5 KB
  • data/sources.json5.5 KB
  • learnings.consolidated.md589 B
  • learnings.md395 B
  • references/index.md2.2 KB
  • references/ivr-design.md23.6 KB
  • references/latency-engineering.md21.2 KB
  • references/livekit-agents-patterns.md17.4 KB
  • references/pipecat-patterns.md24.1 KB
  • references/production-deployment.md17.9 KB
  • references/queueing-theory-applied.md30.7 KB
  • references/s2s-and-native-voice-apis.md12.3 KB
  • references/telephony-platform-selection.md13.2 KB
  • references/voice-pipeline-architecture.md23.1 KB
  • references/voice-quality-metrics.md18.4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…