Skip to content
Back to skills

Vad Tuner

ASecurity

Pick VAD model, threshold, silence hangover, pre-roll, and turn-detection strategy for a voice agent. Use when you need help with vad tuner.

  • 8 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 8, 2026
developmentgo

Security analysis

A100/100

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill vad-tuner --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Vad Tuner?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Vad Tuner
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-vad-tuner/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-vad-tuner)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: vad-tuner
description: Pick VAD model, threshold, silence hangover, pre-roll, and turn-detection strategy for a voice agent. Use when you need help with vad tuner.
license: CC-BY-NC-SA-4.0
phase: 6
lesson: 14
metadata:
  version: 1.0.0
  tags: [vad, silero, cobra, turn-detection, flush-trick]
---

Given the workload (consumer / call-center / edge / accessibility; noise profile; language mix; latency), output:

1. VAD. Silero VAD (default) · Cobra (commercial accuracy) · pyannote segmentation (diarization-grade) · WebRTC VAD (legacy / tiny). One-sentence reason.
2. Parameters. Threshold (0.3-0.5), min speech (200-300 ms), silence hangover (400-800 ms), pre-roll (250-500 ms).
3. Semantic turn detection. Enabled (LiveKit turn-detector or custom MLP) or not. Reason tied to expected user speech patterns.
4. Flush trick. Enabled (if STT supports it — Kyutai / Deepgram) or not. Expected latency savings.
5. Guards. Reject speech shorter than min duration; always keep pre-roll; cap per-user silence-hangover override; fail-open if VAD service is down (treat everything as speech).

Refuse energy-only VAD for production — too noisy. Refuse zero silence-hangover — will interrupt users. Refuse Whisper-based VAD when dedicated Silero is available (slower, less accurate).

Example input: "Call-center IVR for airline rebooking. Noisy background (airport). English + Spanish. < 500 ms turn detection."

Example output:
- VAD: Cobra (commercial) for the noise-resistance advantage. Fall-back to Silero if cost prohibitive.
- Parameters: threshold 0.4 (airport noise floor is high); min speech 300 ms; silence hangover 600 ms (users often pause during IVR to read flight numbers); pre-roll 400 ms.
- Semantic turn: LiveKit turn-detector enabled — mid-sentence pauses common ("I need to change my flight... to tomorrow").
- Flush trick: enabled on Deepgram streaming. Expected savings: 400 ms → 150 ms turn-end latency.
- Guards: fail-open if Cobra/Deepgram unreachable; audit log every VAD-fire event for tuning.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…