Skip to content
Back to skills

Nemotron Voice Agent Builder

ASecurity

Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. Use when building, scaffolding, or iterating on a real-time voice-agent pipeline, including speech (ASR/TTS) customization and cloud or local deployment. Not for offline/batch speech-to-text, text-only chat or RAG, or generic Docker or CUDA work unrelated to a voice agent.

  • 3,503 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 24, 2026
ai-agentsdockerdocumentation

Works with

  • cli

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned September 24, 2026

npx -y skills add NVIDIA/skills --skill nemotron-voice-agent-builder --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Nemotron Voice Agent Builder?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Nemotron Voice Agent Builder
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/nvidia-nemotron-voice-agent-builder/badge)](https://www.skillsdirectory.com/skills/nvidia-nemotron-voice-agent-builder)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: nemotron-voice-agent-builder
description: Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. Use when building, scaffolding, or iterating on a real-time voice-agent pipeline, including speech (ASR/TTS) customization and cloud or local deployment. Not for offline/batch speech-to-text, text-only chat or RAG, or generic Docker or CUDA work unrelated to a voice agent.
version: "2.2.0"
license: CC-BY-4.0 AND Apache-2.0
metadata:
  author: NVIDIA Voice Agent Team <nemotron-voice-agent@nvidia.com>
  tags: "voice-agent, nvidia, nemotron, magpie, pipecat, livekit, nim, vllm, nemo-speech, omni"
---

# Create Voice Agent

Creates a working NVIDIA voice agent or updates an existing project:

- **Cascaded**: ASR transcribes, a text LLM answers, TTS speaks.
- **Omni**: one audio-in LLM replaces ASR and the text LLM. TTS still speaks.

## When to Use This Skill

Use this skill when the user wants to build, scaffold, configure, refine, or fix an
NVIDIA voice agent — a real-time speech pipeline with audio input and spoken output — on
Pipecat or LiveKit. This covers Cascaded (ASR → LLM → TTS) and Omni (audio-in LLM + TTS)
pipelines, speech customization (ASR word boosting, TTS pronunciation), multilingual
routing, and cloud or local deployment (NIM, vLLM, NeMo-Speech.cpp) on workstations, DGX
Spark, or Jetson Thor.

Do not use this skill for text-only chatbots or RAG, standalone batch speech-to-text
transcription, generic Docker or infrastructure help, or unrelated CUDA or model work that
has no voice-agent pipeline.

## Workflow

Classify the starting state before following the phases:

- For an empty project, follow all phases in order.
- For a working existing project, read `references/operations/iterate.md` first, and then route only the references needed for the requested change.
- For a broken existing project, read `references/operations/troubleshoot.md` first. After restoring the baseline, continue with `references/operations/iterate.md`.

For an empty project, follow these phases in order:

| Phase | Read |
| --- | --- |
| 1. Intake | `references/intake.md` resolves pipeline and framework first, then routes preflight, models, language, and domain files |
| 2. Build | `references/platforms/readiness.md` if self-hosted → exact profile or quantization variant in the routed model file → `references/output-contract.md` → routed deployment path → `references/operations/observability.md` → selected framework files |
| 3. Hand over | `references/operations/run.md` gates the client on the generated `scripts/smoke.sh`, then `references/operations/troubleshoot.md` if the spoken exchange fails |
| 4. Iterate | `references/operations/iterate.md` for changes to a working agent |

Wait for approval before writing files.

`references/preflight.md` §4 probes the host and selects the deployment path. The routed
platform file owns that path from there.

Also read `references/networking/remote-webrtc.md` when Pipecat WebRTC crosses hosts or
networks.

## Rules

- Resolve every model id, image tag, profile, and serve flag from current NVIDIA
  documentation at build time. Never from memory.
- Read every routed file before generating code, and keep the language and behavior locks
  approved in intake.
- Handover requires a successful spoken exchange. A running process is not a working voice
  agent.

Files in this skill

  • BENCHMARK.md7.9 KB
  • SKILL.md3.3 KB
  • evals/EVAL.md2.8 KB
  • evals/evals.json4.6 KB
  • references/domain/agent-behavior.md3.1 KB
  • references/domain/speech-customization.md8.2 KB
  • references/frameworks/livekit.md6 KB
  • references/frameworks/omni.md7.8 KB
  • references/frameworks/pipecat.md5.3 KB
  • references/intake.md8.3 KB
  • references/models/asr.md6.8 KB
  • references/models/catalog.md4.2 KB
  • references/models/language-routing.md5.6 KB
  • references/models/llm.md18.3 KB
  • references/models/tts.md5.7 KB
  • references/networking/remote-webrtc.md5 KB
  • references/operations/iterate.md4.5 KB
  • references/operations/observability.md3.1 KB
  • references/operations/run.md4.5 KB
  • references/operations/troubleshoot.md11.4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…