Text-to-speech: make APX speak out loud — "la voz", "que me hable", "en voz alta", "motor de voz", "leelo" — Piper local, ElevenLabs, OpenAI or Gemini. Load when configuring a voice engine, adding a custom TTS server, enabling emotion tags, or fixing silent output.
12 stars
0 votes
0 copies
0 views
Added September 27, 2026
ai-agentsrustgobashtestinggitapibackend
Works with
cli
api
Security analysis
C71/100
mediumUses curl or wget to download content
criticalExfiltrates credentials via HTTP — exact pattern from Snyk ToxicSkills study
criticalSends environment variables or credentials to an external URL
Installs into .claude/skills of the current project.
Are you the author of Apx Voice?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/agentprojectcontext-apx-voice)
---
name: apx-voice
scope: optional
description: Text-to-speech: make APX speak out loud — "la voz", "que me hable", "en voz alta", "motor de voz", "leelo" — Piper local, ElevenLabs, OpenAI or Gemini. Load when configuring a voice engine, adding a custom TTS server, enabling emotion tags, or fixing silent output.
---
# apx-voice
TTS facade in `core/voice/` with the built-in engines below plus any number of custom OpenAI-compatible endpoints. STT lives separately in `core/voice/transcription.js` (Whisper). The "voice channel" combines both for mic→agent→speaker.
## Engines
| id | Local? | Needs key? | Notes |
|---|---|---|---|
| `piper` | yes | no | Local, offline. Requires `piper` CLI + `.onnx` model. es_AR-daniela-high recommended. |
| `elevenlabs` | no | yes | Excellent. Free tier 10k chars/mo. `eleven_multilingual_v2`. |
| `openai` | no | yes | Reuses `engines.openai.api_key`. `tts-1`. Set `base_url` to point it at a custom endpoint instead. |
| `gemini` | no | yes | Returns raw L16 PCM — APX wraps in WAV automatically. Supports emotion tags. |
| `custom:<slug>` | no | maybe | Any OpenAI-compatible TTS server (e.g. QVox / Qwen3-TTS). Config-only `base_url`, own key (often keyless). Supports emotion tags. |
| `mock` | yes | no | Silent WAV; placeholder for tests. |
`auto` probes: piper → elevenlabs → openai → gemini → mock (custom providers are never auto-selected — pick one explicitly).
## Custom OpenAI-compatible endpoints (QVox / Qwen3-TTS)
Point APX at a local or remote OpenAI-compatible `/v1/audio/speech` server without any hardcoding — it's pure config. Two shapes:
- `voice.tts.openai.base_url` — reuse the `openai` engine against a custom endpoint.
- `voice.tts.custom.<slug>` — a named custom provider, surfaces as engine id `custom:<slug>`, backed by the openai adapter.
When `base_url` is set, APX additionally forwards the non-OpenAI fields the server understands — `instruct` (base voice, from the style arg), `language`, `temperature` — and defaults the response format to `wav` (stock OpenAI stays `mp3`). A custom endpoint uses **only its own `api_key`** (often none) and never leaks `engines.openai.api_key` / `OPENAI_API_KEY`.
## Emotion tags (per-engine capability)
Some backends (QVox/Qwen3-TTS, Gemini) accept inline `[tag]` markers and switch speaking emotion per segment while keeping the base voice. This is a **generic per-engine toggle**, not hardcoded to any adapter:
```bash
apx config set --global voice.tts.custom.qvox.emotions.enabled true
# optional: restrict the tag set (defaults to the canonical QVox set)
apx config set --global voice.tts.custom.qvox.emotions.tags '["happy","sad","excited","calm","whisper","laugh","neutral"]'
```
Default tags: `happy, sad, excited, angry, calm, whisper, shout, laugh, cry, narrator, neutral`. The voice-mode prompt learns the tag syntax **only when the engine that will actually speak has emotions enabled**. On any engine without tag support, stray `[tags]` are stripped from the displayed text and never read aloud (kept only for the speaking engine's audio).
## Concrete CLI calls
```bash
apx voice providers # what's configured + available
apx voice say "Hello from APX" --provider piper
apx voice say "Hello from APX" --provider gemini --voice Aoede
apx voice say "..." --no-play # generate WAV, don't play
apx voice listen # mic → STT, records until silence (sox) or Ctrl+C
apx voice listen --seconds 5 # fixed-duration capture
apx voice listen --provider <id> # override STT provider
```
Playback uses system binaries (`afplay`, `paplay`, `aplay`, `play`, `ffplay`). If none found, you get the file path and no playback.
## Spoken replies while driving
When the daemon knows a trip is in progress (Android reports it — see
`rules/android.md`), every automatic Telegram reply goes out twice: a voice note
first, then the same words as text under a transcript header. The turn itself
runs in voice mode, so the reply is one or two spoken sentences rather than a
paragraph read aloud.
The audio is OGG/Opus — the only container Telegram renders as a voice note —
converted from whatever the engine returned via `ffmpeg`. A failing engine or a
missing `ffmpeg` costs the audio, never the message: it falls back to plain
text. The `mock` engine is silence, not speech, and is refused.
```bash
apx config set --global voice.mobility_replies false # keep replies text-only in the car
```
## Configuration
`~/.apx/config.json → voice.tts.<engine>`:
```json
{
"voice": {
"tts": {
"provider": "gemini",
"piper": { "bin": "piper", "model": "/Users/.../es_AR-daniela-high.onnx" },
"elevenlabs": { "api_key": "...", "model": "eleven_multilingual_v2", "voice_id": "..." },
"openai": { "api_key": "...", "model": "tts-1", "voice": "alloy", "format": "mp3" },
"gemini": { "api_key": "...", "model": "gemini-2.5-flash-preview-tts", "voice": "Aoede", "emotions": { "enabled": true } },
"custom": {
"qvox": {
"label": "QVox local",
"base_url": "http://127.0.0.1:5111/v1",
"api_key": "",
"format": "wav",
"language": "es",
"emotions": { "enabled": true }
}
}
}
}
}
```
`apx config set --global voice.tts.provider <name>` to switch. Voice config is read from the global config, so `--global` is required — without it the value lands in one project's config and the voice stack never sees it.
## Quick setup: Piper local (recommended, no internet)
```bash
# 1. Install binary (macOS arm64)
curl -L https://github.com/rhasspy/piper/releases/latest/download/piper_macos_aarch64.tar.gz \
-o /tmp/piper.tar.gz
sudo tar xzf /tmp/piper.tar.gz -C /usr/local/bin --strip-components=1
# 2. Voice model (es_AR, "daniela")
mkdir -p ~/.apx/voices && cd ~/.apx/voices
curl -LO https://huggingface.co/rhasspy/piper-voices/resolve/main/es/es_AR/daniela/high/es_AR-daniela-high.onnx
curl -LO https://huggingface.co/rhasspy/piper-voices/resolve/main/es/es_AR/daniela/high/es_AR-daniela-high.onnx.json
# 3. Configure + test
apx config set --global voice.tts.provider piper
apx config set --global voice.tts.piper.model "$HOME/.apx/voices/es_AR-daniela-high.onnx"
apx voice say "hola, soy APX" --provider piper
```
## Quick setup: Gemini cloud
```bash
apx config set --global voice.tts.provider gemini
apx config set --global voice.tts.gemini.api_key '<GEMINI_KEY>'
apx config set --global engines.gemini.api_key '<GEMINI_KEY>' # reuse for LLM router
apx voice say "Hello from APX" --provider gemini
```
## Unified voice channel
`POST /api/voice/turn` is one round-trip: send audio (or text), get back `{ user_text, reply_text, reply_audio_path }`. STT in, agent loop, TTS out. For overlay and future "voice room" clients.
```bash
curl -X POST http://127.0.0.1:7430/api/voice/turn \
-H "Authorization: Bearer $(cat ~/.apx/daemon.token)" \
-H "Content-Type: application/json" \
-d '{"text":"Hello APX","channel":"voice"}'
```
`channel` names the surface the turn arrives on: `voice` (the deck's spoken mode), `deck`, `desktop`, `telegram`. Anything else — including omitting it — is treated as `api`, the bounded default: the route will not adopt another surface's behaviour just because a caller names it.
Telegram voice messages and overlay mascot still have their own STT pipelines — they don't go through `/api/voice/turn` yet.
## Anti-examples
- DON'T trust `apx voice providers` saying "mock available" as green light — mock is silence. Configure a real provider.
- DON'T set `voice.tts.provider` to a provider with no key. It falls through `auto` to the next, but that's not what you asked.
- DON'T expect Gemini TTS to return MP3 — it returns raw L16 PCM; APX wraps in WAV. Files are `.wav`, mime `audio/wav`. Convert with ffmpeg if you need MP3.
## Troubleshooting silent output
1. `apx voice providers` — what's actually available?
2. `apx voice say "test" --provider <engine> --no-play` — file exists?
3. `file <path>` — valid container? Gemini output should be `RIFF WAVE Microsoft PCM`.
4. `afplay <path>` — does the OS player open it?
5. If 3 fails for Gemini, you may be on APX before the PCM-wrap fix (commit `ba5c416`+).
## Don't
- Paste base64 audio into chat. Use file paths or `send_voice` / `send_audio`.
- Switch providers mid-routine without testing — quality varies a lot across Piper voices and cloud engines.
- Expect TTS streaming yet — `apx voice say` returns a complete file. `/api/tts/stream` is open work.