Skip to content
Back to skills

Tts Generation

ASecurity

AI text-to-speech generation using OpenAI TTS, ElevenLabs, and Google TTS backends. Converts text to audio files with voice selection, speed control, and format options.

  • 40 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 6, 2026
developmentpythonrustgobashapibackenddocumentation

Works with

  • cli
  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 10 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add oimiragieo/agent-studio --skill tts-generation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Tts Generation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Tts Generation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/oimiragieo-tts-generation/badge)](https://www.skillsdirectory.com/skills/oimiragieo-tts-generation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: tts-generation
description: AI text-to-speech generation using OpenAI TTS, ElevenLabs, and Google TTS backends. Converts text to audio files with voice selection, speed control, and format options.
version: 1.0.0
model: sonnet
invoked_by: both
user_invocable: true
tools: [Read, Write, Bash, WebFetch]
agents: [developer, ai-ml-expert]
category: 'AI/ML'
tags: [tts, text-to-speech, audio, openai, elevenlabs, google-tts, voice]
best_practices:
  - Choose model tier based on latency vs quality requirements (tts-1 for speed, tts-1-hd for quality)
  - Cache generated audio files to avoid re-generating identical text
  - Handle API rate limits with exponential backoff
  - Validate text length before API call (OpenAI max 4096 chars per request)
error_handling: strict
source: builtin
trust_score: 100
provenance_sha: e5bb24bfa4f8179d
---

# TTS Generation

## Overview

Generate speech audio from text using AI backends.

- **OpenAI TTS** — `tts-1` (low latency) / `tts-1-hd` (studio quality), 6 voices, 57 languages
- **ElevenLabs** — `eleven_turbo_v2` / `eleven_multilingual_v2`, cloneable voices, 29 languages
- **Google TTS** — `gTTS` Python library, 40+ languages, free tier

## Backend Comparison

| Feature   | OpenAI TTS    | ElevenLabs    | Google TTS     |
| --------- | ------------- | ------------- | -------------- |
| Quality   | High          | Highest       | Medium         |
| Latency   | Low (tts-1)   | Medium        | Low            |
| Cost      | ~$15/1M chars | ~$22/1M chars | Free (limited) |
| Voices    | 6 preset      | Cloneable     | 40+ languages  |
| Max chars | 4096/request  | Unlimited     | ~5000/request  |
| Streaming | Yes           | Yes           | No             |

## Quick Start

### OpenAI TTS (Recommended)

```python
from pathlib import Path
from openai import OpenAI

client = OpenAI()

response = client.audio.speech.with_streaming_response.create(
    model="tts-1-hd",  # tts-1 for speed, tts-1-hd for quality
    voice="nova",       # alloy | echo | fable | onyx | nova | shimmer
    input="Hello world",
    speed=1.0,          # 0.25 to 4.0
)
response.stream_to_file(Path("output.mp3"))
```

### ElevenLabs

```python
from elevenlabs import ElevenLabs

client = ElevenLabs(api_key="YOUR_API_KEY")
audio = client.text_to_speech.convert(
    voice_id="21m00Tcm4TlvDq8ikWAM",  # Rachel
    model_id="eleven_turbo_v2",
    text="Hello world",
    output_format="mp3_44100_128",
)
with open("output.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)
```

### Google TTS (Free)

```python
from gtts import gTTS
gTTS(text="Hello world", lang="en", slow=False).save("output.mp3")
```

## Long-Text Chunking

For text exceeding limits, split at sentence boundaries and concatenate with `pydub`. Pattern: iterate sentences, accumulate into `current` until `max_chars` (4000), flush to `chunks` on overflow.

## Output Formats

`mp3` (general), `opus` (streaming), `flac` (lossless archival), `wav` (editing), `pcm` (raw pipeline).

## Installation

```bash
pip install openai elevenlabs gtts pydub
export OPENAI_API_KEY="sk-..."
export ELEVENLABS_API_KEY="..."
```

## Agent Usage Pattern

- OpenAI TTS: documentation/demos narration
- ElevenLabs: cloned voices or highest quality
- Google TTS: multilingual free-tier
- Chunk at sentence boundaries; cache by content hash

## Related Skills

- `transcription` — Reverse: audio to text via Whisper
- `ai-ml-expert` — Advanced ML pipeline integration

## Memory Protocol (MANDATORY)

**Before starting:**
Read `.claude/context/memory/learnings.md`

**After completing:**

- New pattern → `.claude/context/memory/learnings.md`
- Issue found → `.claude/context/memory/issues.md`
- Decision made → `.claude/context/memory/decisions.md`

> ASSUME INTERRUPTION: If it's not in memory, it didn't happen.

Files in this skill

  • SKILL.md3.7 KB
  • commands/tts-generation.md114 B
  • hooks/post-execute.cjs181 B
  • hooks/pre-execute.cjs287 B
  • references/research-requirements.md289 B
  • rules/tts-generation.md287 B
  • schemas/input.schema.json490 B
  • schemas/output.schema.json258 B
  • scripts/main.cjs1009 B
  • templates/implementation-template.md188 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…