Skip to content
Back to skills

Transcribe

ASecurity

Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.

  • 70 stars
  • 0 votes
  • 4 copies
  • 11 views
  • Added September 12, 2026
ai-agentspythonshellbashapi

Works with

  • cli
  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 7 files and shows the line behind each finding

Scanned September 12, 2026

npx -y skills add seaworld008/Commonly-used-high-value-skills --skill transcribe --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Transcribe?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Transcribe
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/seaworld008-transcribe/badge)](https://www.skillsdirectory.com/skills/seaworld008-transcribe)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: transcribe
description: 'Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.'
zh_description: "用于将音频或视频中的语音转写为文本,并可结合说话人分离和已知说话人提示。"
version: "1.0.0"
author: "seaworld008"
source: "in-house"
source_url: ""
tags: '["transcribe"]'
created_at: "2026-03-04"
updated_at: "2026-06-29"
quality: 3
complexity: "intermediate"
---

# Audio Transcribe

Transcribe audio using OpenAI, with optional speaker diarization when requested. Prefer the bundled CLI for deterministic, repeatable runs.

## Workflow
1. Collect inputs: audio file path(s), desired response format (text/json/diarized_json), optional language hint, and any known speaker references.
2. Verify `OPENAI_API_KEY` is set. If missing, ask the user to set it locally (do not ask them to paste the key).
3. Run the bundled `transcribe_diarize.py` CLI with sensible defaults (fast text transcription).
4. Validate the output: transcription quality, speaker labels, and segment boundaries; iterate with a single targeted change if needed.
5. Save outputs under `output/transcribe/` when working in this repo.

## Decision rules
- Default to `gpt-4o-mini-transcribe` with `--response-format text` for fast transcription.
- If the user wants speaker labels or diarization, use `--model gpt-4o-transcribe-diarize --response-format diarized_json`.
- If audio is longer than ~30 seconds, keep `--chunking-strategy auto`.
- Prompting is not supported for `gpt-4o-transcribe-diarize`.

## Output conventions
- Use `output/transcribe/<job-id>/` for evaluation runs.
- Use `--out-dir` for multiple files to avoid overwriting.

## Dependencies (install if missing)
Prefer `uv` for dependency management.

```
uv pip install openai
```
If `uv` is unavailable:
```
python3 -m pip install openai
```

## Environment
- `OPENAI_API_KEY` must be set for live API calls.
- If the key is missing, instruct the user to create one in the OpenAI platform UI and export it in their shell.
- Never ask the user to paste the full key in chat.

## Skill path (set once)

```bash
export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
export TRANSCRIBE_CLI="$CODEX_HOME/skills/transcribe/scripts/transcribe_diarize.py"
```

User-scoped skills install under `$CODEX_HOME/skills` (default: `~/.codex/skills`).

## CLI quick start
Single file (fast text default):
```
python3 "$TRANSCRIBE_CLI" \
  path/to/audio.wav \
  --out transcript.txt
```

Diarization with known speakers (up to 4):
```
python3 "$TRANSCRIBE_CLI" \
  meeting.m4a \
  --model gpt-4o-transcribe-diarize \
  --known-speaker "Alice=refs/alice.wav" \
  --known-speaker "Bob=refs/bob.wav" \
  --response-format diarized_json \
  --out-dir output/transcribe/meeting
```

Plain text output (explicit):
```
python3 "$TRANSCRIBE_CLI" \
  interview.mp3 \
  --response-format text \
  --out interview.txt
```

## Reference map
- `references/api.md`: supported formats, limits, response formats, and known-speaker notes.

## Quality Checklist

Before returning a transcript:

- Confirm whether the user needs verbatim text, cleaned notes, speaker labels, timestamps, or action items.
- Preserve uncertainty markers for unclear words instead of silently guessing.
- Use diarization only when speaker separation matters; otherwise prefer the simpler text path.
- Keep private audio local except for the transcription API call required by the task.
- For long recordings, segment outputs by time or topic so the transcript remains navigable.
- If known-speaker samples are used, label them by role or name exactly as provided by the user.

## Output Options

Choose the output format based on the task:

- `text`: fastest path for simple transcripts.
- `diarized_json`: structured speaker turns for meetings and interviews.
- `markdown`: readable transcript with headings, timestamps, and notes.

Always state the model used, whether diarization was enabled, and any audio-quality limitations that affect confidence.

Files in this skill

  • LICENSE.txt10.5 KB
  • SKILL.md4 KB
  • agents/openai.yaml414 B
  • assets/transcribe-small.svg750 B
  • assets/transcribe.png1.3 KB
  • references/api.md457 B
  • scripts/transcribe_diarize.py8.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…