Skip to content
Back to skills

Assemblyai

BSecurity

AssemblyAI is a hosted speech-to-text API that transcribes audio and video files or live streams and adds speaker labels, sentiment, entity detection, PII redaction and LLM analysis of the transcript. Use when a user asks to transcribe a recording, label who said what, analyze call sentiment, redact personal data, stream live transcription, or summarize and question a transcript with LLM Gateway (the replacement for LeMUR).

  • 142 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added May 27, 2026
data-aipythongobashnoderailskubernetesapi

Works with

  • cursor
  • terminal
  • cli
  • api

Security analysis

B84/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add TerminalSkills/skills --skill assemblyai --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Assemblyai?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Assemblyai
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-assemblyai/badge)](https://www.skillsdirectory.com/skills/terminalskills-assemblyai)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: assemblyai
description: >-
  AssemblyAI is a hosted speech-to-text API that transcribes audio and video
  files or live streams and adds speaker labels, sentiment, entity detection,
  PII redaction and LLM analysis of the transcript. Use when a user asks to
  transcribe a recording, label who said what, analyze call sentiment, redact
  personal data, stream live transcription, or summarize and question a
  transcript with LLM Gateway (the replacement for LeMUR).
license: Apache-2.0
compatibility: "Python 3.8+ (pip install assemblyai, SDK 1.x) or Node.js 18+ (npm install assemblyai). Requires an AssemblyAI API key."
metadata:
  author: terminal-skills
  version: "1.1.0"
  category: data-ai
  tags: ["assemblyai", "transcription", "speech-recognition", "audio-ai", "diarization"]
  use-cases:
    - "Transcribe a podcast episode with speaker labels and generate show notes"
    - "Run sentiment analysis on customer support call recordings"
    - "Extract key insights and action items from meeting recordings using LLM Gateway"
  agents: [claude-code, openai-codex, gemini-cli, cursor]
---

# AssemblyAI

## Overview

AssemblyAI turns audio and video into text over a REST API and a streaming WebSocket, with official Python and Node.js SDKs. The pieces you combine:

- **Pre-recorded transcription**: submit a URL or file, the SDK polls until the job is `completed` or `error`. Models are chosen with `speech_models` (default `["universal-3-5-pro", "universal-2"]`, an ordered fallback list).
- **Speech Understanding and Guardrails**: options on the same request — speaker labels, sentiment, entities, key phrases, topics, content moderation, PII redaction.
- **LLM Gateway**: an OpenAI-style chat completions endpoint for summaries, chapters and questions about a transcript. It replaces LeMUR, which was shut down on 2026-03-31.
- **Real-time STT**: a WebSocket session (`/v3/ws`) that returns `Turn` events while audio is being spoken.

This skill targets Python SDK 1.x (1.6.1 at the time of writing). Version 1.0 removed every `aai.Lemur*` class and the microphone helper in `assemblyai.extras`; the old `aai.RealtimeTranscriber` (streaming v2) is gone as well.

## Instructions

### Step 1: Install and authenticate

```bash
pip install -U assemblyai
read -rs ASSEMBLYAI_API_KEY && export ASSEMBLYAI_API_KEY   # paste the key from the dashboard
```

```python
import os
import assemblyai as aai

aai.settings.api_key = os.environ["ASSEMBLYAI_API_KEY"]
# EU data residency: aai.settings.base_url = "https://api.eu.assemblyai.com"
```

### Step 2: Transcribe a file or URL

```python
config = aai.TranscriptionConfig(speech_models=["universal-3-5-pro", "universal-2"])
transcriber = aai.Transcriber(config=config)

# A public URL, a local path, a pathlib.Path, bytes or an open binary file all work.
transcript = transcriber.transcribe("https://assembly.ai/wildfires.mp3", poll_timeout=600)

if transcript.status == aai.TranscriptStatus.error:
    raise RuntimeError(f"Transcription failed: {transcript.error}")

print(transcript.id)
print(transcript.text[:300])
```

A failed job is **returned, not raised** — always check `status` before reading `text`. `poll_timeout` (seconds) raises `aai.TranscriptError` if the job is still running; fetch it later with `aai.Transcript.get_by_id(transcript_id)`. Use `transcriber.submit(...)` plus `webhook_url=` in the config to skip polling.

### Step 3: Speaker labels and Speech Understanding

```python
config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro", "universal-2"],
    speaker_labels=True,        # who said what -> transcript.utterances
    sentiment_analysis=True,    # POSITIVE / NEUTRAL / NEGATIVE per sentence
    entity_detection=True,      # people, places, organizations
    auto_highlights=True,       # key phrases
    iab_categories=True,        # topics -> transcript.iab_categories.summary
    content_safety=True,        # content moderation
    language_detection=True,
)
t = aai.Transcriber().transcribe("https://assembly.ai/wildfires.mp3", config)
if t.status == aai.TranscriptStatus.error:
    raise RuntimeError(t.error)

for utt in t.utterances:
    print(f"[{utt.start // 1000:>5}s] Speaker {utt.speaker}: {utt.text}")

for s in t.sentiment_analysis[:5]:
    print(s.sentiment.value, s.speaker, s.text[:80])

for entity in t.entities:
    print(entity.entity_type.value, entity.text)

for phrase in t.auto_highlights.results[:10]:
    print(phrase.rank, phrase.count, phrase.text)

for result in t.content_safety.results:
    for label in result.labels:
        print(label.label.value, f"{label.confidence:.2f}", result.text[:60])
```

If you know the number of speakers, pass `speakers_expected=2`; otherwise leave it out. Timestamps (`start`, `end`) are milliseconds.

### Step 4: Redact personal data

```python
config = aai.TranscriptionConfig(speaker_labels=True).set_redact_pii(
    policies=[
        aai.PIIRedactionPolicy.person_name,
        aai.PIIRedactionPolicy.phone_number,
        aai.PIIRedactionPolicy.email_address,
        aai.PIIRedactionPolicy.credit_card_number,
    ],
    substitution=aai.PIISubstitutionPolicy.entity_name,   # "[PERSON_NAME]" instead of "####"
    redact_audio=True,                                    # also produce a beeped audio file
)
t = aai.Transcriber().transcribe("./calls/support-2026-09-14.mp3", config)
print(t.text)
print(t.get_redacted_audio_url())    # link is valid for 24 hours
```

### Step 5: Summaries, chapters and questions with LLM Gateway

The transcript parameters `auto_chapters`, `summarization`, `summary_model` and `summary_type` are deprecated. Send the transcript to LLM Gateway instead. The literal tag `{{ transcript }}` is replaced server-side with the text of `transcript_id`:

```python
gateway = aai.LLMGateway()

completion = gateway.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{
        "role": "user",
        "content": "List the decisions and action items from this meeting, "
                   "one bullet each, with the owner's name.\n\n{{ transcript }}",
    }],
    transcript_id=transcript.id,
    max_tokens=1000,
)
print(completion.choices[0].message.content)

for model in gateway.models.list().data:    # exact, current model ids
    print(model.id)
```

Model ids are exact strings (`claude-sonnet-4-6`, `gpt-5-mini`, `gemini-2.5-flash`); list them rather than guessing. Without the SDK, POST the same JSON to `https://llm-gateway.assemblyai.com/v1/chat/completions` with the header `Authorization: $ASSEMBLYAI_API_KEY` (no `Bearer`). Errors raise `aai.LLMGatewayError` with `.status_code` and `.request_id`.

### Step 6: Real-time transcription

The SDK no longer captures the microphone; read 16-bit mono PCM yourself and pass the chunks to `stream()`. This example uses `pip install pyaudio`, which compiles against PortAudio on Linux and macOS: install `portaudio19-dev` (Debian/Ubuntu) or `brew install portaudio` first.

```python
import os
import pyaudio
from assemblyai.streaming.v3 import (
    RealTimeError, RealTimeEvents, RealTimeParameters, RealTimeTranscriber,
    TerminationEvent, TurnEvent,
)

RATE, FRAMES = 16_000, 1_600      # 100 ms chunks; the API accepts 50-1000 ms

def microphone():
    audio = pyaudio.PyAudio()
    mic = audio.open(format=pyaudio.paInt16, channels=1, rate=RATE,
                     input=True, frames_per_buffer=FRAMES)
    try:
        while True:
            yield mic.read(FRAMES, exception_on_overflow=False)
    finally:
        mic.stop_stream(); mic.close(); audio.terminate()

def on_turn(client, event: TurnEvent):
    if event.end_of_turn:
        print(f"[final] {event.transcript}")
    else:
        print(f"\r{event.transcript}", end="")

def on_terminated(client, event: TerminationEvent):
    print(f"\nSession closed after {event.audio_duration_seconds}s of audio")

def on_error(client, error: RealTimeError):
    print(f"Streaming error {error.code}: {error}")

client = RealTimeTranscriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])
client.on(RealTimeEvents.Turn, on_turn)
client.on(RealTimeEvents.Termination, on_terminated)
client.on(RealTimeEvents.Error, on_error)
client.connect(RealTimeParameters(sample_rate=RATE, speech_model="universal-3-6-pro"))
try:
    client.stream(microphone())   # Ctrl+C to stop
except KeyboardInterrupt:
    pass
finally:
    client.disconnect(terminate=True)
```

Streaming models: `universal-3-6-pro` (default), `universal-3-5-pro`, `universal-streaming-english`, `universal-streaming-multilingual`. The former `StreamingClient`/`StreamingParameters`/`StreamingEvents` names still import as aliases of the `RealTime*` classes.

### Feature reference

| Feature | Config | Read the result from |
|---------|--------|----------------------|
| Speaker labels | `speaker_labels=True` | `transcript.utterances` |
| Sentiment analysis | `sentiment_analysis=True` | `transcript.sentiment_analysis` |
| Entity detection | `entity_detection=True` | `transcript.entities` |
| Key phrases | `auto_highlights=True` | `transcript.auto_highlights.results` |
| Topic detection | `iab_categories=True` | `transcript.iab_categories` |
| Content moderation | `content_safety=True` | `transcript.content_safety` |
| Language detection | `language_detection=True` | `transcript.language_code` |
| PII redaction | `.set_redact_pii(policies=[...])` | `transcript.text`, `get_redacted_audio_url()` |
| Domain vocabulary | `keyterms_prompt=["Kubernetes", "Grafana"]` | better spelling in `text` |
| Subtitles | — | `transcript.export_subtitles_srt()`, `export_subtitles_vtt()` |
| Chapters, summaries, Q&A | LLM Gateway (Step 5) | `completion.choices[0].message.content` |

## Examples

### Example 1: Podcast episode to speaker-labeled transcript and show notes

**User prompt:** "Transcribe episode 42 of our podcast with speaker labels and write show notes with the key takeaways."

```python
import os
import assemblyai as aai

aai.settings.api_key = os.environ["ASSEMBLYAI_API_KEY"]

config = aai.TranscriptionConfig(speaker_labels=True, speakers_expected=2)
episode = aai.Transcriber().transcribe("./episodes/ep42-observability.mp3", config, poll_timeout=1800)
if episode.status == aai.TranscriptStatus.error:
    raise RuntimeError(episode.error)

with open("ep42-transcript.txt", "w") as out:
    for utt in episode.utterances:
        out.write(f"Speaker {utt.speaker}: {utt.text}\n\n")
with open("ep42.srt", "w") as out:
    out.write(episode.export_subtitles_srt())

notes = aai.LLMGateway().chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content":
        "Write podcast show notes in Markdown: a one-paragraph summary, five key "
        "takeaways as bullets, and the tools mentioned.\n\n{{ transcript }}"}],
    transcript_id=episode.id,
    max_tokens=1500,
)
print(notes.choices[0].message.content)
```

Result: `ep42-transcript.txt` with turns labeled `Speaker A:` and `Speaker B:`, an `ep42.srt` subtitle file, and Markdown show notes printed to the terminal.

### Example 2: Sentiment report for a support call with personal data removed

**User prompt:** "Analyze yesterday's support call: who was negative and when? Customer names and card numbers must not appear in the output."

```python
import os
from collections import Counter
import assemblyai as aai

aai.settings.api_key = os.environ["ASSEMBLYAI_API_KEY"]

config = aai.TranscriptionConfig(
    speaker_labels=True, speakers_expected=2, sentiment_analysis=True,
).set_redact_pii(
    policies=[aai.PIIRedactionPolicy.person_name, aai.PIIRedactionPolicy.credit_card_number],
    substitution=aai.PIISubstitutionPolicy.entity_name,
)
call = aai.Transcriber().transcribe("./calls/support-2026-09-30-1412.wav", config)
if call.status == aai.TranscriptStatus.error:
    raise RuntimeError(call.error)

totals = Counter((s.speaker, s.sentiment.value) for s in call.sentiment_analysis)
for (speaker, sentiment), count in sorted(totals.items()):
    print(f"Speaker {speaker}: {sentiment} x{count}")

for s in call.sentiment_analysis:
    if s.sentiment == aai.SentimentType.negative:
        print(f"{s.start // 60000:02d}:{s.start // 1000 % 60:02d} Speaker {s.speaker}: {s.text}")
```

Result: a per-speaker count such as `Speaker B: NEGATIVE x7`, followed by each negative sentence with its timestamp, where every name reads `[PERSON_NAME]` and card numbers are replaced the same way.

## Guidelines

- **Do not write LeMUR code.** The LeMUR API was shut down on 2026-03-31 and SDK 1.0 removed `transcript.lemur.*`, `aai.Lemur` and `aai.LemurModel`. Use LLM Gateway.
- **Do not use `aai.RealtimeTranscriber`** or `wss://api.assemblyai.com/v2/realtime/ws`; that is streaming v2. Use `assemblyai.streaming.v3`.
- **Streaming is billed for the time the session is open**, not for audio sent. Always call `disconnect(terminate=True)`; an abandoned session runs until the 3-hour cap.
- **Keep the API key on the server.** For browsers and mobile apps, mint a short-lived token with `RealTimeTranscriber(api_key=...).create_temporary_token(expires_in_seconds=60)` and send only the token.
- **REST authentication is the raw key** in the `Authorization` header, without a `Bearer` prefix.
- **Limits:** up to 5 GB and 10 hours per file (2.2 GB for a local upload); redacted audio needs a source file under 1 GB. Free accounts run 5 transcription jobs in parallel, paid accounts 200 or more; extra jobs are queued, not rejected.
- **Cost:** new accounts get $50 of free credit for transcription features; LLM Gateway usage is not covered by it and is billed per token.
- **Audio leaves your machine.** For recordings that may not be sent to a third party, use a local model instead, and use `set_redact_pii` when transcripts are stored or shared.
- **Parameters change between models.** Check a flag against https://www.assemblyai.com/docs/llms.txt before relying on it — for example `prompt` and `keyterms_prompt` behave differently on Universal-3.5 Pro and Universal-2.

Files in this skill

  • SKILL.md10.1 KB
  • _scores.json1.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…