Skip to content
Back to skills

Tokenizers

ASecurity

Understand and work with tokenizers — BPE, WordPiece, Unigram, special tokens, and the practical effects of tokenization on cost and quality. Use when token counts, chunking, or model inputs misbehave.

  • 2 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 29, 2026
ai-agentsgodebuggingdocumentation

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill tokenizers --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Tokenizers?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Tokenizers
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-tokenizers/badge)](https://www.skillsdirectory.com/skills/aicodedecode-tokenizers)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: tokenizers
description: Understand and work with tokenizers — BPE, WordPiece, Unigram, special tokens, and the practical effects of tokenization on cost and quality. Use when token counts, chunking, or model inputs misbehave.
category: ai-research
---

# Tokenizers

Tokenizers turn text into the integer sequences models actually process. They're invisible until 
they cause problems: surprising costs, broken chunking, or corrupted inputs. Understanding them 
turns those mysteries into mechanics.

## Overview

Modern tokenizers learn subword vocabularies from data (BPE merges frequent pairs; WordPiece and 
Unigram have their own strategies), then split text into those pieces. Key consequences: token 
count ≠ word count (varies by language and content — code and non-English text tokenize less 
efficiently), special tokens mark boundaries and roles, and detokenization isn't always the exact 
inverse. Every model has its own tokenizer — using the wrong one corrupts inputs silently.

## When to use

- Estimating costs: tokens are the billing unit — count with the right tokenizer.
- Debugging "the model saw something different": encoding issues, truncation surprises.
- Designing chunking for RAG: chunk by tokens, not characters.
- Working with structured formats: JSON, code, and special tokens need care.

## Core concepts

- **Subword algorithms**: BPE (merge frequent pairs), WordPiece (likelihood-based merges), Unigram 
(probabilistic pruning). Different trade-offs; what matters is which one your model uses.
- **Vocabulary**: the learned piece inventory, typically 30k–200k entries. Larger vocabs tokenize 
more efficiently (fewer tokens per text) but need bigger embedding tables.
- **Special tokens**: BOS/EOS markers, padding, separators, and chat-template tokens (`<|user|>`, 
etc.). Chat templates are tokenizer-level — apply them via the tokenizer, not by hand.
- **Fertility**: tokens per word — varies wildly by language and domain. English prose ~1.3; code 
higher; some languages 2–3×. Budget multilingual contexts accordingly.
- **Encode/decode round-trip**: decoding isn't always exact (whitespace handling, byte fallbacks). 
Test round-trips for pipelines that reconstruct text.
- **Pre-tokenization**: how text is split before BPE (whitespace, punctuation rules). Explains many 
"why did it split there?" mysteries.

## Practical workflow

1. Identify the model's actual tokenizer — never assume; load it from the model's own artifacts.
2. For cost estimation: tokenize representative inputs with that tokenizer, not a generic counter.
3. For chunking: chunk on token boundaries with the model's tokenizer; account for special tokens 
added by chat templates.
4. When debugging inputs: decode what the model actually received — compare against what you 
intended.
5. Apply chat templates through the tokenizer's template support; verify the rendered format 
matches documentation.
6. For multilingual work: measure fertility per language; adjust context budgets per language.

```text
Tokenizer debugging checklist:
[ ] Using the model's own tokenizer (not a generic one)
[ ] Chat template applied via tokenizer, output verified
[ ] Token counts measured on real inputs per language
[ ] Special tokens accounted for in budgets
[ ] Round-trip tested for reconstruction pipelines
```

## Common pitfalls

- **Wrong tokenizer**: counting or chunking with a different tokenizer than the model uses. 
Silently wrong everywhere.
- **Hand-rolled chat templates**: formatting `<|user|>` blocks manually and getting subtle details 
wrong. Use the tokenizer's template.
- **Character-based chunking**: splitting mid-token or mid-word for RAG. Chunk by tokens with 
overlap.
- **Ignoring fertility**: budgeting English token counts for multilingual content. Measure per 
language.
- **Truncation surprises**: tokenizers truncate silently at limits. Check lengths; truncate 
deliberately with the right strategy (head, tail, or middle).
- **Byte-fallback blindness**: rare characters becoming multi-token byte sequences. Affects cost 
and sometimes quality for unusual scripts.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…