Skip to content
Back to skills

Tokenizers

ASecurity

R tokenizers package for text tokenization. Use for fast, consistent tokenization of text.

  • 5 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added June 4, 2026
data

Security analysis

A100/100

Scanned June 4, 2026

npx -y skills add LeoLin990405/r-analytics-skill --skill tokenizers --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Tokenizers?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Tokenizers
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/leolin990405-tokenizers/badge)](https://www.skillsdirectory.com/skills/leolin990405-tokenizers)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: tokenizers
description: R tokenizers package for text tokenization. Use for fast, consistent tokenization of text.
---

# tokenizers

Fast, consistent tokenization of natural language text.

## Word Tokenization

```r
library(tokenizers)

# Tokenize into words
tokenize_words("This is a test sentence.")

# Multiple texts
texts <- c("First sentence.", "Second sentence.")
tokenize_words(texts)
```

## Options

```r
# Lowercase
tokenize_words(text, lowercase = TRUE)

# Keep punctuation
tokenize_words(text, strip_punct = FALSE)

# Keep numbers
tokenize_words(text, strip_numeric = FALSE)

# Stopwords
tokenize_words(text, stopwords = stopwords::stopwords("en"))
```

## Sentence Tokenization

```r
# Tokenize into sentences
tokenize_sentences("First sentence. Second sentence!")

# With abbreviations
tokenize_sentences(text, strip_punct = FALSE)
```

## Character Tokenization

```r
# Single characters
tokenize_characters("hello")

# Character shingles
tokenize_character_shingles("hello", n = 3)
```

## N-grams

```r
# Word n-grams
tokenize_ngrams("This is a test", n = 2)

# Range of n-grams
tokenize_ngrams("This is a test", n = 2, n_min = 1)

# Skip-grams
tokenize_skip_ngrams("This is a test", n = 2, k = 1)
```

## Paragraph Tokenization

```r
# Split by paragraphs
text <- "First paragraph.\n\nSecond paragraph."
tokenize_paragraphs(text)
```

## Line Tokenization

```r
# Split by lines
tokenize_lines("Line 1\nLine 2\nLine 3")
```

## Regex Tokenization

```r
# Custom pattern
tokenize_regex(text, pattern = "\\s+")
```

## Word Stems

```r
# Tokenize and stem
tokenize_word_stems("running cats jumping")

# With language
tokenize_word_stems(text, language = "english")
```

## PTB Tokenization

```r
# Penn Treebank style
tokenize_ptb("It's a test.")
```

## Tweet Tokenization

```r
# Twitter-aware tokenization
tokenize_tweets("Hello @user! Check out #rstats http://example.com")
```

## Count Tokens

```r
# Count words
count_words("This is a test sentence.")

# Count sentences
count_sentences("First. Second. Third.")

# Count characters
count_characters("hello")
```

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…