Skip to content
Back to skills

R Nlp Text

ASecurity

R text mining with tidytext, tm, quanteda. Use for tokenization, TF-IDF, document-term matrices.

  • 5 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added June 4, 2026
data

Security analysis

A100/100

Pro scans all 8 files and shows the line behind each finding

Scanned June 4, 2026

npx -y skills add LeoLin990405/r-analytics-skill --skill r-nlp-text --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of R Nlp Text?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for R Nlp Text
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/leolin990405-r-nlp-text/badge)](https://www.skillsdirectory.com/skills/leolin990405-r-nlp-text)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: r-nlp-text
description: R text mining with tidytext, tm, quanteda. Use for tokenization, TF-IDF, document-term matrices.
---

# R Text Mining

Text processing and analysis.

## tidytext

```r
library(tidytext)
library(dplyr)

# Tokenize
df %>% unnest_tokens(word, text)

# Remove stop words
df %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words)

# Word counts
df %>%
  unnest_tokens(word, text) %>%
  count(word, sort = TRUE)

# TF-IDF
df %>%
  unnest_tokens(word, text) %>%
  count(document, word) %>%
  bind_tf_idf(word, document, n)

# N-grams
df %>% unnest_tokens(bigram, text, token = "ngrams", n = 2)

# Cast to DTM
dtm <- df %>%
  unnest_tokens(word, text) %>%
  count(document, word) %>%
  cast_dtm(document, word, n)
```

## quanteda

```r
library(quanteda)

# Create corpus
corpus <- corpus(texts)

# Tokenize
tokens <- tokens(corpus, remove_punct = TRUE, remove_numbers = TRUE)
tokens <- tokens_tolower(tokens)
tokens <- tokens_remove(tokens, stopwords("en"))
tokens <- tokens_wordstem(tokens)

# Document-feature matrix
dfm <- dfm(tokens)
dfm <- dfm_trim(dfm, min_termfreq = 5, min_docfreq = 2)

# TF-IDF weighting
dfm_tfidf <- dfm_tfidf(dfm)

# Top features
topfeatures(dfm, 20)

# Keyword in context
kwic(tokens, pattern = "economy", window = 5)
```

## tm

```r
library(tm)

# Create corpus
corpus <- Corpus(VectorSource(texts))

# Preprocessing
corpus <- tm_map(corpus, content_transformer(tolower))
corpus <- tm_map(corpus, removePunctuation)
corpus <- tm_map(corpus, removeNumbers)
corpus <- tm_map(corpus, removeWords, stopwords("english"))
corpus <- tm_map(corpus, stemDocument)
corpus <- tm_map(corpus, stripWhitespace)

# Document-term matrix
dtm <- DocumentTermMatrix(corpus)
dtm <- removeSparseTerms(dtm, 0.99)

# Term-document matrix
tdm <- TermDocumentMatrix(corpus)

# Find frequent terms
findFreqTerms(dtm, lowfreq = 10)

# Find associations
findAssocs(dtm, "economy", corlimit = 0.3)
```

## text2vec

```r
library(text2vec)

# Iterator
it <- itoken(texts, tokenizer = word_tokenizer, progressbar = FALSE)

# Vocabulary
vocab <- create_vocabulary(it)
vocab <- prune_vocabulary(vocab, term_count_min = 5)

# Vectorizer
vectorizer <- vocab_vectorizer(vocab)

# DTM
dtm <- create_dtm(it, vectorizer)

# TF-IDF
tfidf <- TfIdf$new()
dtm_tfidf <- fit_transform(dtm, tfidf)
```

Files in this skill

  • SKILL.md2.3 KB
  • hunspell/SKILL.md1.7 KB
  • quanteda/SKILL.md1.6 KB
  • stringdist/SKILL.md2.1 KB
  • text2vec/SKILL.md1.5 KB
  • tidytext/SKILL.md3.3 KB
  • tm/SKILL.md1.5 KB
  • tokenizers/SKILL.md2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…