Skip to content
Back to skills

R Nlp

ASecurity

R natural language processing packages. Use for text mining, sentiment analysis, topic modeling, tokenization, and text vectorization.

  • 5 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added June 4, 2026
datagitapi

Works with

  • api

Security analysis

A100/100

Pro scans all 19 files and shows the line behind each finding

Scanned June 4, 2026

npx -y skills add LeoLin990405/r-analytics-skill --skill r-nlp --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of R Nlp?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for R Nlp
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/leolin990405-r-nlp/badge)](https://www.skillsdirectory.com/skills/leolin990405-r-nlp)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: r-nlp
description: R natural language processing packages. Use for text mining, sentiment analysis, topic modeling, tokenization, and text vectorization.
---

# R NLP Skill

## Sub-skills

| Sub-skill | Description |
|-----------|-------------|
| [r-nlp-text](r-nlp-text/SKILL.md) | tidytext, quanteda, text2vec |
| [r-nlp-topic](r-nlp-topic/SKILL.md) | LDA, STM, topic modeling |
| [r-nlp-sentiment](r-nlp-sentiment/SKILL.md) | sentimentr, syuzhet, lexicons |

Natural Language Processing and text mining in R.

## Core NLP Packages

| Package | Description |
|---------|-------------|
| **tidytext** ★ | Tidy text mining (tidyverse style) |
| **quanteda** ★ | Quantitative text analysis |
| **tm** | Comprehensive text mining framework |
| **text2vec** ★ | Fast vectorization & word embeddings |
| **NLP** | Basic NLP functions |
| **openNLP** | Apache OpenNLP interface |

## Text Processing

| Package | Description |
|---------|-------------|
| **stringr** | Consistent string manipulation |
| **stringi** | ICU-based string processing |
| **SnowballC** | Snowball stemmers |
| **koRpus** | Text analysis package |
| **utf8** | UTF-8 text handling |

## Sentiment Analysis

| Package | Description |
|---------|-------------|
| **syuzhet** | Sentiment extraction (3 dictionaries) |
| **sentimentr** | Sentence-level sentiment |
| **tidytext** | Sentiment lexicons (AFINN, Bing, NRC) |

## Topic Modeling

| Package | Description |
|---------|-------------|
| **topicmodels** | LDA and CTM topic models |
| **LDAvis** | Interactive topic model visualization |
| **stm** | Structural topic models |

## Other

| Package | Description |
|---------|-------------|
| **zipfR** | Word frequency distributions |
| **MonkeyLearn** | MonkeyLearn API interface |
| **corporaexplorer** | Dynamic text collection exploration |

## Quick Examples

```r
# tidytext workflow
library(tidytext)
library(dplyr)

# Tokenize
df %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words) %>%
  count(word, sort = TRUE)

# Sentiment analysis
df %>%
  unnest_tokens(word, text) %>%
  inner_join(get_sentiments("bing")) %>%
  count(sentiment)

# TF-IDF
df %>%
  unnest_tokens(word, text) %>%
  count(document, word) %>%
  bind_tf_idf(word, document, n)

# quanteda
library(quanteda)
corpus <- corpus(texts)
tokens <- tokens(corpus, remove_punct = TRUE)
dfm <- dfm(tokens) %>%
  dfm_remove(stopwords("en")) %>%
  dfm_trim(min_termfreq = 5)

# Topic modeling
library(topicmodels)
dtm <- cast_dtm(df, document, word, n)
lda <- LDA(dtm, k = 5, control = list(seed = 1234))
topics <- tidy(lda, matrix = "beta")

# text2vec word embeddings
library(text2vec)
it <- itoken(texts, tokenizer = word_tokenizer)
vocab <- create_vocabulary(it)
vectorizer <- vocab_vectorizer(vocab)
tcm <- create_tcm(it, vectorizer, skip_grams_window = 5)
glove <- GloVe$new(rank = 50)
word_vectors <- glove$fit_transform(tcm, n_iter = 10)
```

## Common Workflows

### Text Preprocessing
```r
library(tidytext)
library(dplyr)
library(stringr)

clean_text <- df %>%
  mutate(
    text = str_to_lower(text),
    text = str_remove_all(text, "[[:punct:]]"),
    text = str_remove_all(text, "[[:digit:]]"),
    text = str_squish(text)
  ) %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words) %>%
  mutate(word = SnowballC::wordStem(word))
```

### Document-Term Matrix
```r
# tidytext to DTM
dtm <- df %>%
  unnest_tokens(word, text) %>%
  count(doc_id, word) %>%
  cast_dtm(doc_id, word, n)

# quanteda DFM
dfm <- corpus(texts) %>%
  tokens() %>%
  dfm()
```

## Resources

- tidytext book: https://www.tidytextmining.com/
- quanteda: https://quanteda.io/
- text2vec: https://text2vec.org/
- tm: https://tm.r-forge.r-project.org/

Files in this skill

  • SKILL.md3.6 KB
  • r-nlp-sentiment/SKILL.md3.1 KB
  • r-nlp-sentiment/sentimentr/SKILL.md3.5 KB
  • r-nlp-sentiment/syuzhet/SKILL.md3.1 KB
  • r-nlp-text/SKILL.md2.3 KB
  • r-nlp-text/hunspell/SKILL.md1.7 KB
  • r-nlp-text/quanteda/SKILL.md1.6 KB
  • r-nlp-text/stringdist/SKILL.md2.1 KB
  • r-nlp-text/text2vec/SKILL.md1.5 KB
  • r-nlp-text/tidytext/SKILL.md3.3 KB
  • r-nlp-text/tm/SKILL.md1.5 KB
  • r-nlp-text/tokenizers/SKILL.md2 KB
  • r-nlp-topic/LDAvis/SKILL.md3.8 KB
  • r-nlp-topic/SKILL.md2.1 KB
  • r-nlp-topic/stm/SKILL.md3.6 KB
  • r-nlp-topic/topicmodels/SKILL.md3.7 KB
  • sub-skills/r-nlp-text/sub-skills/quanteda/SKILL.md1.6 KB
  • sub-skills/r-nlp-text/sub-skills/text2vec/SKILL.md1.5 KB
  • sub-skills/r-nlp-text/sub-skills/tm/SKILL.md1.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…