Skip to content
Back to skills

Sparklyr

ASecurity

R sparklyr package for Apache Spark. Use for distributed data processing with dplyr interface.

  • 5 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added June 4, 2026
datagosql

Security analysis

A100/100

Scanned June 4, 2026

npx -y skills add LeoLin990405/r-analytics-skill --skill sparklyr --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sparklyr?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Sparklyr
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/leolin990405-sparklyr/badge)](https://www.skillsdirectory.com/skills/leolin990405-sparklyr)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: sparklyr
description: R sparklyr package for Apache Spark. Use for distributed data processing with dplyr interface.
---

# sparklyr Package

R interface to Apache Spark.

## Connect

```r
library(sparklyr)
library(dplyr)

# Local mode
sc <- spark_connect(master = "local")

# Cluster
sc <- spark_connect(master = "yarn")
sc <- spark_connect(master = "spark://host:7077")
```

## Data Transfer

```r
# Copy to Spark
spark_df <- copy_to(sc, local_df, "my_table")

# Read from Spark
local_df <- collect(spark_df)

# Read files
spark_df <- spark_read_csv(sc, "data", "path/to/file.csv")
spark_df <- spark_read_parquet(sc, "data", "path/to/file.parquet")
spark_df <- spark_read_json(sc, "data", "path/to/file.json")
```

## dplyr Operations

```r
result <- spark_df %>%
  filter(year == 2023) %>%
  group_by(category) %>%
  summarise(
    count = n(),
    avg_value = mean(value)
  ) %>%
  arrange(desc(count))

# View SQL
show_query(result)

# Execute
collected <- collect(result)
```

## Machine Learning

```r
# Split data
partitions <- sdf_random_split(spark_df, training = 0.8, test = 0.2)

# Linear regression
model <- partitions$training %>%
  ml_linear_regression(y ~ x1 + x2)

# Predictions
predictions <- ml_predict(model, partitions$test)

# Random forest
model <- partitions$training %>%
  ml_random_forest(y ~ ., type = "classification")
```

## Write Data

```r
spark_write_csv(spark_df, "output.csv")
spark_write_parquet(spark_df, "output.parquet")
```

## Disconnect

```r
spark_disconnect(sc)
```

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…