Back to skills
SKILL.md
Perplexity Filter
ASecurity困惑度过滤器。过滤以保留困惑度分数小于特定最大值的样本。 本SKILL使用依赖data_juicer,请在调用前安装好python环境并安装data_juicer,你可用以下指令进行安装: pip install py-data-juicer 当用户提到困惑度过滤、文本困惑度检测、语言模型分数过滤、文本质量过滤等需求时使用此skill。
- 539 stars
- 0 votes
- 0 copies
- 0 views
- Added September 11, 2026
Security analysis
92/100- Installs packages at runtime which could introduce malicious dependencies
Pro scans all 5 files and shows the line behind each finding
npx -y skills add cas-bigdatalab/piflow --skill perplexity_filter --agent claude-codeAre you the author of Perplexity Filter?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/cas-bigdatalab-perplexity-filter)---
name: perplexity_filter
description: |
困惑度过滤器。过滤以保留困惑度分数小于特定最大值的样本。
本SKILL使用依赖data_juicer,请在调用前安装好python环境并安装data_juicer,你可用以下指令进行安装:
pip install py-data-juicer
当用户提到困惑度过滤、文本困惑度检测、语言模型分数过滤、文本质量过滤等需求时使用此skill。
name_zh: 困惑度过滤器算子
input_params:
- name: input_path
type: string
required: true
description: 输入数据文件路径(JSON/JSONL格式)
- name: output_path
type: string
required: true
description: 输出数据文件路径(JSONL格式)
- name: lang
type: string
required: false
default: en
description: 语言代码(影响分词和语言模型)
- name: max_ppl
type: float
required: false
default: 1500
description: 最大困惑度阈值(低于此值保留)
- name: batch_size
type: int
required: false
default: 1
description: 批处理大小
- name: num_proc
type: int
required: false
default: 1
description: 并行处理的进程数
- name: text_key
type: string
required: false
default: text
description: 要操作的文本字段名
output_params:
- name: output_path
type: jsonl_file
description: 过滤后的JSONL文件,包含困惑度低于阈值的样本
tag: 数据筛选
publisher: COMMUNITY
---
## 功能概述
该算子使用语言模型(KenLM)计算文本的困惑度分数,过滤保留困惑度在指定范围内的样本。低困惑度表示文本流畅自然。
## 核心参数
| 参数 | 类型 | 必填 | 默认值 | 说明 |
|------|------|------|--------|------|
| input_path | string | 是 | - | 输入数据文件路径 (JSON/JSONL格式) |
| output_path | string | 是 | - | 输出数据文件路径 (JSONL格式) |
| lang | string | 否 | 'en' | 语言代码(影响分词和语言模型) |
| max_ppl | float | 否 | 1500 | 最大困惑度阈值(低于此值保留) |
| batch_size | int | 否 | 1 | 批处理大小 |
| num_proc | int | 否 | 1 | 并行处理的进程数 |
| text_key | string | 否 | text | 要操作的文本字段名 |
## 输入数据格式
输入文件应为 JSON 或 JSONL 格式,每行包含一个样本,样本需包含 `text_key` 指定的字段(默认 `text`):
```json
{"<text_key>": "这是一段需要检测困惑度的文本内容"}
```
## 输出数据格式
输出为 JSONL 格式,每行一个符合条件的样本,仅包含 text 字段。
## 使用示例
### 命令行调用
```bash
# 英语文本困惑度过滤
python scripts/run_perplexity_filter.py \
--input_path /path/to/input.jsonl \
--output_path /path/to/output.jsonl \
--lang en \
--max_ppl 900
# 中文文本困惑度过滤
python scripts/run_perplexity_filter.py \
--input_path /path/to/input.jsonl \
--output_path /path/to/output.jsonl \
--lang zh \
--max_ppl 1500
```
### 参数说明
- `--input_path`: 输入文件路径
- `--output_path`: 输出文件路径
- `--lang`: 语言代码(默认 en,支持 en, zh 等)
- `--max_ppl`: 最大困惑度阈值(默认1500)
- `--batch_size`: 批处理大小(默认1)
- `--num_proc`: 并行进程数,默认1
- `--text_key`: 要操作的文本字段名(默认text)
## 注意事项
1. 该算子使用 sentencepiece 进行分词,kenlm 计算困惑度
2. 低困惑度表示文本流畅自然,高困惑度可能表示文本混乱或随机
3. 支持的语言包括:en, zh 等
4. 处理大量数据时可适当增加 batch_size 和 num_proc 提高效率Files in this skill
- SKILL.md
- assets/icon.png
- scripts/example_input.json
- scripts/run_perplexity_filter.py
- skill.json
Attribution
Comments
Loading comments…