Skip to content
Back to skills

Data And Features

ASecurity

"ASRT data configuration, datalist, dictionary, WAV, and speech

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
toolsapi

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 7 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill data-and-features --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data And Features?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Data And Features
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-data-and-features/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-data-and-features)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: data-and-features
description: "ASRT data configuration, datalist, dictionary, WAV, and speech
  feature extraction operating guidance."
metadata:
  disco-role: operating
disable-model-invocation: true
license: GPL 3.0
---

# data-and-features

Use this sub-skill when an ASRT task involves dataset configuration, pinyin dictionaries, WAV metadata, or speech-feature extraction before model training, evaluation, prediction, or serving.

## Owns

- `asrt_config.json` structure and dataset section semantics.
- `dict.txt` tab-separated pinyin-to-symbol rows and pinyin-to-index mapping.
- ASRT wav-list and syllable-label line schemas.
- `DataLoader` data assembly behavior and `utils.config` cache behavior.
- WAV readers and byte decoding used before feature extraction.
- MFCC, Logfbank, Spectrogram, and SpecAugment feature behavior, including the 16 kHz constraint for spectrogram-style inputs.

## Route away

- Training, evaluation, prediction scripts, acoustic-model class selection, model weights, CTC, and tensor shapes beyond data/feature compatibility: use `acoustic-models`.
- Pinyin sequence to Chinese text decoding or language-model internals: use `language-model`.
- HTTP/gRPC request payloads, server/client execution, and SDK calls: use `serving-clients`.
- Full dataset downloads and interactive download helpers: treat as reference-only, not as a bundled runtime workflow.

## Operating path

1. Read [references/data-and-config.md](references/data-and-config.md) before editing or validating config, dict, datalist, or syllable labels.
2. Run [scripts/validate_asrt_config.py](scripts/validate_asrt_config.py) for self-contained structural validation of config/list/dict files.
3. Read [references/audio-and-features.md](references/audio-and-features.md) before diagnosing WAV metadata, 16 kHz failures, spectrogram shape, MFCC/Logfbank behavior, or SpecAugment randomness.
4. Run [scripts/inspect_audio_features.py](scripts/inspect_audio_features.py) for safe WAV and feature-shape inspection without importing ASRT.
5. Use [references/api-reference.md](references/api-reference.md) for quick schema/API reminders and [references/troubleshooting.md](references/troubleshooting.md) for failure diagnosis.

## Safety and boundaries

The bundled scripts are read-only validators/inspectors. They do not download corpora, record microphone audio, call servers, train models, or require an ASRT checkout. Provide user-owned config, dict, datalist, label, or WAV paths explicitly.

Files in this skill

  • SKILL.md2.4 KB
  • references/api-reference.md4.3 KB
  • references/audio-and-features.md6.1 KB
  • references/data-and-config.md6.8 KB
  • references/troubleshooting.md6.3 KB
  • scripts/inspect_audio_features.py14 KB
  • scripts/validate_asrt_config.py14.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…