Skip to content
Back to skills

Gtars

ASecurity

Inspects and plans work with Gtars, the Rust/Python/CLI toolkit for genomic intervals: BED RegionSet set algebra (reduce, setdiff, intersect, closest, cluster, gaps), overlap counts, consensus peak sets, coverage, tokenizers, fragment processing, and refget/BEDbase stores. Use when merging or comparing BED interval sets, building a consensus or universe from peaks, tokenizing regions for ML models, working with fragment files, or planning refget sequence-collection and BEDbase cache use. Use ...

  • 7 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 4, 2026
researchpythonrustgobashtestinggitapidocumentation

Works with

  • cli
  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 15 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add KalarisLabs/research-agent-skills --skill gtars --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Gtars?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Gtars
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/kalarislabs-gtars/badge)](https://www.skillsdirectory.com/skills/kalarislabs-gtars)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: gtars
description: 'Inspects and plans work with Gtars, the Rust/Python/CLI toolkit for genomic intervals: BED RegionSet set algebra (reduce, setdiff, intersect, closest, cluster, gaps), overlap counts, consensus peak sets, coverage, tokenizers, fragment processing, and refget/BEDbase stores. Use when merging or comparing BED interval sets, building a consensus or universe from peaks, tokenizing regions for ML models, working with fragment files, or planning refget sequence-collection and BEDbase cache use. Use when checking BED coordinates, contigs, or strand before running gtars. Not for general BAM/VCF processing; use samtools or bedtools for that.'
license: MIT
compatibility: Python bindings require Python 3.10+ and gtars 0.9.2. The Rust meta-crate and gtars-cli are 0.9.0 and require a Rust toolchain supporting Edition 2024; upstream declares no rust-version. Bundled audit CLIs use only Python 3.10+ standard library and are local/network-free. Remote constructors, pretrained tokenizers, refget, and BEDbase caching require explicit network and storage approval.
allowed-tools: Read Write Edit Bash Glob
metadata:
  version: '1.3'
  category: life-sciences
  maintainer: Kalaris Labs
---

# Gtars

Gtars provides native Rust implementations, Python bindings, and a feature-gated
`gtars` binary for genomic interval and reference-sequence work. Start with the
bundled local inspectors; call upstream code only after the data contract,
provenance, resource bounds, and side effects are explicit.

## Verified snapshot (2026-07-23)

- Python: [`gtars==0.9.2`](https://pypi.org/project/gtars/), released
  2026-06-17, `Requires-Python >=3.10`.
- Rust meta-crate: [`gtars=0.9.0`](https://crates.io/crates/gtars), released
  2026-06-15. Its default feature set is empty.
- CLI crate/binary: [`gtars-cli=0.9.0`](https://crates.io/crates/gtars-cli);
  the installed binary is named `gtars`.
- Direct refget crate: [`gtars-refget=0.9.1`](https://crates.io/crates/gtars-refget),
  released 2026-06-17. `gtars=0.9.0` itself pins its component release set, which
  includes refget 0.9.0.
- Upstream intentionally versions workspace crates, Python bindings, and CLI
  independently. Do not assume matching numbers mean matching artifacts.
- The published docs changelog stops at 0.5.1. API examples here were checked
  against the 0.9.2 Python stubs/runtime and the `v0.9.0` CLI/Rust source.

The `license: MIT` field covers this skill. Published `gtars` crates declare MIT,
while the GitHub repository currently displays BSD-2-Clause at the root; verify
the exact artifact's license before redistribution.

## Native-code trust gate and exact pins

The Python wheel contains a PyO3 native extension. Cargo installation compiles a
native binary and can run dependency build scripts. Treat either path as code
execution:

1. Confirm the official PyPI/crates.io/GitHub owner and immutable version.
2. Review filenames, platform tags, release provenance, license, and SHA-256.
   GitHub's v0.9.0 binary release includes per-archive `.sha256` sidecars.
3. Never run an untrusted prebuilt binary, wheel, source tree, Cargo build script,
   or archive installer. Use isolation and CPU/RAM/disk/time limits.
4. Keep a lockfile and artifact hashes with the analysis manifest.

After that review, create an isolated Python environment:

```bash
uv venv --python 3.11 .venv-gtars
uv pip install --dry-run --python .venv-gtars/bin/python "gtars==0.9.2"
uv pip install --python .venv-gtars/bin/python "gtars==0.9.2"
.venv-gtars/bin/python -c \
  "import gtars; assert gtars.__version__ == '0.9.2'; print(gtars.__version__)"
```

For the reviewed CLI source release:

```bash
cargo install gtars-cli --version 0.9.0 --locked
gtars --version
gtars --help
```

For a Rust project, pin the wrapper exactly and enable only required features:

```toml
[dependencies]
gtars = { version = "=0.9.0", default-features = false, features = [
  "core", "overlaprs", "uniwig", "tokenizers", "refget"
] }
```

Use `gtars-refget = "=0.9.1"` directly only when the newer direct component API is
required and compatibility has been tested. Do not replace these pins with a Git
branch or an unreviewed release.

## Genomic data contract

Apply this contract before every operation:

1. **Coordinates:** BED intervals are 0-based and half-open: `[start, end)`.
   Require `0 <= start < end <= contig_length`. Gtars coordinates are `u32`, so
   reject values above `4,294,967,295`.
2. **Assembly:** record an assembly accession/version and the SHA-256 of the exact
   chromosome-sizes or refget sequence-collection metadata. Never infer assembly
   from filenames or `chr` prefixes.
3. **Contigs:** compare names exactly. `1` and `chr1`, alternate loci, decoys, and
   mitochondrial aliases are not interchangeable. Rename or liftover only as a
   separately reviewed transformation.
4. **Sorting:** preserve the original file, then sort a copy by chromosome-sizes
   order and numeric start/end when the operation requires it. Python
   `RegionSet(path)` currently sorts lexicographically by contig and start while
   loading; do not rely on original row order afterward.
5. **Strand:** BED6 uses `+`, `-`, or `.`. `Region.rest` retains trailing BED
   fields, but a file-backed Python `RegionSet` currently initializes its separate
   `strands` vector to `*`. Several set operations drop strand. Preserve and
   validate strand externally when it is scientifically meaningful.
6. **Duplicates/adjacency:** choose policies explicitly. `reduce()` and consensus
   merge overlapping **and adjacent** intervals; ordinary half-open overlap does
   not treat `[0,10)` and `[10,20)` as overlapping.

Run the local validator first:

```bash
python3 -B scripts/bed_validator.py \
  --input data.bed.gz \
  --assembly GRCh38.p14 \
  --chrom-sizes GRCh38.p14.chrom.sizes \
  --require-sorted
```

## Safe local workflow

1. Inventory local files, checksums, assembly, contig dictionary, coordinate
   system, strand policy, patient/replicate groups, and intended outputs.
2. Validate BED/fragments and estimate work. Pilot a small synthetic file.
3. Choose Python, CLI, or Rust from the documented surface; do not translate API
   names by guesswork.
4. Set hard limits for input bytes/records/files, threads/jobs, memory, temporary
   disk, output size, and wall time.
5. Run in a dedicated output directory. Refuse collisions unless overwrite was
   explicitly approved.
6. Revalidate output sorting, bounds, row counts, checksums, and provenance.

## Current Python core

Imports are from submodules, not the `gtars` top level:

```python
from gtars.models import Region, RegionSet

query = RegionSet.from_regions(
    [
        Region(chr="chr1", start=100, end=200, rest=None),
        Region(chr="chr1", start=300, end=400, rest=None),
    ],
    strands=["+", "-"],
)
universe = RegionSet.from_vectors(
    ["chr1", "chr1"],
    [150, 500],
    [350, 600],
)

counts = query.count_overlaps(universe)       # one count per query region
flags = query.any_overlaps(universe)          # one bool per query region
indices = query.find_overlaps(universe)       # indices into universe
pieces = query.intersect_all(universe)        # all intersection fragments
fraction = query.coverage(universe)           # fraction of query bp covered
```

`RegionSet.sort()` mutates and returns `None`. Set algebra includes `reduce`,
`setdiff`, `pintersect` (pairs by index), `concat`, `union`, `jaccard`,
`coverage`, `overlap_coefficient`, `intersect_all`, `closest`, `cluster`, and
`gaps`. Read `references/python-api.md` before relying on ordering or strand.

Consensus is a Python binding in a different module:

```python
from gtars.genomic_distributions import consensus

rows = consensus([query, universe])
# rows: [{"chr": ..., "start": ..., "end": ..., "count": ...}, ...]
```

Signal-track generation is **not** exposed as `gtars.uniwig` in Python 0.9.2;
use the reviewed CLI or Rust API. `RegionSet.coverage()` is a base-pair set metric,
not a WIG/bigWig generator.

## Tokenizers, fragments, and reference stores

Use only local constructors by default:

```python
from gtars.models import RegionSet
from gtars.tokenizers import Tokenizer

tokenizer = Tokenizer.from_bed("reviewed-universe.bed")
regions = RegionSet("local-query.bed")
tokens = tokenizer.tokenize(regions)
encoding = tokenizer(regions)
ids = encoding["input_ids"]
```

`Tokenizer.from_pretrained(name)` contacts Hugging Face and writes its cache when
the argument is not an existing local directory; it exposes no revision or cache
argument. Obtain explicit approval, fetch an immutable revision through a reviewed
mechanism, verify checksums, then pass the local snapshot directory. See
`references/tokenizers.md`.

For refget, prefer `RefgetStore.in_memory()` or `RefgetStore.open_local(path)`.
`open_remote(cache_path, remote_url)` contacts a remote service, creates/uses a
local cache, and performs on-demand range reads. See `references/refget.md`.

## Network and cache gate

No download or cache write is implicit in this skill. Before any network-capable
upstream call:

- obtain explicit user approval for the exact host, endpoint, data, and cache;
- allowlist HTTPS hosts and reject unreviewed redirects;
- record immutable revision/identifier, retrieval time, expected SHA-256 and
  domain digest, assembly accession, size quota, and provenance;
- disclose sensitive BED coordinates, barcodes, sample labels, and reference
  choices that could leave the approved environment;
- validate downloaded content as untrusted before using it.

Important side effects:

- `RegionSet(path)` has HTTP support; a nonexistent local string may be treated as
  a URL. Check that the local path exists before construction.
- `Tokenizer.from_pretrained` may download `universe.bed.gz` into the Hugging Face
  cache.
- `RefgetStore.on_disk` creates/writes a store. `open_remote` loads remote metadata
  and enables persistence by default.
- `gtars bbcache` creates cache directories even when constructing the client.
  Cache/download commands use `BBCLIENT_CACHE` (default `~/.bbcache`) and
  `BEDBASE_API` (default `https://api.bedbase.org`).

## Sensitive metadata and leakage

Genomic intervals, rare loci, barcodes, sample names, phenotypes, and assembly
choices can be identifying. Keep full paths and raw coordinates out of logs;
default bundled reports redact paths and emit only counts/checksums.

Freeze splits by patient/donor first, then keep all technical and biological
replicates in the same split. Fit consensus sets, universes, tokenizers, scaling,
thresholds, and QC rules on training data only. Do not create a universe from all
samples and then split: that leaks validation/test locus support. Record excluded
samples and replicate aggregation separately.

## Bundled deterministic CLIs

All six helpers reject URLs, traversal, symlinks, and special files; apply byte,
record, file, coordinate, and worker caps; use no network or gtars import; and
write no output files. Plans contain fixed argv templates and never launch them.

```bash
python3 -B scripts/bed_validator.py --help
python3 -B scripts/execution_plan.py --help
python3 -B scripts/tokenizer_manifest.py --help
python3 -B scripts/refget_digest_plan.py --help
python3 -B scripts/coverage_preflight.py --help
python3 -B scripts/artifact_inspector.py --help
```

Run synthetic tests without bytecode:

```bash
PYTHONDONTWRITEBYTECODE=1 python3 -B -m unittest discover \
  -s tests/gtars -p 'test_*.py' -v
```

## Migration traps removed in 1.1

Do not use stale examples containing `gtars.RegionSet`,
`RegionSet.from_bed`, `TreeTokenizer`, `gtars.igd.build_index`,
`gtars.uniwig.coverage_from_bed`, `gtars.RefgetStore`, global
`set_option`/`set_log_level`, `parallel_apply`, or invented exception classes.
CLI forms such as `uniwig generate`, `igd build`, `scoring score`, and
`fragsplit cluster-split` are also stale for 0.9.0.

Upstream's published docs and stubs have some drift (for example the older
`GlobalRefgetStore` tutorial and incomplete 0.9.2 stubs). Prefer installed
signature smoke tests plus immutable tagged source when they conflict.

## Bundled references

These are the only six bundled references; all links are local and present:

- `references/python-api.md` — exact Python 0.9.2 imports and behavior
- `references/overlap.md` — overlap/count/set algebra and consensus semantics
- `references/coverage.md` — uniwig, bigWig, coverage, sorting, and resources
- `references/tokenizers.md` — tokenizer/universe and fragment compatibility
- `references/refget.md` — digests, stores, BEDbase, network/cache controls
- `references/cli.md` — CLI 0.9.0 commands, features, and migrations

## Agent operating procedure

1. **Check the environment.** Confirm tool versions, the reference genome/annotation build and the input formats (FASTQ, BAM, VCF, h5ad).
2. **Pin down the inputs.** Confirm formats, identifiers and parameters from the data or the user. Ask rather than guess any value that changes the result.
3. **Run a small version first.** Run the pipeline on a small subset (one sample, one chromosome, a few thousand cells) first.
4. **Execute the full task** using the instructions and references above.
5. **Validate the result.** Check QC metrics, sample identities, genome build consistency and batch effects before interpreting results.
6. **Report.** State what was run (versions, commands, parameters), what was checked, and what is still uncertain.

| If this happens | Do this |
|---|---|
| Genome builds or identifiers do not match between inputs | Stop and harmonize (liftover, ID mapping) before continuing. |
| A function, flag or endpoint in these instructions is missing in the installed version | Check the installed version's own documentation (`help()`, `--help`, official docs), adapt, and tell the user. Never invent an API. |
| A required input, identifier or parameter is ambiguous | Ask the user, or state the assumption explicitly before running. |

**Integrity rules**

- Never fabricate results, parameters, identifiers, citations or statistics. If something cannot be run or verified, say so plainly.
- Do not interpret biological significance beyond what the statistics support; report multiple-testing correction.
- Treat version-specific details here as possibly outdated: confirm them against the official documentation for the installed version.
- Ask before actions that cost money, consume shared GPUs or cloud quota, touch personal or patient data, or cannot be undone.

## Related skills

- `geniml`: Plans and audits local genomic-interval machine learning workflows with Geniml (0.8.4) and Gtars: validates BED files against chromosome si…
- `polars-bio`: Python library polars-bio for genomic interval operations and bioinformatics file I/O on Polars DataFrames, built on Arrow and DataFusion.
- `genomic-coordinates`: Convert genomic intervals between coordinate conventions, normalise and compare variant representations, and detect assembly or contig-nami…

Files in this skill

  • SKILL.md14.7 KB
  • references/cli.md9.7 KB
  • references/coverage.md7.5 KB
  • references/overlap.md7.3 KB
  • references/python-api.md9.4 KB
  • references/refget.md10.3 KB
  • references/tokenizers.md8 KB
  • scripts/__init__.py57 B
  • scripts/_common.py14.6 KB
  • scripts/artifact_inspector.py10.6 KB
  • scripts/bed_validator.py5.3 KB
  • scripts/coverage_preflight.py8.1 KB
  • scripts/execution_plan.py11.6 KB
  • scripts/refget_digest_plan.py10.7 KB
  • scripts/tokenizer_manifest.py8.4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…