Skip to content
Back to skills

Sequence Architecture Picker

ASecurity

Pick sequence architecture (RNN, transformer, SSM, hybrid) given length, throughput, and training budget. Use when you need help with sequence architecture picker.

  • 8 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 8, 2026
developmentgo

Security analysis

A100/100

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill sequence-architecture-picker --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sequence Architecture Picker?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Sequence Architecture Picker
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-sequence-architecture-picker/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-sequence-architecture-picker)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: sequence-architecture-picker
description: Pick sequence architecture (RNN, transformer, SSM, hybrid) given length, throughput, and training budget. Use when you need help with sequence architecture picker.
license: CC-BY-NC-SA-4.0
phase: 7
lesson: 1
metadata:
  version: 1.0.0
  tags: [transformers, architecture, rnn, ssm]
---

Given a sequence problem (max length, batch shape, training tokens budgeted, inference latency target, device class), output:

1. Primary architecture. One of: transformer, state-space model (Mamba/RWKV), hybrid SSM+attention, RNN. One-sentence reason tied to the dominant constraint.
2. Context length strategy. If transformer: full attention cutoff, sliding window size, RoPE scaling factor. If SSM: scan chunk size. If RNN: hidden width.
3. Training FLOP profile. Approximate FLOPs per token from architecture + context; note whether the spec fits the compute budget.
4. Inference memory profile. KV cache for transformers, state size for SSMs, per-token memory for RNNs. Flag if the target device can hold a single batch of 1.
5. Risk note. One specific failure mode that this choice is known to have at the scale of the spec (e.g. transformer OOM at 64K context on a 24GB GPU without Flash Attention).

Refuse to recommend a pure RNN for any training run above 1B tokens without explicitly stating the gradient-flow and parallelism penalties. Refuse to recommend a full-attention transformer for >64K context without stating the `O(N^2)` memory cost. Refuse to recommend a brand-new architecture (published <12 months ago) for production without a named fallback.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…