Skip to content
Back to skills

Adaflash Adaptive Speculative Decoding

ASecurity

Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

  • 3 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 11, 2026
developmentperformance

Security analysis

A100/100

Scanned September 11, 2026

npx -y skills add hiyenwong/ai_collection --skill adaflash-adaptive-speculative-decoding --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Adaflash Adaptive Speculative Decoding?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Adaflash Adaptive Speculative Decoding
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hiyenwong-adaflash-adaptive-speculative-decoding/badge)](https://www.skillsdirectory.com/skills/hiyenwong-adaflash-adaptive-speculative-decoding)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: adaflash-adaptive-speculative-decoding
version: 1.0.0
description: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
trigger_words:
  - adaflash
  - adaptive speculative decoding
  - diffusion drafters
  - on-policy distillation
arxiv_id: 2607.19223
---

# AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

## Overview
AdaFlash is a framework for accelerating large language model inference through adaptive speculative decoding using diffusion drafters. It addresses the high variance issues in diffusion drafters by combining on-policy distillation with adaptive length selection.

## Key Components

### 1. On-Policy Distillation (OPD) for Diffusion Drafters
- Uses reverse-KL divergence tailored specifically for diffusion drafters
- Provides stable convergence and reduces domain-level variance
- Brings consistent acceptance rates across different domains

### 2. Adaptive Length Head
- Dynamically adjusts candidate sequence length during inference
- Substantially lowers verification cost of the target model
- Handles token-level variance effectively

## Implementation Steps

1. **Setup Diffusion Drafter**: Implement or adapt a diffusion-based drafter model that can generate draft sequences in parallel
2. **Apply OPD Training**: Train the drafter using on-policy distillation with reverse-KL divergence
3. **Add Adaptive Length Head**: Implement a mechanism to predict optimal draft length based on context
4. **Integrate with Target Model**: Combine the adaptive drafter with your target LLM for speculative decoding
5. **Tune Hyperparameters**: Adjust temperature, length prediction thresholds, and verification parameters

## Benefits
- Up to 66% higher throughput compared to previous state-of-the-art methods
- Consistent performance across different domains
- Especially effective in high-concurrency scenarios
- Reduces both domain-level and token-level variance

## Use Cases
- LLM inference acceleration in production systems
- High-throughput text generation services
- Real-time conversational AI applications
- Batch processing of large text corpora

## References
- Paper: [AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters](https://arxiv.org/abs/2607.19223)
- Related work: DFlash, speculative decoding, on-policy distillation

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…