Skip to content
Back to skills

Diffusionvl Ar To Diffusion

ASecurity

Convert pre-trained autoregressive vision-language models into diffusion VLMs without architectural modifications. Use block diffusion strategy enabling arbitrary-length generation and KV-cache reuse. Hybrid attention enforces bidirectional within blocks, causal between blocks. Requires less than 5% of data compared to prior diffusion VLM methods.

  • 6 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 9, 2026
code-quality

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add ADu2021/skillXiv --skill diffusionvl-ar-to-diffusion --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Diffusionvl Ar To Diffusion?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Diffusionvl Ar To Diffusion
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/adu2021-diffusionvl-ar-to-diffusion/badge)](https://www.skillsdirectory.com/skills/adu2021-diffusionvl-ar-to-diffusion)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: diffusionvl-ar-to-diffusion
title: "DiffusionVL: Converting Autoregressive Vision-Language Models into Efficient Diffusion VLMs"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: https://arxiv.org/abs/2512.15713
keywords: [vision-language, diffusion-models, paradigm-shift, parallel-generation, model-conversion]
description: "Convert pre-trained autoregressive vision-language models into diffusion VLMs without architectural modifications. Use block diffusion strategy enabling arbitrary-length generation and KV-cache reuse. Hybrid attention enforces bidirectional within blocks, causal between blocks. Requires less than 5% of data compared to prior diffusion VLM methods."
---

## Skill Summary

DiffusionVL demonstrates direct conversion of already vision-language aligned autoregressive models into diffusion VLMs through full-parameter diffusion finetuning. The approach converts the "next-token prediction paradigm into a diffusion paradigm." Block diffusion strategy enables arbitrary-length generation with intra-block parallel denoising and inter-block autoregressive decoding. Hybrid attention mechanism enforces bidirectional attention within blocks and causal attention between blocks. Remarkably, requires less than 5% of data compared to prior diffusion VLM methods, proving "the gap between dVLMs and AR-VLMs is minimal."

## When To Use

- Converting existing AR vision-language models to parallel-generation diffusion models
- Scenarios requiring efficient VLM inference with minimal retraining
- Projects where parallel generation benefits outweigh minor quality trade-offs
- Research on paradigm-agnostic model conversion techniques

## When NOT To Use

- Building VLMs from scratch (end-to-end diffusion training may be preferable)
- Applications where AR properties are critical (causal attention constraints)
- Scenarios with very limited training data (even 5% is substantial for some projects)
- Domains where hybrid attention mechanism causes artifacts

## Core Technique

Two main pathways enable vision-language paradigm conversion:

**1. AR-VLM to dVLM (Paradigm Shift)**
Direct conversion of already vision-language aligned autoregressive models through full-parameter diffusion finetuning. Convert the "next-token prediction paradigm into a diffusion paradigm" using minimal data.

**2. AR-LM to dVLM (Modality + Paradigm Shift)**
Two-stage approach: connector first aligns vision and text spaces using autoregressive training, then diffusion finetuning completes the conversion. Enables VLM creation from vision and language separately.

**3. Block Diffusion Strategy**
Enable arbitrary-length generation and KV-cache reuse through:
- Intra-block parallel denoising: generate block contents in parallel
- Inter-block autoregressive decoding: generate blocks sequentially

**4. Hybrid Attention Mechanism**
Enforce bidirectional attention within blocks (full context for parallel denoising) and causal attention between blocks (respecting generation order). This balances efficiency with coherence.

## Implementation Notes

Start with pre-trained AR-VLM. Implement block diffusion strategy with hybrid attention pattern. Fine-tune entire model with diffusion objective using your target data (only 5% of original AR training needed). Monitor accuracy preservation and measure parallel generation speedup. Validate on your target VLM tasks.

## References

- Original paper: DiffusionVL (Dec 2025)
- Block diffusion language models
- Vision-language model architectures

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…