Skip to content
Back to skills

Dirl Diffusion Rl

ASecurity

Enable effective RL for diffusion language models via DiPO (unbiased GRPO for dLLMs) and framework optimizations. FlexAttention accelerates blockwise training, LMDeploy optimizes inference, achieving training-inference consistency—improving dLLM math performance to rival larger autoregressive models.

  • 6 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
devopspythongogitperformance

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add ADu2021/skillXiv --skill dirl-diffusion-rl --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Dirl Diffusion Rl?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Dirl Diffusion Rl
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/adu2021-dirl-diffusion-rl/badge)](https://www.skillsdirectory.com/skills/adu2021-dirl-diffusion-rl)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: dirl-diffusion-rl
title: "DiRL: Efficient Post-Training for Diffusion Language Models"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: https://arxiv.org/abs/2512.22234
keywords: [diffusion-language-model, reinforcement-learning, policy-optimization]
description: "Enable effective RL for diffusion language models via DiPO (unbiased GRPO for dLLMs) and framework optimizations. FlexAttention accelerates blockwise training, LMDeploy optimizes inference, achieving training-inference consistency—improving dLLM math performance to rival larger autoregressive models."
---

## Overview

DiRL introduces RL infrastructure tailored for diffusion language models.

## Core Technique

**DiPO Algorithm:**
First unbiased Group Relative Policy Optimization for dLLMs.

```python
def dipo_training(model, dataset):
    # Blockwise attention for efficient computation
    # Unbiased logit computation (fixes prior biases)
    # GRPO with dLLM-specific optimizations
```

## Performance

- State-of-the-art dLLM math performance
- Outperforms larger autoregressive models

## References

- DiPO: unbiased GRPO for dLLMs
- Blockwise training optimization

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…