Back to skills
SKILL.md
Dirl Diffusion Rl
ASecurityEnable effective RL for diffusion language models via DiPO (unbiased GRPO for dLLMs) and framework optimizations. FlexAttention accelerates blockwise training, LMDeploy optimizes inference, achieving training-inference consistency—improving dLLM math performance to rival larger autoregressive models.
- 6 stars
- 0 votes
- 0 copies
- 2 views
- Added September 9, 2026
Security analysis
100/100npx -y skills add ADu2021/skillXiv --skill dirl-diffusion-rl --agent claude-codeAre you the author of Dirl Diffusion Rl?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/adu2021-dirl-diffusion-rl)---
name: dirl-diffusion-rl
title: "DiRL: Efficient Post-Training for Diffusion Language Models"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: https://arxiv.org/abs/2512.22234
keywords: [diffusion-language-model, reinforcement-learning, policy-optimization]
description: "Enable effective RL for diffusion language models via DiPO (unbiased GRPO for dLLMs) and framework optimizations. FlexAttention accelerates blockwise training, LMDeploy optimizes inference, achieving training-inference consistency—improving dLLM math performance to rival larger autoregressive models."
---
## Overview
DiRL introduces RL infrastructure tailored for diffusion language models.
## Core Technique
**DiPO Algorithm:**
First unbiased Group Relative Policy Optimization for dLLMs.
```python
def dipo_training(model, dataset):
# Blockwise attention for efficient computation
# Unbiased logit computation (fixes prior biases)
# GRPO with dLLM-specific optimizations
```
## Performance
- State-of-the-art dLLM math performance
- Outperforms larger autoregressive models
## References
- DiPO: unbiased GRPO for dLLMs
- Blockwise training optimization
Attribution
Comments
Loading comments…