{"repo":"Gen-Verse/dLLM-RL","free":true,"listed":false,"github":"https://github.com/Gen-Verse/dLLM-RL","clone":"git clone https://github.com/Gen-Verse/dLLM-RL.git","description":"[ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.","language":"Python","stars":520,"topics":["code-generation","diffusion-language-models","large-language-models","llm-reasoning","mathmatical-reasoning","reinforcement-learning-algorithms","rlhf"],"license":"Apache-2.0","category":"dev-tools","readme_excerpt":"Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models Most comprehensive framework for dLLM's and multimodal dLLM's post-training 🌱 Features - Model Support : TraDo, SDAR, Dream, LLaDA, MMaDA, LLaDA-V, and Diffu-Coder Almost all open-sourced discrete diffusion language models are supported here. - Diverse Settings : We support deployment, SFT , RL (with optional value model for variance reduction and process reward model for fine-grained supervision), and RLHF across diverse settings ( math, coding, multimodal ) and different architectures ( both full/block attention dLLMs ). - Inference Acceleration : improved KV-cache, jetengine (based on nano-vllm), different sampling strategies, support multi-nodes, easy to build your own accelerated inference methods. - RL Training : TraceRL (support diffusion value model), coupled RL, random masking RL, accelerated sampling, including Math, coding, and general RL tasks, support multi-nodes, easy to build your reinforcement learning methods across diverse settings - SFT : Block SFT, semi-AR SFT, random masking SFT, support multi-nodes and long-CoT finetune. 🧠 RL Methods (TraceRL) & Models (TraDo) We propose TraceRL , a trajectory-aware reinforcement learning method for diffusion language models, which demonstrates the best performance among RL approaches for DLMs. We also introduce a diffusion-based value model that reduces variance and improves stability during optimization. Based on TraceRL, we derive a","default_branch":null,"files":null,"tree":[],"storefront":"/r/Gen-Verse","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Gen-Verse/dLLM-RL/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}