this post was submitted on 04 Sep 2026
3 points (100.0% liked)

Stable Diffusion

5708 readers
14 users here now

Discuss matters related to our favourite AI Art generation technology

Also see

Other communities

founded 3 years ago
MODERATORS
 

Abstract

We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.

Paper: https://arxiv.org/abs/2609.03796

Code: https://github.com/inclusionAI/LLaDA-Image

Model: https://huggingface.co/inclusionAI/LLaDA-Image

Turbo Model: https://huggingface.co/inclusionAI/LLaDA-Image-Turbo

no comments (yet)
sorted by: hot top controversial new old
there doesn't seem to be anything here