Tech & AI News
Hacker News

Training Text-to-Image Models 3.6× Faster

The new JiT-DDT encoder-decoder architecture overcomes the compression limitations of standard Latent Diffusion Models by integrating compression directly into the diffusion process. This approach enables 3.6× faster training compared to the Linum v2 baseline while generating images with 4× more pixels. The model code and weights are released under the Apache 2.0 license.