Hacker News
Training Text-to-Image Models 3.6× Faster
The new JiT-DDT encoder-decoder architecture overcomes the compression limitations of standard Latent Diffusion Models by integrating compression directly into the diffusion process. This approach enables 3.6× faster training compared to the Linum v2 baseline while generating images with 4× more pixels. The model code and weights are released under the Apache 2.0 license.