TAE Qwen-Image 2.1 · AcademiaSD
Tiny AutoEncoder (TAESD-style decoder) for fast, sharp live previews of Qwen-Image 2.1 in ComfyUI.
Left: Qwen-Image 2.1 VAE · Center: TAEQwenImage21_AcademiaSD · Right: Latent2RGB (ComfyUI's default preview for this model)
ComfyUI previews Qwen-Image 2.1 only with Latent2RGB, a linear 64→3 color projection that looks blurry and blocky. This decoder is distilled from the Qwen-Image 2.1 VAE: it turns the sampler's latents into a real preview image in about 15 ms, with no need for the full VAE.
| File | TAEQwenImage21_AcademiaSD.safetensors |
| Size | 3.3 MB (fp16) |
| Parameters | 1.63 M |
| Input | Qwen-Image 2.1 latents, 64 channels, as the sampler sees them (x0) |
| Output | RGB at 16× the latent resolution, range [0, 1] |
| Decode time | ~15 ms for a 1024×1024 preview (RTX 5080, fp16) |
| Validation PSNR | 30.1 dB vs 20.9 dB for Latent2RGB |
🚀 Usage in ComfyUI
- Download
TAEQwenImage21_AcademiaSD.safetensorstoComfyUI/models/vae_approx/. - Install ComfyUI-KJNodes.
- Add the Model Preview Override node between your Qwen-Image 2.1 model and the sampler.
- In its
tiny_vaeinput, chooseTAEQwenImage21_AcademiaSD.safetensors.
ComfyUI's built-in TAESD preview method does not pick up this file: core has no tiny decoder entry for Qwen-Image 2.1 and its TAESD loader only supports 8× decoders. Use the KJNodes node.
🔧 Training details
- Teacher: Qwen-Image 2.1 VAE (
qwen_image_2.1_vae_bf16.safetensors). - Architecture: flat TAESD decoder,
Clamp → conv → 4 × (3 blocks + 2× upsample + conv) → block → conv, width 64. - Latent space: latents normalized with ComfyUI's
QwenImage21latent format (mean/std), the same space as the sampler'sx0, so no extra scaling is needed. - Data: 4,072 varied images, 512×512 crops, encoded once with the real VAE.
- Training: 30,000 steps, batch 8, 256×256 tiles, AdamW LR 5e-4 with cosine decay, L1 + FFT loss, EMA 0.999, latent noise augmentation for stable previews on the noisy
x0of the first steps.
⚠️ Limitations
- Preview only: use the real VAE for the final decode.
- Decoder only, there is no encoder.
- RGB only: the Qwen-Image 2.1 VAE also decodes an alpha channel, which this decoder skips.
🇪🇸 Español
Tiny AutoEncoder (decoder tipo TAESD) para previsualizar Qwen-Image 2.1 en ComfyUI con nitidez y en tiempo real. Sustituye a Latent2RGB, la previsualización por defecto, borrosa y pixelada. Destilado del VAE de Qwen-Image 2.1: decodifica una previsualización de 1024 px en unos 15 ms, con 30,1 dB de PSNR frente a 20,9 dB de Latent2RGB.
Uso:
- Descarga
TAEQwenImage21_AcademiaSD.safetensorsenComfyUI/models/vae_approx/. - Instala ComfyUI-KJNodes.
- Pon el nodo Model Preview Override entre el modelo de Qwen-Image 2.1 y el sampler.
- En su entrada
tiny_vae, eligeTAEQwenImage21_AcademiaSD.safetensors.
El método de previsualización TAESD que trae ComfyUI no lo usa, porque no tiene entrada para Qwen-Image 2.1. Usa el nodo de KJNodes.
Solo sirve para previsualizar: la imagen final se decodifica con el VAE real.
🙏 Credits
- TAESD architecture: madebyollin/taesd
- Model Preview Override node: kijai/ComfyUI-KJNodes
- Qwen-Image 2.1: Qwen team
🔗 AcademiaSD
YouTube · X / Twitter · Discord · Ko-fi
