TAE Qwen-Image 2.1 · AcademiaSD

Tiny AutoEncoder (TAESD-style decoder) for fast, sharp live previews of Qwen-Image 2.1 in ComfyUI.

Example

Original · TAE · Latent2RGB

Left: Qwen-Image 2.1 VAE · Center: TAEQwenImage21_AcademiaSD · Right: Latent2RGB (ComfyUI's default preview for this model)

ComfyUI previews Qwen-Image 2.1 only with Latent2RGB, a linear 64→3 color projection that looks blurry and blocky. This decoder is distilled from the Qwen-Image 2.1 VAE: it turns the sampler's latents into a real preview image in about 15 ms, with no need for the full VAE.

File TAEQwenImage21_AcademiaSD.safetensors
Size 3.3 MB (fp16)
Parameters 1.63 M
Input Qwen-Image 2.1 latents, 64 channels, as the sampler sees them (x0)
Output RGB at 16× the latent resolution, range [0, 1]
Decode time ~15 ms for a 1024×1024 preview (RTX 5080, fp16)
Validation PSNR 30.1 dB vs 20.9 dB for Latent2RGB

🚀 Usage in ComfyUI

  1. Download TAEQwenImage21_AcademiaSD.safetensors to ComfyUI/models/vae_approx/.
  2. Install ComfyUI-KJNodes.
  3. Add the Model Preview Override node between your Qwen-Image 2.1 model and the sampler.
  4. In its tiny_vae input, choose TAEQwenImage21_AcademiaSD.safetensors.

ComfyUI's built-in TAESD preview method does not pick up this file: core has no tiny decoder entry for Qwen-Image 2.1 and its TAESD loader only supports 8× decoders. Use the KJNodes node.

🔧 Training details

  • Teacher: Qwen-Image 2.1 VAE (qwen_image_2.1_vae_bf16.safetensors).
  • Architecture: flat TAESD decoder, Clamp → conv → 4 × (3 blocks + 2× upsample + conv) → block → conv, width 64.
  • Latent space: latents normalized with ComfyUI's QwenImage21 latent format (mean/std), the same space as the sampler's x0, so no extra scaling is needed.
  • Data: 4,072 varied images, 512×512 crops, encoded once with the real VAE.
  • Training: 30,000 steps, batch 8, 256×256 tiles, AdamW LR 5e-4 with cosine decay, L1 + FFT loss, EMA 0.999, latent noise augmentation for stable previews on the noisy x0 of the first steps.

⚠️ Limitations

  • Preview only: use the real VAE for the final decode.
  • Decoder only, there is no encoder.
  • RGB only: the Qwen-Image 2.1 VAE also decodes an alpha channel, which this decoder skips.

🇪🇸 Español

Tiny AutoEncoder (decoder tipo TAESD) para previsualizar Qwen-Image 2.1 en ComfyUI con nitidez y en tiempo real. Sustituye a Latent2RGB, la previsualización por defecto, borrosa y pixelada. Destilado del VAE de Qwen-Image 2.1: decodifica una previsualización de 1024 px en unos 15 ms, con 30,1 dB de PSNR frente a 20,9 dB de Latent2RGB.

Uso:

  1. Descarga TAEQwenImage21_AcademiaSD.safetensors en ComfyUI/models/vae_approx/.
  2. Instala ComfyUI-KJNodes.
  3. Pon el nodo Model Preview Override entre el modelo de Qwen-Image 2.1 y el sampler.
  4. En su entrada tiny_vae, elige TAEQwenImage21_AcademiaSD.safetensors.

El método de previsualización TAESD que trae ComfyUI no lo usa, porque no tiene entrada para Qwen-Image 2.1. Usa el nodo de KJNodes.

Solo sirve para previsualizar: la imagen final se decodifica con el VAE real.


🙏 Credits

🔗 AcademiaSD

YouTube · X / Twitter · Discord · Ko-fi

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support