I noticed that in Generator Mode, the conditioning image is only passed through the VAE and is not fed through the ViT for high-level semantic understanding. I’m curious about the reasoning behind this design choice. Were any ablation studies conducted to validate it? Thank you!
I noticed that in Generator Mode, the conditioning image is only passed through the VAE and is not fed through the ViT for high-level semantic understanding. I’m curious about the reasoning behind this design choice. Were any ablation studies conducted to validate it? Thank you!