As an Amazon Associate we earn from qualifying purchases.
I've been digging into stylesynthesis lately, and the choice between using GANs and VAEs keeps coming up. It seems like GANs often produce sharper, more realistic results, especially when you need fine-grained control over the style transfer. However, they can be notoriously arduous to train and stabilize. I’ve seen a lot of cases where a GAN training run just collapses, yielding nothing useful.
Conversely, VAEs are generally more stable and easier to train, and their latent space representation can be really useful for exploring and manipulating different styles. You can get smoother transitions between styles, which is great for some applications, but the output frequently enough ends up being a bit blurry compared to what a GAN can achieve.It feels like VAEs are a safer bet if you just need something that works, but GANs have the potential to produce the best results if you can get them to behave.
Has anyone else worked with both GANs and VAEs for style synthesis? What are your go-to strategies for mitigating the instability issues with GANs, and are there any tricks you've found to improve the sharpness of VAE-based style transfer? I'm particularly interested in techniques people are using to evaluate the "quality" of the synthesized style beyond just visual inspection.