As an Amazon Associate we earn from qualifying purchases.
Okay, so I saw a TIL about StyleSynthesis the other day, and it led me down a rabbit hole. I understand the general concept,generating images or text that match a particular style,but I'm really curious about how it incorporates user feedback to improve. Is it all just reinforcement learning, trial and error with a reward function that aligns with "good" feedback? Like, if lots of people consistently downvote images with a specific kind of artifact, does it simply learn to avoid generating those artifacts?
It feels like it must be more nuanced than that, especially when you're talking about subjective aesthetic qualities. How does it deal with contradictory feedback? Does it try to cluster users based on their preferences to create specialized style models, or is it all averaged out? And what kind of feedback is most valuable - explicit ratings, or implicit feedback like time spent viewing an image? I'm hoping someone with a better understanding of the inner workings can ELI5!