Style consistency across a set
Generating one good image is close to solved. Generating forty that belong together is not. We are working on conditioning that holds a look stable across seeds, subjects and aspect ratios.
Home / Research
ResearchWe train and evaluate our own models, so the agenda comes from the failures people actually hit rather than from a leaderboard.
Ordered by how often users raise them, which is also the order we work on them.
Generating one good image is close to solved. Generating forty that belong together is not. We are working on conditioning that holds a look stable across seeds, subjects and aspect ratios.
Hands, teeth, reflections, and text inside an image. People forgive a soft background; they do not forgive six fingers. Most of our evaluation weight sits here.
Every step is compute someone pays for. Work on schedulers and distillation aims to hold output quality while cutting the number of passes needed to reach it.
Content credentials are easy to attach and easy to strip. We are interested in marking that survives a screenshot, a re-encode and a crop.
Automated image-quality scores reward images that look plausible at a glance. They are poor at catching the failures that actually cause someone to throw a generation away and start again.
So alongside standard metrics we run structured human review on a fixed prompt set, weighted toward the failure categories above, and we track how often a generation gets regenerated — the most honest quality signal we have.
We have not published papers. When we do they will be listed here rather than announced and never linked. In the meantime the work shows up in the model, and the gallery is the honest version of a results table.
If you break our model in an interesting way, we want the prompt that did it.