Visual Quality Intelligence, Part 3: The Judge That Decided What AI Finds Beautiful
Part 3 of a series on how AI learns to see quality. Read Part 2 here.
Visual generative models are trained on billions of images and videos, but only the highest-quality data makes it into the training set. The selected data determines what a model will be able to generate, and models perform poorly on concepts that are rare in their data. The quality bar applied to the training set quietly shapes everything the model can do.
So what kinds of images make it into foundation model training, and who decides?
The starting corpus for a foundation model spans billions of images. It is impossible for humans to review data at that scale. So researchers built models that automatically assign each image a quality score, say on a scale of 1 to 10. These models are called aesthetic predictors.
How do you train an aesthetic predictor? With human preference data. AVA collected aesthetic ratings from photography contests. Simulacra Aesthetic Captions collected 1-to-10 ratings of early AI-generated images. The most influential aesthetic predictor, the LAION-Aesthetics Predictor (LAP), was trained on these datasets and then used to filter the images that trained models like Stable Diffusion and Midjourney.
As generated images became a dominant share of online visual content, researchers started to look more closely at aesthetic predictors. Their role is hard to overstate: generative models can only learn to generate what survives the filter, and new models are often ranked on scores from an aesthetic predictor. The judge shapes the model twice: once through the data, once through the leaderboard.
So what exactly does the predictor find beautiful?
Researchers answered this question earlier this year. The Algorithmic Gaze traced the origin of the LAP and audited its preferences. The audit found that LAP disproportionately filters in images whose captions mention women, while filtering out images whose captions mention men or LGBTQ+ people. Scoring roughly 330k artworks, the authors also found that LAP rates realistic landscapes, cityscapes, and portraits from western and Japanese artists most highly, while abstract and African art score lowest, reinforcing what they describe as the imperial and male gaze of western art history.
Why these biases? The paper paired its audit with a trace ethnography: audits identify algorithmic biases, while ethnographies help explain where the biases came from by focusing on the people behind the algorithm. The aesthetic scores used to train LAP came primarily from English-speaking photographers and western AI enthusiasts. A W.E.I.R.D. rater pool: Western, Educated, Industrialized, Rich, Democratic. LAP is effectively a representation of the average aesthetic preference of that small group of raters.
Most striking of all: LAP was trained by a single individual, LAION co-founder Christoph Schuhmann. During training, he picked the final model architecture based on his own visual preference, even though other candidates scored lower error.
One person's eye effectively overrode the collective vote, and that model went on to filter the data behind image generators used by millions of people.
Last week I asked how you judge output quality: a metric, your eyes, or a vibe check. For the most consequential data filter in generative AI, the answer was a vibe check.
Why does this matter? A model trained on average opinion learns to generate average outputs. Creative and innovative images tend to divide raters, and disagreement pulls down the mean, so a filter built on average preference systematically suppresses exactly the outputs that are most original. This might explain why artists are so critical of the increasingly photorealistic outputs of visual generative models: as models tune toward average beauty, creative and unusual outputs get deprioritized, which limits how well these models can serve creative and artistic expression.
The core issue is that most aesthetic predictors assume quality can be expressed as a single numerical score, valid for any image and void of any context. That assumption is fundamentally flawed, because quality is highly dependent on context.
Aesthetic predictors need to redefine quality: not as a single numerical score, but as a context-dependent, explainable judgment. Context-aware judgment can accommodate average opinion as well as more demanding quality requirements. Adding context can help broaden the output space of visual generators beyond average aesthetics. And explainable, context-based judgments create the possibility of aesthetic predictors being useful inside agentic visual generation loops. Recall the example from Part 2: for a restoration task, a blurry but semantically accurate photo of a dog in a park is a better output than a gorgeous photo of a boat. The degraded input and the task are the context that makes the right judgment possible.
The audit's authors land in a similar place, calling on AI developers to shift away from prescriptive measures of aesthetics toward more pluralistic evaluation.
Judging quality with a single numerical score is flawed. Context-rich, explainable quality judgment is the next era: it is what will make aesthetic predictors truly useful for visual generation, and what will unlock high-quality creative agentic visual generation.
Next week, Part 4: the 2004 metric still running inside your pipeline.
What do you find beautiful that AI keeps missing? Which quality filter shaped your training data, and do you know what it prefers? Let me know in the comments, or find me on LinkedIn.
References
- J. Taylor, W. Agnew, M. Sap, S. E. Fox, and H. Zhu, "The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor," FAccT 2026.
- LAION, "LAION-Aesthetics," 2022.
- N. Murray, L. Marchesotti, and F. Perronnin, "AVA: A Large-Scale Database for Aesthetic Visual Analysis," CVPR 2012.
- C. Schuhmann et al., "LAION-5B: An open large-scale dataset for training next generation image-text models," NeurIPS 2022.
This is Part 3 of the Visual Quality Intelligence series. Follow along on LinkedIn or subscribe for the full bibliography with each part.
Member discussion