Ending AI Slop — Thais Castello Branco, Taste Labs
Ending AI Slop
AI Engineer
At AI Engineer World's Fair 2026, Thais Castello Branco argues that ending AI slop requires treating taste as a measurement and routing problem. Subjective work should be decomposed into parts that can be verified, while the remainder stays attached to contextual human preference instead of being averaged into a false universal score.
AI Engineer's official schedule lists the session as Ending AI Slop, presented by Castello Branco on the Data Quality track on June 30, 2026. It is distinct from her next-day Design Engineering session, Training Taste.
Key Points Covered
- Route failures by layer and method: Taste Labs works on model evaluation and training with frontier labs, while some agent and application problems may be better addressed through context or user intent. Castello Branco frames decomposition as the way to decide which method fits each part [00:00:13]-[00:02:19].
- Capability follows what a domain makes measurable: Castello Branco argues that models perform well on code partly because code decomposes, verifies, and executes. Design and writing are harder to train because good output lacks one equally stable definition [00:02:19]-[00:03:25].
- Context and time prevent one universal taste score: A design can fit a startup and fail a finance firm, while accepted styles change over time. Evaluation therefore has to preserve audience, situation, and temporal context [00:03:25]-[00:04:27].
- Decompose fuzzy goals before choosing a verifier: Brand adherence becomes more tractable when broken into colors, typography, motion, animation, and texture. The ground truth can grade valid use of those components without requiring the agent to copy the original artifact [00:04:27]-[00:07:24].
- Route each component to the method that fits it: Programmatically observable properties can become checks or reinforcement-learning environments; creativity, style, and contextual preference still need human judgment. The goal is to pull only the measurable parts toward verification, not pretend the whole task is objective [00:06:21]-[00:10:29].
- The most likely answer is not the best subjective answer: Castello Branco argues that a single optimization stream can collapse creative output toward the mean. Useful variation has to break patterns intentionally rather than reward randomness or one average style [00:07:24]-[00:09:26].
- Preference disagreement can be signal rather than noise: Aggregating ratings without knowing who preferred what collapses distinct tastes into a misleading average. Castello Branco suggests attaching a preference vector to preference data to preserve plural preferences instead of treating them as noise [00:10:29]-[00:12:25].
- High-signal feedback is specific and anchored: Expert selection, precise evaluative language, and tying commentary to the exact code component being judged make subjective feedback less noisy and more useful for training or evaluation [00:11:23]-[00:14:26].
- Consensus should depend on the attribute: Expert disagreement about alignment may expose defective data because alignment is comparatively objective; disagreement about style or aesthetics can be valid evidence of different preferences [00:13:22]-[00:15:28].
- Quality matters more than annotation volume: Castello Branco closes by advocating for high-quality data created by people with deep domain understanding rather than larger collections of noisy judgments [00:15:28]-[00:16:10].
Full video: https://www.youtube.com/watch?v=lCBf9slCanI(opens in a new tab)