Notes · 05
Beautiful is
not a boolean.
Ask a model "is this beautiful?" and you will get an answer, delivered with total confidence, worth almost nothing. The failure is not that the model is dumb. The failure is that the question is malformed. Beautiful is not a property a design has or lacks, like a missing semicolon. This note is about what beauty actually is, and how that one idea shaped the way uistash decided to ask.
01 · The malformed question
Beauty is a direction, not a classification.
Every design is a bundle of trade-offs: density against calm, character against familiarity, whitespace against information. There is no point in that space labeled correct. What a person calls beautiful is a direction through it, the trades they consistently prefer, and the direction differs from person to person in exactly the ways that matter.
A classifier needs a shared ground truth. Beauty has none. So any system that asks a machine "is this good?" is really asking "is this near the average of everyone?", which is the one place a personal taste, by definition, is not.
02 · Ask the person
Keep or kill, blind, repeatedly.
So uistash never asks the model to judge beauty. It asks the person, and it stacks the deck against self-deception. A blind comparison ran one brief across up to six sides, each an engine paired with a taste context: nothing, or your tastestack, or a file you dropped in. The results came back in a shuffled order that never revealed which side made which design. You looked at the work, only the work, and made a pick.
The vocabulary was deliberately small and deliberately honest: pick what you would keep, or say neither, or say unsure. The escape hatches are first-class answers, because a forced choice between two designs you would never ship is not taste data, it is noise wearing taste's clothes.
03 · One coordinate at a time
The labels reveal only after you choose.
Only after the pick landed did the comparison reveal which engine and which taste context made what. The order matters: judge first, learn afterward, so the reveal can inform you without ever steering you. Each pick was recorded as evidence, one preference at a time.
One pick says almost nothing, and that is fine. It is one coordinate. Enough coordinates and a direction emerges that no single answer contained: the trades you keep making when nobody, including the interface, is telling you what you are choosing between.
04 · The instrument
The human judges. The machine measures.
There is a machine judge in the Lab, and the way it is caged is the whole point. It never answers "which is better?". It answers a different question: "which would this specific person save?", anchored to the invariants distilled from their actual library. It runs only after the human has picked, compares each pair twice with positions swapped, and when its two calls disagree it abstains instead of guessing.
Its picks live in a separate ledger from the human's, and what gets tracked is agreement, not authority. The machine is calibrating against the person, never substituting for them. Beauty stays a question only the person can answer. The machine's job is to make the answering rigorous: blind, repeated, recorded. The human stays the judge. The machine stays the instrument.
Where the instrument went.
The room that ran these comparisons is no longer in the app. The argument it was built on did not change: the person stays the judge, and the instrument's only job is to keep them honest.