SkinDeepRESEARCHSteve Seguin

How the model learns

Your ratings tell it which details to pay attention to.

1. Start with the settings behind a drawing

A drawing can be described by a few numbers: colour, width, roundness. A renderer turns those settings into the picture you see.

  1. Settings: blue, wide, round
  2. Renderer draws the mug
  3. You choose Like or Pass
For a learned image generator, similar internal numbers are called a latent vector; its controls are usually less tidy.

2. Learn which settings go with your ratings

LikeBlue, wide
PassRed, wide
Illustration: the shape stays the same, so these ratings suggest colour matters. Real preferences need more examples.

The small preference model learns a weight for each setting. A setting associated with likes raises its score; one associated with passes lowers it. Each rating updates those weights. Only the small preference model changes; the drawing renderer stays the same.

3. Use that score to make a choice

  1. New settings
  2. Preference model scores them
  3. Rank, suggest, or make a small edit
To suggest something, search the allowed settings for a high score, then render those settings as a picture.

The generator makes the image. The preference model predicts whether you will like it. Your next rating is the check on that prediction.

What the tests show

The maths and how we test it

Learn from the latent vector

For a latent vector z, the browser model predicts p(z) = sigmoid(w · z + b). Ratings fit the weights and bias; the renderer turns the vector into a face or composition.

On the box [-t,t]^d, a linear score is maximized by z[i] = t · sign(w[i]). A zero-weight dimension can remain unchanged. Choosing a smaller box is an explicit constraint, not proof of realism.

A target score usually has many solutions. To make a minimal edit, minimize ||z − z₀||² subject to the bounds and a target logit. The bounded solution has the form clip(z₀ + λw, -1, 1); a monotone search finds λ. Targets above the attainable maximum are reported as unreachable.

These are statements about the fitted model. They do not guarantee that a person likes the result, that a generator produces valid content, or that latent distance measures visible change.

Several preferences need a richer model

A linear-plus-sigmoid classifier has a linear decision boundary. Over a convex latent box, it cannot represent several disconnected preferred regions. Different points at the same score are not automatically distinct preference modes.

Nonlinear features, a small neural head, separate context profiles, or a mixture can represent richer preferences. They need additional data and their own held-out evaluation.

Evaluate what matters

Training fit measures how well the model matches ratings it has already seen. The demo also offers blind evaluation on uniformly sampled examples; those ratings do not update the model. Confidence remains an uncalibrated model score unless independently tested.

Multiple preference modes · Archived 2D teaching widget

Two useful variations