SkinDeepRESEARCHSteve Seguin

Which example should come next?

Some ratings teach the model more than others.

Something you may like
Something it is unsure about
Something different
The demo mixes these choices rather than showing only its favourites.

If the model only shows things it already thinks you like, it may never learn what it is missing. A few uncertain or different examples can help.

The actual demo sampler has now been tested on made-up preferences. Its mixture did not consistently beat random sampling, and performed worse with noisy ratings. Enjoyment and rating effort still need a human study.

Sampling methods and tests

Four useful baselines

  • Random: an unbiased reference distribution.
  • Uncertainty: sample near the current boundary.
  • Diversity: cover distinct parts of the representation.
  • Mixed: combine predicted likes, uncertain samples, and exploration.

The browser currently selects 5 predicted likes, 4 uncertain samples, and 3 exploratory samples per deck of 12. The queue is refreshed after every six ratings, so only half of each shuffled deck is rated. This exact behavior is included in the new comparison.

What counts as improvement

Measure held-out quality at equal label budgets, candidate generation cost, duplicates, and selection latency. User fatigue and satisfaction require their own study; classifier confidence is not a substitute.

The previous page’s fixed efficiency percentages are unvalidated and have been removed from current guidance. P03 records synthetic comparisons; no human satisfaction result is claimed.

Sampling results · Archived discussion