SkinDeepRESEARCHSteve Seguin

Research FAQ

What the experiments show, how the methods work, and what is still untested.

Learning preferences

Can it learn what I like?

The demo learns patterns in your ratings and uses them to score or adjust drawings. Its calculations and simple made-up preferences have been tested; we have not yet shown that people prefer its suggestions.

Does choosing examples carefully save ratings?

Not consistently in our tests. The demo’s mix of likely likes, uncertain examples and random examples performed worse with noisy ratings than random sampling; selecting uncertain examples did better on the simple test rules.

Can it learn different tastes or changing moods?

A single linear model struggles with separated tastes. A nonlinear feature model helped in a synthetic test; keeping separate models for explicitly named contexts also helped in a small alternating-context test. Neither demonstrates learning a person’s changing moods.

Does a higher model score mean a better suggestion?

No. A model can confidently prefer the wrong thing. We tested selection from the same candidate pool and retained cases where maximizing the score disagrees with the known preference; a person’s new rating is still the real check.

Decisions and speed

Does the model need to write an answer?

No. A small classifier can read its internal numbers and return a label directly. Returning a number instead of one trained token did not, by itself, show a reliable speed advantage.

How early can it stop, and does that save time?

A shared four-layer BERT model can answer after layer 2. On 500 unused messages, its frozen rule preserved every full-depth decision while proposing 400 early exits, or 40% fewer transformer blocks. Actual browser execution also works. Only 18 new test messages were toxic, so rare-error reliability remains uncertain.

Can it go faster without changing its answers?

On 100 tested messages, reusing the fixed instructions cut time by 37.9% with identical decisions. Processing messages in groups of four cut queue time by 9.7%. Both still use all 24 layers; these checks do not establish identical behavior on every possible input.

Would a tiny model with a larger fallback be simpler?

The original pilot roughly halved time, but introduced a toxic-message miss. On 100 new messages, a stricter fallback matched Qwen’s decisions while avoiding Qwen only twice; its measured time saving was less than 1%.

Does it follow changing rules or understand a live conversation?

Not reliably in the tests so far: opposite rules exposed failures even at full depth. We have not tested these decision classifiers on live conversations with audio or video context.

Clicks and actions

Can it return a click location directly?

Yes, a pointer model can return a location without writing coordinates as text. It found 25/30 targets in the first study, but only 8/25 on a harder new benchmark. A confidence cutoff withheld all 25 infeasible requests there, while also withholding 21 valid ones and still allowing two wrong clicks.

Can those decisions finish a task?

A separate small grid model completed 7 of 10 mazes after more training and a rule against wall moves. That is not a completed GUI task: the screenshot tests return points and do not execute clicks.

Longer context

Can a nearly full GPU use disk for a 200K-token context?

Saving text to files can let a long task use a short working context. It does not give a GPU unlimited active memory. This model needs about 12.2 GiB of attention cache for 200,000 tokens, plus other working state. Parking an inactive numerical cache and streaming it during generation are different designs; neither has a completed benchmark in this research snapshot.

Can self-editing let a task continue beyond its context budget?

The revised self-editing agent processed a 121K-token ledger within a 32K working budget. It got 24/24 in 15 minutes on one seed and 21/24 in 21 minutes on another. The three misses came from a bug in its update script, with every batch retained. The 480K-stream result is still pending. Keeping the data available does not by itself ensure correct computation.

What changes the wait, writing speed and accuracy?

Prefix caching cuts repeated reading: cached 200K questions started in about 2 seconds, versus 114 seconds in a separate cold probe. Standard-gate decode stayed near 89 tokens/s with or without caching. Longer active context still slowed generation; the recall probes ranged from about 127 tokens/s at 8K to 27 at 250K. Recall was 708/720 across the test, including 120/120 at 60K. See the tables for sample sizes and differences between the runs.

How are external notes different from an editable conversation?

A notes file is stored outside the active conversation and must be loaded when needed. A CLM mirror edit changes the message history sent to the model on the next call. On the 121K-token ledger, both file-using agents got 24/24 in about 1.9 minutes while holding under 9K active tokens. That is one task result, not a guarantee that every detail was archived or that every long document will be equally fast.

Can a CPU cleaner turn a 32K context into 100K?

The measured cleanup rules did not: conservative deduplication and formatting cleanup removed 1.83% of roughly five million tokens. Moving large outputs to files was a separate 7.50% reduction, with a proxy indicating some text would need retrieval later. A trained CPU pruning classifier was not benchmarked. Removing earlier reasoning blindly left a separate model run stuck without an answer.

Scope and limits

Can preferences stay within hard requirements?

Explicit bounds can prevent a known forbidden edit; new conflict cases check that mechanism. A high preference score is not a safety guarantee, and personalized instruction-following has not been validated.

Are the ratings and representations private?

The browser demo keeps its ratings on your device. That does not prove an encoded image hides private information; we have not run a reconstruction or privacy-attack study.

Has this been tested on real generators, music or other applications?

Not in the current experiments. Generator transfer, music, feeds, mutual preferences and engineering designs are proposed applications; validating them needs task data, appropriate models and, where relevant, participating users.

Will the results hold on another model or device?

A separate implementation replay reproduced saved decisions and layer stopping, including a published runtime version. We have CPU timings and a full-depth browser demo, but no matched GPU/NPU comparison or evidence that the trained classifiers transfer to another model.

Evidence behind these answers · Technical study protocols