SkinDeepRESEARCHSteve Seguin

Exploratory results

Find names to redact

Return marked text spans instead of rewriting a document.

Back to decisions without text

Look for names that may need masking before sharing a document. The demo now checks the entire input; this is name detection, not complete anonymization.

Same classifierComplete name occurrences markedOther tokens incorrectly marked
First 1,200 tokens290 / 445339
Entire input319 / 445378

Reading the whole input found 29 additional complete name occurrences, while marking 39 more unrelated tokens. Some names are still missed or only partly marked.

What was tested?

The same 50 TAB court documents, with the original trained weights and cutoff. Both rows count all annotated name occurrences, including the old unprocessed tail. The original model used 100 training and 25 development documents; this follow-up reuses the earlier test documents. All fit within the demo’s 30,000-character limit.

Full-input comparison and every outcome · Original model and dataset