Exploratory results
Find names to redact
Return marked text spans instead of rewriting a document.
Back to decisions without text
Look for names that may need masking before sharing a document. The demo now checks the entire input; this is name detection, not complete anonymization.
| Same classifier | Complete name occurrences marked | Other tokens incorrectly marked |
|---|---|---|
| First 1,200 tokens | 290 / 445 | 339 |
| Entire input | 319 / 445 | 378 |
Reading the whole input found 29 additional complete name occurrences, while marking 39 more unrelated tokens. Some names are still missed or only partly marked.
What was tested?
The same 50 TAB court documents, with the original trained weights and cutoff. Both rows count all annotated name occurrences, including the old unprocessed tail. The original model used 100 training and 25 development documents; this follow-up reuses the earlier test documents. All fit within the demo’s 30,000-character limit.
Full-input comparison and every outcome · Original model and dataset