York University · M.A. research
ASC-PIE: An Evaluation Framework for PII-Aware Named-Entity Recognition
I completed my M.A. in Information Systems & Technology at York University, and it was officially awarded in 2026.
ThesisOfficial degree status · 2026
SPRINT-PPPaper accepted to IEEE CASCON 2026 (Toronto, 10–12 Nov 2026); the method is also part of my M.A. thesis.
Thesis abstract
What the thesis set out to do
The robust extraction of Personally Identifiable Information (PII) is essential for privacy protection in modern text-processing systems, where users often share sensitive details in prompts, emails, chat logs, and support tickets. As PII categories and deployment domains evolve, updating extraction models can improve coverage but may also cause catastrophic forgetting of previously learned types. This thesis investigates how PII extraction can remain accurate, reliable, and maintainable as task scope expands. It introduces ASC-PIE, a unified English corpus and evaluation framework that combines public datasets with a synthetic component to improve coverage of rare and challenging PII cases without using real personal data. Using ASC-PIE, the thesis compares encoder-based, encoder-decoder, and decoder-only models under supervised fine-tuning and in-context prompting. It also proposes SPRINT-PP, a privacy-safe continual-learning method that mitigates forgetting without storing raw historical examples. The evaluation covers extraction quality, robustness, output validity, computational efficiency, and knowledge retention.
Research at a glance
A shared corpus, protocol, and reproducible benchmark
Validity is the share of generated outputs that follow the required structure. The four strongest models produce usable output every time. The decoder-only and BART-base models do not.
SPRINT-PP achieved the strongest tested accuracy while remaining privacy-safe and storing no raw historical PII.
Findings
Answers to the four research questions
RQ1Supervised model families
RoBERTa-large is the strongest supervised model, with a strict F1 of 0.991 and perfect output validity. FLAN-T5-base is the closest generative alternative at 0.987, also with perfect validity. The decoder-only models fall behind mainly on recall: when they produce a valid prediction it is usually right, but they recover fewer entities overall. Validity matters in a privacy pipeline, because a model whose outputs are often structurally unusable is not a dependable extractor however good its best answers are, so RoBERTa-large became the backbone for the continual-learning work.
The chart above ranks all seven models by strict F1 and shows each one's validity.
RQ2Fine-tuning versus prompting
For decoder-only models, in-context prompting is a strong alternative that needs no training. The best key-value three-shot prompts beat the matching fine-tuned baselines on the reported subsets, most clearly for Qwen2.5-7B (0.778 against 0.400). The output format matters as much as the number of examples: key-value is consistently stronger than JSON on quality, validity, and speed, at about 1.76 s/query against 26.8 s/query for Qwen. Three shots is the best operating point, because five shots cost more time without consistent gains.
| Configuration | Strict F1 | Validity |
|---|---|---|
| Qwen2.5-7B · fine-tuned | 0.400 | 0.543 |
| Qwen2.5-7B · JSON, 3-shot | 0.709 | 0.906 |
| Qwen2.5-7B · key-value, 3-shot | 0.778 | 0.950 |
| Llama3.1-8B · fine-tuned | 0.690 | 0.814 |
| Llama3.1-8B · JSON, 5-shot | 0.705 | 0.856 |
| Llama3.1-8B · key-value, 3-shot | 0.725 | 0.914 |
RQ3Forgetting and mitigation
Naive sequential fine-tuning collapses. Each new stage is learned well, but every earlier stage falls to an F1 of 0.000 in the stage-restricted view, so the model replaces what it knew instead of adding to it. SPRINT-PP is the strongest mitigation, reaching 83.5% final accuracy against 83.0% for distillation and 76.5% for replay. It stores no raw historical PII, and it trains in 25.6% less time than distillation.
The chart above compares the four update strategies on final accuracy.
RQ4Corpus design and synthetic augmentation
Corpus design is a major driver of extraction quality, not a preprocessing detail. With the same RoBERTa-large backbone and the same evaluation, training on public data alone reaches a strict F1 of 0.415, with recall limited to 0.307. Training on the full ASC-PIE corpus reaches 0.992, with recall at 0.995, a gain of +0.577. Strict and normalised F1 are nearly identical in both settings, so the gain reflects real coverage of entities, not formatting. The gains are broad but uneven, and largest where public data alone gave the least coverage.
| PII type | Public-only | Full ASC-PIE | Gain |
|---|---|---|---|
| JOBTITLE | 0.0000 | 0.9982 | +0.9982 |
| 0.0786 | 0.9940 | +0.9154 | |
| SIN | 0.2236 | 0.9980 | +0.7744 |
| CITY | 0.2826 | 0.9995 | +0.7169 |
| NAME | 0.2716 | 0.9824 | +0.7108 |
| STATE | 0.2891 | 0.9986 | +0.7095 |
| LOCATION | 0.4477 | 0.9102 | +0.4625 |
| PHONENUM | 0.7340 | 0.9938 | +0.2598 |
| COUNTRY | 0.8000 | 1.0000 | +0.2000 |
| CREDITCARD | 0.9786 | 0.9911 | +0.0125 |
| DRIVER-LICENCE | 0.9744 | 0.9756 | +0.0012 |
| TIME | 1.0000 | 0.8333 | -0.1667 |
| TAXNUM | 1.0000 | 0.7500 | -0.2500 |
TIME and TAXNUM are slightly lower under full ASC-PIE, but the matched subset holds only 5 TIME mentions and 3 TAXNUM mentions, so those two differences should not be over-read.
Research framework
A comparable path through privacy datasets
PII-aware NER evaluation
ASC-PIE evaluates named-entity recognition for personally identifiable information across privacy-focused datasets.
Research data pipeline
The ASC-PIE experiments prepare real and synthetic privacy datasets for NER training and evaluation.
PII label standardization
The research pipeline maps differing PII label schemes into a shared representation for comparison.
ASC-PIE experiment stack
ASC-PIE experiments use Python, PyTorch, Hugging Face Transformers, scikit-learn, and seqeval.
Research links