York University · M.A. research

ASC-PIE: An Evaluation Framework for PII-Aware Named-Entity Recognition

I completed my M.A. in Information Systems & Technology at York University, and it was officially awarded in 2026.

ThesisOfficial degree status · 2026

SPRINT-PPPaper accepted to IEEE CASCON 2026 (Toronto, 10–12 Nov 2026); the method is also part of my M.A. thesis.

Supervisor
Prof. Marin Litoiu · CERAS Lab, York University
Paper
ASC-PIE and SPRINT-PP: Evaluating Privacy-Safe Continual Learning for PII ExtractionAccepted to IEEE CASCON 2026

Thesis abstract

What the thesis set out to do

The robust extraction of Personally Identifiable Information (PII) is essential for privacy protection in modern text-processing systems, where users often share sensitive details in prompts, emails, chat logs, and support tickets. As PII categories and deployment domains evolve, updating extraction models can improve coverage but may also cause catastrophic forgetting of previously learned types. This thesis investigates how PII extraction can remain accurate, reliable, and maintainable as task scope expands. It introduces ASC-PIE, a unified English corpus and evaluation framework that combines public datasets with a synthetic component to improve coverage of rare and challenging PII cases without using real personal data. Using ASC-PIE, the thesis compares encoder-based, encoder-decoder, and decoder-only models under supervised fine-tuning and in-context prompting. It also proposes SPRINT-PP, a privacy-safe continual-learning method that mitigates forgetting without storing raw historical examples. The evaluation covers extraction quality, robustness, output validity, computational efficiency, and knowledge retention.

Abstract from the York University thesis record

Research at a glance

A shared corpus, protocol, and reproducible benchmark

Total examples333,109
Entity mentions2,025,878
Canonical PII types19
Corpus sources5
Supervised fine-tuning · strict F1 and output validity
RoBERTa-large99.1%validity 1.000
FLAN-T5-base98.7%validity 1.000
ModernBERT-large84.6%validity 1.000
BERT-large-cased80.7%validity 1.000
Llama 3.1 8B69.0%validity 0.814
Qwen 2.5 7B40.0%validity 0.543
BART-base34.8%validity 0.561

Validity is the share of generated outputs that follow the required structure. The four strongest models produce usable output every time. The decoder-only and BART-base models do not.

Continual-learning accuracy by strategy
SPRINT-PP83.5%
Distillation83.0%
Replay76.5%
Baseline14.6%

SPRINT-PP achieved the strongest tested accuracy while remaining privacy-safe and storing no raw historical PII.

Findings

Answers to the four research questions

RQ1Supervised model families

RoBERTa-large is the strongest supervised model, with a strict F1 of 0.991 and perfect output validity. FLAN-T5-base is the closest generative alternative at 0.987, also with perfect validity. The decoder-only models fall behind mainly on recall: when they produce a valid prediction it is usually right, but they recover fewer entities overall. Validity matters in a privacy pipeline, because a model whose outputs are often structurally unusable is not a dependable extractor however good its best answers are, so RoBERTa-large became the backbone for the continual-learning work.

The chart above ranks all seven models by strict F1 and shows each one's validity.

RQ2Fine-tuning versus prompting

For decoder-only models, in-context prompting is a strong alternative that needs no training. The best key-value three-shot prompts beat the matching fine-tuned baselines on the reported subsets, most clearly for Qwen2.5-7B (0.778 against 0.400). The output format matters as much as the number of examples: key-value is consistently stronger than JSON on quality, validity, and speed, at about 1.76 s/query against 26.8 s/query for Qwen. Three shots is the best operating point, because five shots cost more time without consistent gains.

Decoder-only models: fine-tuned versus in-context prompting, strict F1 and output validity. Fine-tuning and prompting were reported on different-sized subsets, so read the comparison as indicative.
ConfigurationStrict F1Validity
Qwen2.5-7B · fine-tuned0.4000.543
Qwen2.5-7B · JSON, 3-shot0.7090.906
Qwen2.5-7B · key-value, 3-shot0.7780.950
Llama3.1-8B · fine-tuned0.6900.814
Llama3.1-8B · JSON, 5-shot0.7050.856
Llama3.1-8B · key-value, 3-shot0.7250.914

RQ3Forgetting and mitigation

Naive sequential fine-tuning collapses. Each new stage is learned well, but every earlier stage falls to an F1 of 0.000 in the stage-restricted view, so the model replaces what it knew instead of adding to it. SPRINT-PP is the strongest mitigation, reaching 83.5% final accuracy against 83.0% for distillation and 76.5% for replay. It stores no raw historical PII, and it trains in 25.6% less time than distillation.

The chart above compares the four update strategies on final accuracy.

RQ4Corpus design and synthetic augmentation

Corpus design is a major driver of extraction quality, not a preprocessing detail. With the same RoBERTa-large backbone and the same evaluation, training on public data alone reaches a strict F1 of 0.415, with recall limited to 0.307. Training on the full ASC-PIE corpus reaches 0.992, with recall at 0.995, a gain of +0.577. Strict and normalised F1 are nearly identical in both settings, so the gain reflects real coverage of entities, not formatting. The gains are broad but uneven, and largest where public data alone gave the least coverage.

Public-only vs. full ASC-PIE · RoBERTa-large, strict scores
Public-only vs. full ASC-PIE · RoBERTa-large, strict scores
LabelPublic-onlyFull ASC-PIE
Precision0.6400.989
Recall0.3070.995
F10.4150.992
Selected per-type strict F1, public-only versus full ASC-PIE training, on the matched comparison subset.
PII typePublic-onlyFull ASC-PIEGain
JOBTITLE0.00000.9982+0.9982
EMAIL0.07860.9940+0.9154
SIN0.22360.9980+0.7744
CITY0.28260.9995+0.7169
NAME0.27160.9824+0.7108
STATE0.28910.9986+0.7095
LOCATION0.44770.9102+0.4625
PHONENUM0.73400.9938+0.2598
COUNTRY0.80001.0000+0.2000
CREDITCARD0.97860.9911+0.0125
DRIVER-LICENCE0.97440.9756+0.0012
TIME1.00000.8333-0.1667
TAXNUM1.00000.7500-0.2500

TIME and TAXNUM are slightly lower under full ASC-PIE, but the matched subset holds only 5 TIME mentions and 3 TAXNUM mentions, so those two differences should not be over-read.

Research framework

A comparable path through privacy datasets

Four public corpora and a synthetic component are mapped to a shared 19-type schema, split once, exported in three formats, and scored through one canonical evaluator across three model families. Relationships: Data sources to Schema mapping; Schema mapping to Fixed splits; Fixed splits to Synced exports; Synced exports to Model families; Model families to Canonical parser; Canonical parser to Shared evaluator.

PII-aware NER evaluation

ASC-PIE evaluates named-entity recognition for personally identifiable information across privacy-focused datasets.

Research data pipeline

The ASC-PIE experiments prepare real and synthetic privacy datasets for NER training and evaluation.

PII label standardization

The research pipeline maps differing PII label schemes into a shared representation for comparison.

ASC-PIE experiment stack

ASC-PIE experiments use Python, PyTorch, Hugging Face Transformers, scikit-learn, and seqeval.

Research links

Explore the research