AI/ML · Research · 5 min read

ASC-PIE

Every PII dataset spoke its own language. I built a shared one, then put seven models through the same exam.

I developed this PII-aware named-entity recognition corpus and evaluation framework for my M.A. thesis, harmonizing four public corpora plus my own synthetic augmentation under a shared 19-type schema.

My role
I designed and built the research pipeline: corpus standardization, the shared schema, the evaluation framework, and the reproducible experiment artifacts, supervised by Prof. Marin Litoiu at York University’s CERAS Lab.
Context
M.A. thesis, York University (CERAS Lab)
Period
Sep 2024 – Apr 2026
Status
Thesis · awarded 2026

Thesis · awarded 2026

  • 333,109examples in the corpus
  • 19canonical PII types
  • 99.1%best strict F1RoBERTa-large, held-out subset
  • +0.577F1 gained from corpus designsame backbone, public-only vs full ASC-PIE
  • Python
  • PyTorch
  • Hugging Face Transformers
  • Hugging Face Datasets
  • scikit-learn
  • Jupyter / Colab

Roles: AI/ML Engineer · Software Engineer · TA / Instructor.

The problem

People paste personal details into prompts, emails, chat logs, and support tickets constantly, and organizations are expected to find that PII before it causes harm. Regulations such as GDPR, PIPEDA, and CCPA all assume a system can locate personally identifying information reliably.

The trouble was never a shortage of PII datasets. It was that public PII datasets each spoke a different language: different label sets, different formats, different annotation conventions. Comparing an encoder model trained on one corpus against a generative model trained on another told you almost nothing about which approach actually extracts PII better, because the comparison itself was not fair.

Who it was for

Anyone choosing a model for PII extraction: a privacy team weighing a fast token classifier against a flexible generative one, or a researcher asking whether prompting can replace fine-tuning. I built ASC-PIE around four research questions: how model families compare under one protocol, how in-context prompting compares with fine-tuning, how badly a model forgets when its label space grows, and how much corpus design itself changes the outcome.

My role

I designed and built the ASC-PIE research pipeline for my M.A. thesis at York University’s CERAS Lab, supervised by Prof. Marin Litoiu: the corpus itself, the shared 19-type schema, the evaluation framework that scores every model family the same way, and the reproducible artifacts behind every result on this page. My thesis is completed, and the degree was officially awarded in 2026, conferred on 24 Jul 2026.

What I built

  • Four public corpora, unified

    AI4Privacy, Gretel, CoNLL++, and PII-DD are remapped onto one schema, then extended with my own synthetic augmentation for rare PII types.

  • A shared 19-type schema

    Every corpus is normalized to the same 19 canonical PII types, from NAME and EMAIL through SIN, PASSPORT, and ORGANIZATION.

  • Synchronized exports for every model family

    The same fixed splits are exported as IOB2 tags, key-value text, and JSON, so encoder, encoder-decoder, and decoder-only models train and evaluate on matched data.

  • One evaluator for every model family

    IOB2 tags, key-value text, and JSON are all normalized to the same (type, text) pairs before scoring, so strict F1, normalized F1, and validity mean the same thing regardless of model family.

  • Reproducible experiment artifacts

    Dataset manifests, configurations, per-type reports, and prediction traces make every result traceable back to the run that produced it.

Underneath the schema sits a chip cloud of the 19 canonical PII types, grouped the way a privacy reviewer would read them: identity documents (NAME, GENDER, SIN, TAXNUM, NATIONALID, PASSPORT, DRIVER-LICENCE), contact details (EMAIL, PHONENUM), financial data (CREDITCARD), geography (LOCATION, CITY, STATE, COUNTRY, ZIPCODE), and context (DATE, TIME, JOBTITLE, ORGANIZATION).

How it works

Every corpus is remapped to one schema, split once, exported in synchronized formats, and scored by one evaluator, regardless of which model family reads it.

Four public corpora and a synthetic component are mapped to a shared 19-type schema, split once, exported in three formats, and scored through one canonical evaluator across three model families. Relationships: Data sources to Schema mapping; Schema mapping to Fixed splits; Fixed splits to Synced exports; Synced exports to Model families; Model families to Canonical parser; Canonical parser to Shared evaluator.

Decisions that mattered

  • I chose a realistic long-tail label distribution over an artificially balanced benchmark because production PII inventories are lopsided too. NAME appears in 445K spans and COUNTRY in only 914, and smoothing that away would have hidden exactly the categories that are hardest to extract.

  • I chose strict entity-text matching as the primary metric because a partial or fuzzy match still leaves a PII span unresolved for a downstream redaction or auditing pipeline.

  • I chose fixed, disjoint train, validation, and test splits because every model family, and every later continual-learning stage, needed to be compared on the same unseen data.

Hard problems I solved

  • ProblemThe four public corpora used different labels, tag sets, and annotation conventions for what was often the same underlying PII.

    FixI built a deterministic schema-mapping step that remaps every corpus's labels onto the shared 19-type schema, repairs IOB2 tag inconsistencies, deduplicates, and routes anything outside the schema to O.

  • ProblemGenerative models sometimes returned text that could not be parsed back into (type, text) pairs at all.

    FixI normalized every model family's output to the same canonical representation before scoring, and made validity, whether an output is even structurally usable, a first-class metric alongside F1.

  • ProblemRare, highly structured PII types such as PASSPORT and TAXNUM had very little public support.

    FixI added a synthetic augmentation component, Luhn-valid credit card numbers, controlled templates, and Fake Name Generator records, to strengthen the training signal for exactly the categories the public corpora under-cover.

Tech stack

Tooling & evaluation
  • seqeval: Entity-level F1 scoring
  • Jupyter / Colab: Corpus notebooks and experiment runs

Outcome

RoBERTa-large came out on top of the model-family comparison, with FLAN-T5-base close behind as the strongest generative alternative. The decoder-only models, Llama3.1-8B and Qwen2.5-7B, were mainly limited by recall rather than precision, and BART-base struggled with both validity and F1.

Strict F1 and validity by model, on a held-out subset.
ModelFamilyValidityStrict F1
RoBERTa-largeEncoder100.0%99.1%
FLAN-T5-baseEncoder-decoder100.0%98.7%
ModernBERT-largeEncoder100.0%84.6%
BERT-large-casedEncoder100.0%80.7%
Llama3.1-8BDecoder-only81.4%69.0%
Qwen2.5-7BDecoder-only54.3%40.0%
BART-baseEncoder-decoder56.1%34.8%

For the decoder-only models, in-context prompting beat fine-tuning outright. The best 3-shot key-value prompt outperformed the fine-tuned baseline for both models tested, and it did so far faster than the equivalent JSON prompt: about 1.76 s/query for key-value versus about 26.8 s/query for JSON on the same model.

Fine-tuned strict F1 versus the best 3-shot key-value in-context result, same model.
ModelFine-tuned strict F1Best 3-shot key-value strict F1
Qwen2.5-7B40.0%77.8%
Llama3.1-8B69.0%72.5%

The largest single result, though, came from the corpus itself rather than the model. Training the same RoBERTa-large backbone on the public corpora alone reached 41.5% strict F1, limited mainly by recall. Training it on the full ASC-PIE corpus, public corpora plus my own synthetic augmentation, reached 99.2% strict F1 under the identical protocol.

RoBERTa-large trained on public corpora only versus the full ASC-PIE corpus, same backbone and evaluation protocol.
Training corpusPrecisionRecallStrict F1
Public corpora only64.0%30.7%41.5%
Full ASC-PIE98.9%99.5%99.2%

That is a gain of +0.577 strict F1 on the same backbone and the same evaluation protocol, and the gain was uneven in a telling way: the largest jumps landed on categories such as JOBTITLE, EMAIL, and SIN, where public-only training had almost nothing to work with, while already near-ceiling structured types such as CREDITCARD barely moved.

Selected per-type strict F1, public corpora only versus full ASC-PIE.
PII typePublic-onlyFull ASC-PIE
JOBTITLE0.0%99.8%
EMAIL7.9%99.4%
SIN22.4%99.8%
NAME27.2%98.2%
COUNTRY80.0%100.0%
CREDITCARD97.9%99.1%

The same corpus and evaluator also underpin a paper, ASC-PIE and SPRINT-PP: Evaluating Privacy-Safe Continual Learning for PII Extraction, accepted to IEEE CASCON 2026, which walks through the continual-learning half of this work in full on its own page.

What I learned

Corpus design moved results more than changing the model did.

Two RoBERTa-large runs, identical architecture and identical protocol, landed 0.577 strict F1 apart because one saw a properly unified, augmented corpus and the other did not. That is a larger swing than moving between any two of the model families compared here, and it is the reason ASC-PIE treats corpus construction as a first-class contribution rather than a preprocessing footnote.

What I’d do next

The thesis’s own future directions apply here directly: extending ASC-PIE to multilingual and cross-lingual settings under the same unified schema, broadening the model-family study with additional open-weight backbones, and running stronger drift tests to see how the schema and corpus hold up outside this thesis’s own splits.

The research record has the full results, charts, and the answers to the four research questions. The thesis is catalogued on YorkSpace and reachable through its permanent handle. The corpus tooling, training scripts, and evaluation harness behind every number on this page live in the thesis experiments repository.