AI/ML · Research · 5 min read

SPRINT-PP

My model forgot everything it learned. The fix had to work without keeping anyone’s personal data.

I designed and built SPRINT-PP, a privacy-safe continual-learning method that lets a PII-extraction model learn new entity types across stages without storing any raw historical text.

My role
I designed the three-stage protocol, implemented all four compared strategies, developed the SPRINT-PP method, and ran the experiments as part of my M.A. thesis.
Context
M.A. thesis research, York University (CERAS Lab), supervised by Prof. Marin Litoiu
Status
Accepted · CASCON 2026

Accepted · CASCON 2026

  • 83.5%ACC, best of four strategiesSPRINT-PP, shared RoBERTa-large backbone
  • 0raw historical examples stored
  • 25.6%less training time than distillation
  • 14.6%naive sequential fine-tuning ACC
  • Python
  • PyTorch
  • RoBERTa-large
  • Hugging Face Transformers

Roles: AI/ML Engineer.

Paper: “ASC-PIE and SPRINT-PP: Evaluating Privacy-Safe Continual Learning for PII Extraction” — accepted to IEEE CASCON 2026 (Toronto, 10–12 Nov 2026); the method is also part of my M.A. thesis. No author list yet.

The problem

PII requirements do not stay fixed. A system built to find one set of categories today may need to find more next year, as a regulator adds a category or a product adds a new field. The obvious fix is to keep fine-tuning the same model on the new types.

I ran that obvious fix as a baseline, staged over three rounds of new PII types on the same RoBERTa-large backbone. Each stage was learned well on its own, F1 above 96.3% on the fresh types every time, but the previous stage’s F1 collapsed to 0% the moment the next stage began. The model was not accumulating knowledge; it was replacing it.

The reason is what the underlying research calls the poisoned-gradient problem: later-stage text still contains old entities, now relabelled O, and training on those tokens actively teaches the model to stop recognizing what it already knew. The obvious mitigation, replay, fixes this by storing old labelled examples so the model keeps seeing them, but those examples are personal data, and storing them is exactly the retention a privacy-aware system should avoid.

My role

I designed the three-stage protocol, implemented all four compared update strategies, developed the SPRINT-PP method itself, and ran every experiment behind this page as part of my M.A. thesis at York University’s CERAS Lab, supervised by Prof. Marin Litoiu.

What I built

I built the staged benchmark itself, not just the mitigation. Every stage adds new PII types on top of what came before, and I score each strategy two ways: a stage-restricted view that shows what a strategy learned, and a cumulative view that shows what it managed to keep usable once later stages arrived. Separating those two views is what makes it possible to see forgetting and new-type learning as two different numbers instead of one blurred score.

Every one of those scores is entity-level F1, not token accuracy, and that choice matters here. In PII extraction most tokens in any sentence are labelled O, meaning “not an entity”. A model that has forgotten an entire type can still label the large majority of tokens correctly, because it simply calls them O, so token accuracy stays respectable while the model can no longer find a single entity of that type. Entity-level F1 only gives credit when the whole entity is found, so it shows forgetting the way a user would experience it: the thing you needed extracted is missing.

The two views also read differently across strategies, which is why I kept both. In the stage-restricted view, replay’s score on the first stage slips as later stages arrive, while SPRINT-PP holds it almost level. In the cumulative view, where old and new types are scored together, that gap widens, because a model that has quietly lost one type drags down every sentence containing it. Reading the two together is what separates a method that learns new types from one that also keeps the old ones usable.

  • A three-stage class-incremental protocol

    Each stage adds new PII types on top of the ones already learned, while the train, validation, and test splits stay fixed and disjoint throughout.

  • Two evaluation views

    A stage-restricted view scores each stage in isolation to show what was learned; a cumulative view scores old and new labels together to show what stayed usable.

  • Four compared update strategies

    Baseline sequential fine-tuning, replay, distillation, and SPRINT-PP all share the same encoder backbone, data splits, and stage design, so only the update mechanism differs.

How it works

Each new stage’s data is triaged against the frozen previous model, corrected selectively instead of broadly replayed, and anchored so the classifier keeps its old rows.

Each new stage is checked for suspected old entities by the frozen previous model, corrected with soft labels and a prototype memory instead of raw replay, and anchored so the classifier keeps its old rows. Relationships: Stage t data to Student model; Teacher model to Triage; Triage to Selective correction (suspected old entity); Prototype memory to Selective correction (no raw text stored); Selective correction to Student model; Head anchoring to Student model.

Tech stack

Outcome

SPRINT-PP finished ahead of every other strategy on ACC and BWT, and it did so while never storing a raw historical example. Only replay stores raw examples, and replay still forgot more than either privacy-safe alternative.

Continual-learning summary across four update strategies, shared RoBERTa-large backbone.
StrategyACC ↑BWT ↑Forgetting ↓Intransigence ↓Stores raw PII?
SPRINT-PP83.5%−0.20750.20750.0156No
Distillation83.0%−0.21110.21110.0193No
Replay76.5%−0.27260.27260.0278Yes
Baseline14.6%−0.65060.65060.0202No

SPRINT-PP was also 25.6% faster to train than distillation, on the same backbone and the same staged protocol.

Stage-restricted F1: each stage evaluated on its own

Baseline
Baseline stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.960.000.00
Stage 10.000.980.00
Stage 20.000.000.99
Replay
Replay stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.970.000.00
Stage 10.950.970.00
Stage 20.870.990.98
Distillation
Distillation stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.970.000.00
Stage 10.970.980.00
Stage 20.960.981.00
SPRINT-PP
SPRINT-PP stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.970.000.00
Stage 10.970.990.00
Stage 20.970.991.00

Read down the diagonal and every strategy looks fine: each stage is learned well in isolation. Read the lower-left corner and the difference shows up: the baseline’s early stages read 0.000 once later stages arrive, while SPRINT-PP keeps stage 0 above 0.96 all the way through stage 2.

Cumulative F1: old and new labels scored together

Replay
Replay stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.970.000.00
Stage 10.830.950.00
Stage 20.630.740.92
Distillation
Distillation stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.970.000.00
Stage 10.850.970.00
Stage 20.700.810.98
SPRINT-PP
SPRINT-PP stage-restricted F1
Stage 0Stage 1Stage 2
Stage 00.970.000.00
Stage 10.860.970.00
Stage 20.710.820.98

The baseline is left out of the cumulative view above: only its final row is available from the underlying research, and I would rather show three strategies in full than estimate a fourth.

The honest limits: this is one backbone, three stages, English-only, and the privacy protection is operational, no raw replay, not a formal privacy guarantee.

What I learned

How a method treats data is part of its result.

Replay looks strong on paper if you only read the ACC column and ignore what it has to keep to get there. SPRINT-PP and distillation both stay ahead of it while storing nothing raw, which is the comparison that actually matters once the categories being learned are personal data in the first place. A continual-learning method for PII cannot be judged on accuracy alone; what it stores to reach that accuracy is part of the result, not a footnote to it.

What I’d do next

The underlying research’s own future work applies directly here: evaluating SPRINT-PP and other privacy-safe mitigation strategies under stronger domain drift, noisier supervision, and partially overlapping annotation policies, and testing it against other backbones and longer stage sequences than the three used in this thesis. Extending the comparison to more than three stages would also show whether the gap between SPRINT-PP and distillation holds, narrows, or widens as the label space keeps growing.

The continual-learning results here build directly on the corpus and shared evaluator described on the ASC-PIE research record, and the underlying thesis is catalogued on YorkSpace. The IEEE Xplore record will be added once the CASCON 2026 proceedings are published. The experiments behind it run from the thesis experiments repository.