Research series

Training Data Attribution for NLP and LLM Research

A thesis- and interview-oriented path through attribution units, utilities, Shapley values, ablation, causal evidence, instance-level methods, scaling, uncertainty, validation, and research extensions.

How to use this series

The goal is not to sound like a generic tutorial. The goal is to make my thesis explainable, defensible, and expandable.

Interview readiness

Each note ends with a spoken-style answer that can be adapted for project interviews and PhD discussions.

Research precision

The series separates attribution units, utilities, estimators, interventions, and uncertainty so the claims stay precise.

PhD direction

The later notes point toward hierarchical attribution, intervention validation, and LLM factuality/style attribution.

Start by defining what attribution is allowed to claim.

A strong attribution project begins with clear units, clear utilities, and careful claims.