Thesis title: Privacy in Large Language Models: A Unified Study of Threats, Defenses, and Clinical Applications
Artificial Intelligence has progressively shifted from hand-crafted, rule-based systems to data-driven
models that learn general patterns from large corpora. In Natural Language Processing (NLP), this
evolution has culminated in Large Language Models (LLMs): high-capacity neural models that,
by combining large-scale pretraining with modern adaptation techniques, can follow instructions,
generalize across tasks, and generate fluent text in many settings. In just a few years, LLMs
have become a widely used interface for interacting with information and services, with chat-style
assistants turning language into a practical programming and control medium for end users.
This rapid adoption has made privacy a central concern. LLMs are trained on vast amounts of
text, may be adapted to domain-specific corpora, and are routinely prompted with user-provided
context. These choices create multiple avenues for unintended disclosure: models can memorize
rare or sensitive sequences and may reproduce identifying details during generation. Adversaries
can also probe models through inference-time interactions to extract information. At the same
time, mitigating privacy risks is not merely a matter of data removal: protective interventions
often interact with model quality, domain utility, and the practical constraints of deployment in
resource-limited or regulated environments.
This thesis studies privacy-preserving LLMs by combining a unifying perspective on threats with
empirical investigations of defenses, with particular emphasis on Differential Privacy (DP) as a prin-
cipled tool for limiting the influence of any single training record on a model’s behavior. DP offers
a formal notion of protection with explicit parameters and composition properties, but its practical
application requires careful modeling choices and accounting procedures that can materially affect
both privacy and utility.
The work is organized around three complementary directions. First, we provide a structured
survey of privacy risks and mitigation strategies for LLMs, focusing in depth on DP-based ap-
proaches and on recurring methodological issues in measuring privacy-utility trade-offs. Second,
motivated by the high stakes of healthcare text, we investigate LLM-based de-identification for
clinical notes in multilingual contexts through case studies on Italian and Dutch clinical narratives,
where we comparatively evaluate manual anonymization, NER-based pipelines, and DP-based tech-
niques under a consistent evaluation framework. Third, we examine privacy risks that arise specif-
ically from prompting by introducing and studying a membership inference attack that targets in-
context demonstrations, where the adversary aims to infer whether a candidate record was included
among the examples used to condition generation. In addition, we document a DP-ICL pipeline
for synthetic clinical text generation developed in the context of DataTools4Heart, as a building
block towards CardioSynth, a planned open synthetic cardiology dataset containing structured and
unstructured data.
Overall, the thesis aims to clarify which privacy threats are most salient for contemporary LLM
workflows, to provide evidence on the strengths and limitations of existing defenses, particularly
DP in realistic pipelines, and to contribute practical methodologies for privacy-aware language
technologies in sensitive domains.
Keywords: Large Language Models; Privacy; Differential Privacy; Clinical NLP; De-identification;
Membership Inference; In-context Learning.