Malaspina, Francesco
(2026)
Concordance and influence of large language model treatment recommendations versus hematologist decision-making in relapsed/refractory diffuse large B-Cell lymphoma: the Ariadne's thread study, [Dissertation thesis], Alma Mater Studiorum Università di Bologna.
Dottorato di ricerca in
Oncologia, ematologia e patologia, 37 Ciclo.
Documenti full-text disponibili:
Abstract
Large language models (LLMs) achieve high scores on medical benchmarks, yet their clinical reasoning capabilities in complex, real-world scenarios remain poorly characterized. This thesis evaluates LLM performance in treatment selection for relapsed/refractory diffuse large B-cell lymphoma (R/R DLBCL), a scenario requiring integration of incomplete evidence, regulatory constraints, and nuanced judgment across multiple treatment options.
In this study, 50 synthetic clinical cases were independently assessed by one lymphoma subspecialist and one general hematologist. Experts first formulated treatment recommendations independently, then evaluated recommendations generated by “Ariadne,” a commercially available LLM guided by a custom prompt and a curated R/R DLBCL knowledge base. Primary outcomes included overall clinical acceptability and treatment recommendation modification rates; exploratory outcomes assessed inter-rater agreement, harm potential in rejected cases, confidence changes, and error pattern characterisation.
The LLM achieved 83.0% overall acceptability (95% CI: 74.5%–89.1%), with treatment concordance (same treatment category selected independently by LLM and expert) observed in 75% of evaluations. However, physician expertise fundamentally altered the human-AI interaction pattern. General hematologists were five times more likely to modify their treatment after LLM exposure than lymphoma subspecialists (32% vs 6%), and showed higher acceptability rates (90% vs 76%). Subspecialists rejected 8 recommendations that generalists accepted versus only 1 reverse case, and 5 of these 8 asymmetric cases (63%) involved treatment choices deemed excessively toxic or sub-optimal by subspecialists. In 4 of these 8 cases, generalists also modified their treatment toward the LLM recommendation, including 2 with harm potential.
This reveals an expertise-dependent deployment paradox: the clinician population most behaviourally responsive to LLM recommendations is also least equipped to detect errors that carry genuine harm potential. Clinical deployment requires expertise-stratified safeguards, and evaluation frameworks must move beyond aggregate accuracy metrics to assess error reproducibility, harm potential, and safety across diverse user populations.
Abstract
Large language models (LLMs) achieve high scores on medical benchmarks, yet their clinical reasoning capabilities in complex, real-world scenarios remain poorly characterized. This thesis evaluates LLM performance in treatment selection for relapsed/refractory diffuse large B-cell lymphoma (R/R DLBCL), a scenario requiring integration of incomplete evidence, regulatory constraints, and nuanced judgment across multiple treatment options.
In this study, 50 synthetic clinical cases were independently assessed by one lymphoma subspecialist and one general hematologist. Experts first formulated treatment recommendations independently, then evaluated recommendations generated by “Ariadne,” a commercially available LLM guided by a custom prompt and a curated R/R DLBCL knowledge base. Primary outcomes included overall clinical acceptability and treatment recommendation modification rates; exploratory outcomes assessed inter-rater agreement, harm potential in rejected cases, confidence changes, and error pattern characterisation.
The LLM achieved 83.0% overall acceptability (95% CI: 74.5%–89.1%), with treatment concordance (same treatment category selected independently by LLM and expert) observed in 75% of evaluations. However, physician expertise fundamentally altered the human-AI interaction pattern. General hematologists were five times more likely to modify their treatment after LLM exposure than lymphoma subspecialists (32% vs 6%), and showed higher acceptability rates (90% vs 76%). Subspecialists rejected 8 recommendations that generalists accepted versus only 1 reverse case, and 5 of these 8 asymmetric cases (63%) involved treatment choices deemed excessively toxic or sub-optimal by subspecialists. In 4 of these 8 cases, generalists also modified their treatment toward the LLM recommendation, including 2 with harm potential.
This reveals an expertise-dependent deployment paradox: the clinician population most behaviourally responsive to LLM recommendations is also least equipped to detect errors that carry genuine harm potential. Clinical deployment requires expertise-stratified safeguards, and evaluation frameworks must move beyond aggregate accuracy metrics to assess error reproducibility, harm potential, and safety across diverse user populations.
Tipologia del documento
Tesi di dottorato
Autore
Malaspina, Francesco
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
37
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Hematology, large B cell lymphoma treatment, artificial intelligence, large language model, prompt engineering, behavioural influence, domain expertise, health benchmark
Data di discussione
1 Aprile 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Malaspina, Francesco
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
37
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Hematology, large B cell lymphoma treatment, artificial intelligence, large language model, prompt engineering, behavioural influence, domain expertise, health benchmark
Data di discussione
1 Aprile 2026
URI
Statistica sui download
Gestione del documento: