Graciotti, Arianna
(2026)
Long-tail knowledge extraction from historical documents, [Dissertation thesis], Alma Mater Studiorum Università di Bologna.
Dottorato di ricerca in
Computer science and engineering, 38 Ciclo.
Documenti full-text disponibili:
Abstract
Knowledge Extraction (KE) transforms unstructured text into machine-readable representations, enabling automatic Knowledge Graph (KG) generation. Modern KE systems use Large Language Models (LLMs) as core components, leveraging knowledge stored in their parameters through pre-training. LLMs struggle with long-tail knowledge (infrequent topics in pre-training corpora), producing unreliable outputs for niche queries. While KGs structure knowledge transparently and support deterministic SPARQL queries, their coverage remains incomplete due to maintenance complexity, with long-tail entities poorly represented or absent. This thesis addresses long-tail KE through multilingual historical documents: English and Italian music magazines from the 19th and early 20th centuries. Historical documents exemplify long-tail challenges, as historical entities and events are sparsely represented in LLMs' pre-training datasets and KGs. The work first develops KG generation methods intended to generalise across diverse domains, exposing limitations when applied to historical documents. This thesis develops human-curated benchmarks from historical documents that are more challenging than existing datasets, with no temporal or topical overlap. These benchmarks reveal that KE systems rely excessively on statistical priors, returning popular but implausible outputs from temporal and typological perspectives. To address these issues, approaches that combine neural models with heuristics based on information in KGs are evaluated on these and other benchmarks built from historical documents. These approaches mitigate bias towards prevalent knowledge and enhance long-tail KE performance. Entities absent from KGs (NIL entities) create blind spots for KE systems. In historical documents, NIL entities are numerous and include entities prominent in their time but lost to history. Many lost person entities are women whose existence has faced systematic underrepresentation. A methodology for enriching KGs with absent entities and their background knowledge is proposed and evaluated on a novel gold-standard KG built for NIL person entities occurring in historical documents, enabling future systems to leverage historical knowledge more effectively and equitably.
Abstract
Knowledge Extraction (KE) transforms unstructured text into machine-readable representations, enabling automatic Knowledge Graph (KG) generation. Modern KE systems use Large Language Models (LLMs) as core components, leveraging knowledge stored in their parameters through pre-training. LLMs struggle with long-tail knowledge (infrequent topics in pre-training corpora), producing unreliable outputs for niche queries. While KGs structure knowledge transparently and support deterministic SPARQL queries, their coverage remains incomplete due to maintenance complexity, with long-tail entities poorly represented or absent. This thesis addresses long-tail KE through multilingual historical documents: English and Italian music magazines from the 19th and early 20th centuries. Historical documents exemplify long-tail challenges, as historical entities and events are sparsely represented in LLMs' pre-training datasets and KGs. The work first develops KG generation methods intended to generalise across diverse domains, exposing limitations when applied to historical documents. This thesis develops human-curated benchmarks from historical documents that are more challenging than existing datasets, with no temporal or topical overlap. These benchmarks reveal that KE systems rely excessively on statistical priors, returning popular but implausible outputs from temporal and typological perspectives. To address these issues, approaches that combine neural models with heuristics based on information in KGs are evaluated on these and other benchmarks built from historical documents. These approaches mitigate bias towards prevalent knowledge and enhance long-tail KE performance. Entities absent from KGs (NIL entities) create blind spots for KE systems. In historical documents, NIL entities are numerous and include entities prominent in their time but lost to history. Many lost person entities are women whose existence has faced systematic underrepresentation. A methodology for enriching KGs with absent entities and their background knowledge is proposed and evaluated on a novel gold-standard KG built for NIL person entities occurring in historical documents, enabling future systems to leverage historical knowledge more effectively and equitably.
Tipologia del documento
Tesi di dottorato
Autore
Graciotti, Arianna
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Knowledge Extraction, Large Language Models, Knowledge Graphs, Long-tail Knowledge, AI for Cultural Heritage, Natural Language Processing, Semantic Web, Digital Humanities
Data di discussione
26 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Graciotti, Arianna
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Knowledge Extraction, Large Language Models, Knowledge Graphs, Long-tail Knowledge, AI for Cultural Heritage, Natural Language Processing, Semantic Web, Digital Humanities
Data di discussione
26 Marzo 2026
URI
Statistica sui download
Gestione del documento: