Comparing Hallucination Detection Metrics for Multilingual Generation

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

While many hallucination detection techniques have been evaluated on English text, their effectiveness in multilingual contexts remains unknown. This paper assesses how well various factual hallucination detection metrics (lexical metrics like ROUGE and Named Entity Overlap, and Natural Language Inference (NLI)-based metrics) identify hallucinations in generated biographical summaries across languages. We compare how well automatic metrics correlate to each other and whether they agree with human judgments of factuality. Our analysis reveals that while the lexical metrics are ineffective, NLI-based metrics perform well, correlating with human annotations in many settings and often outperforming supervised models. However, NLI metrics are still limited, as they do not detect single-fact hallucinations well and fail for lower-resource languages. Therefore, our findings highlight the gaps in exisiting hallucination detection methods for non-English languages and motivate future research to develop more robust multilingual detection methods for LLM hallucinations.

Related collections

Author and article information

Journal

Publisher: arXiv

Publication date (Electronic): 2024

Publication date Submitted: 16 February 2024

Publication date Updated: 19 February 2024

Publication date Submitted: 16 June 2024

Publication date Updated: 18 June 2024

Publication date Available: February 2024

Article

DOI: 10.48550/ARXIV.2402.10496

SO-VID: 76c4d40e-308f-4527-8d70-f17319741a05

License:

Creative Commons Attribution 4.0 International

History

Keywords: FOS: Computer and information sciences,Artificial Intelligence (cs.AI),Computation and Language (cs.CL)

Data availability:

Keywords: FOS: Computer and information sciences, Artificial Intelligence (cs.AI), Computation and Language (cs.CL)

Comments

Comment on this article

scite_

Smart Citations

Citing PublicationsSupportingMentioningContrasting

View Citations

See how this article has been cited at scite.ai

scite shows how a scientific paper has been cited by providing the context of the citation, a classification describing whether it supports, mentions, or contrasts the cited claim, and a label indicating in which section the citation was made.