Skip to main navigation Skip to search Skip to main content

IRLBench: A multi-modal, culturally grounded, parallel Irish-English benchmark for open-ended LLM reasoning evaluation

Research output: Chapter in Book/Report/Conference proceedingsConference proceedingpeer-review

Abstract

Recent advances in Large Language Models (LLMs) have demonstrated promising capabilities, yet their performance in multilingual and low-resource settings remains modest. Existing benchmarks often exhibit cultural bias, restrict evaluation to text-only, rely on multiple-choice formats, and, more importantly, are ineffectual for extremely low-resource languages. To address these gaps, we introduce IRLBench, presented in parallel English and Irish, which is considered definitely endangered by UNESCO. Our benchmark consists of 12 representative subjects developed from the 2024 Irish Leaving Certificate exam, enabling fine-grained analysis of model capabilities across domains. By framing the task as long-form generation and leveraging the official marking scheme, it supports not only a comprehensive evaluation of correctness but also language fidelity. Our extensive experiments of leading closed-source and open-source LLMs reveal a persistent performance gap between English and Irish, in which models produce valid Irish responses less than 80% of the time, and answer correctly 55.8% of the time compared to 76.2% in English for the best-performing model. With Irish as the case study, our work exposes systemic weaknesses in today's multilingual LLMs and provides a rigorous benchmark for evaluating true multilingual capabilities. We release IRLBench and an accompanying evaluation codebase to enable future research on robust, culturally aware multilingual AI development.

Original languageEnglish
Title of host publicationKDD 2026 - Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
PublisherAssociation for Computing Machinery
Pages2794-2805
Number of pages12
ISBN (Electronic)9798400722585
DOIs
Publication statusPublished - 20 Apr 2026
Event32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026 - Jeju Island, Korea, Republic of
Duration: 9 Aug 202613 Aug 2026

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Volume1-A
ISSN (Print)2154-817X

Conference

Conference32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026
Country/TerritoryKorea, Republic of
CityJeju Island
Period9/08/2613/08/26

UCC Futures

  • Artificial Intelligence and Data Analytics

Keywords

  • Benchmarking
  • Extremely low-resource
  • Large language model
  • Large vision-language model
  • [ComputerScience]
  • [Insight Centre for Data Analytics]

Fingerprint

Dive into the research topics of 'IRLBench: A multi-modal, culturally grounded, parallel Irish-English benchmark for open-ended LLM reasoning evaluation'. Together they form a unique fingerprint.

Cite this