Navigating the Black Box: A Scoping Review of Large Language Models and Explainable Artificial Intelligence in Public Financial Institutions

Authors

  • Usman Yunusa Labaran Baze University https://orcid.org/0000-0002-2196-9588
  • Rufai Aliyu Yauri Baze University
  • Charles Isah Saidu Baze University
  • Amina F. Shehu Baze University
  • Rislana A. Kanya Cosmopolitan University
  • Ahmad N. Tambaya Cyber Safety Alliance (CSA)

Keywords:

Large Language Models, Explainable Artificial Intelligence, Public Financial Institutions, Sovereign Wealth Funds, Retrieval-Augmented Generation, Scoping Review

Abstract

Public financial institutions (PFIs) are implementing general-purpose large language models (LLMs) much faster than the evidence supports. Specifically, sovereign wealth funds (SWFs) operating under fiduciary duties and transparency obligations set out in the Santiago Principles are unable to quickly adapt. This is primarily because generic or unexplained model outputs cannot satisfy these requirements. Objective: This review maps the current literature on adapting LLMs for institutional finance and the explainability of their outputs. The aim is to identify the evidence gaps critical for SWF deployment. Methods: We conducted a corpus-bounded scoping review using a 496-record library assembled during a targeted evaluation of an African SWF. After deduplication, 470 unique records were screened against predefined criteria. This yielded in 230 included studies that we synthesised across three core themes: LLM adaptation, explainable artificial intelligence (xAI), and deployment risk and evaluation. Key Findings: The final synthesis included 77 studies on adaptation, 108 on explainability, and 74 on deployment risk and evaluation with some studies spanning multiple themes. In terms of traceability, adaptation splits into two fundamentally opposing approaches as parametric fine-tuning successfully injects institutional knowledge but obscures its provenance. Whereas retrieval-augmented generation (RAG) exposes its sources, but it offers no guarantee that the model uses them. xAI techniques developed for classifier architectures degrade significantly when applied to free-text generative output. Notably, no included study evaluates an LLM deployed within an SWF. Conclusion: High performance on generic benchmarks does not certify a model for institutional deployment, and African SWFs remain an empirical void. This review establishes the research agenda required to address this gap.

DOI: https://doi.org/10.5281/zenodo.22952993 

Downloads

Published

2026-09-25