Skip to main content

Data science in central banking: applications and tools

Type
Publication
Series
IFC Bulletin 59
Date Published
10 October 2023
Sources
IFC

The Irving Fisher Committee on Central Bank Statistics (IFC) periodically organises workshops on “Data science in central banking” with a diverse audience of practitioners and technicians. The most recent one took place in 2022 and focused on the broad spectrum of data science applications/tools used in central banks.

The concept of data science refers to the study of data and therefore includes the various techniques for extracting insights from them. Yet data science is fundamentally different from traditional data analysis, as it typically applies to large, complex and/or unstructured information sets.

A key factor supporting the development of central banks’ data science projects in recent years has been the sheer volume and complexity of financial data available in today’s societies. This requires more sophisticated techniques for data management and analysis, a trend reinforced by the new opportunities opened up by artificial intelligence (AI) and machine learning (ML). Another factor has been the greater focus on real-time, evidence-based policymaking, which requires authorities to rely on better analytical and forecasting capacities to support their decisions. Additionally, advances in statistical computing infrastructure and enhanced training have enhanced the data skills of official sector staff.

These elements have accelerated efforts to advance data science, helping central banks to quickly adapt to the swiftly evolving financial landscape. In this endeavour, the role of data scientists lies at the intersection of three areas: information technology (IT); mathematical and statistical methods; and business, or “subject-matter” expertise. 

From the outset, IT has absorbed a great deal of attention and resources. Central banks are increasingly aware that a modern IT architecture is crucial in reliably and securely dealing with data. A key objective is to facilitate access to a large and diverse range of sources as well as relevant IT software and tools in a user-friendly way. But implementing such an IT architecture can be challenging, calling as it does for careful implementation, clear governance frameworks, and the application of common standards. The emphasis is on adopting advanced IT tools and engineering practices – including cloud computing, software containers, automation tools, and continuous software integration and delivery pipelines. In particular, there has been a growing interest among central banks on using and producing software that can be shared as open source, either with their peers or with the general public. Such an open source software (OSS) strategy can be instrumental for honing their own IT development and strengthening their data science capabilities.

Once the IT infrastructure is able to support the development and deployment of data science applications, the focus is on performing the various mathematical and statistical operations that are needed to deal with the raw data. Data scientists need not only to access very large and complex information sets, but also to compile statistics via multiple sequential tasks (signal extraction, quality management, dissemination) before using them to extract relevant insights. Many different AI techniques can be used for this purpose, including for conducting textual analysis, reflecting the increasing opportunities offered by natural language processing (NLP) tools and large language models (LLMs).

A third lesson is that data science projects require a good understanding of the business cases and therefore a close cooperation with subject-matter experts. One obvious reason is that economic indicators such as GDP are more than just numbers: analysing them calls for an understanding of the way the statistics have been compiled as well as the complex factors that drive them – say, fiscal policy or geopolitical tensions – and their real-world implications. Moreover, this expertise is essential to support informed policy decisions: in particular, translating data insights into actionable recommendations for central banks cannot be communicated as a “black box” and requires transparent explanations to the various stakeholders involved – from other authorities to the general public. This is even more important with data science applications that may need to be adapted when used to answer economic questions, for instance when analysing causal relationships. Finally, business area expertise can help central banks prioritise effectively between concurrent data science initiatives especially in view of resources constraints.


The views expressed in this publication are those of the authors and do not necessarily represent the official views of the Committee, its members or the BIS.