x
×
BACK TO SCHEDULE REGULAR

Session 3

When Artificial Intelligence Meets Official Statistics

3 June 2026
11:00 – 12:30
ŠIBENIK IV SHOW ON MAP

Session chair
Žaklina Čizmović
Head of Department, Quality, Statistical Standards and GIS Development, Croatian Bureau of Statistics (CBS)

Read more Read less As the Head of the Quality, Statistical Standards, and Geoinformation System Development Department, Žaklina Čizmović works on improving statistical standards and geoinformation systems. She has a master's degree in economics and is an expert in statistics and quality management. Žaklina has led important projects to align Croatian statistical practices with EU standards. She played a key role in setting up the Quality System in the Croatian Bureau of Statistics (CBS) and led the development of a system to improve the quality and management of statistical surveys. Her work ensured that statistical processes are efficient and reliable. At the European level, Žaklina contributed to updating the Code of Practice for European Statistics and helped improve the NACE Rev. 2.1 classification, making economic data more accurate and useful. She also focuses on geoinformation systems and subnational statistics, recognising the importance of location-based data in modern analysis. Her efforts have improved the quality and accuracy of statistical data, making it more useful for decision-making.

Presentation title
Leveraging AI to Enhance Engagement for Istat: Survey Respondent, User Support and Natural Language Access to SDMX Data
PPT PaperPRESENTATION PDF PaperPAPER

In 2025, the Italian National Institute of Statistics (Istat) introduced artificial intelligence across several of its user facing systems to enhance interaction with external audiences, both in the context of survey participation and in the dissemination of statistical information.

Read more Read less Within this broader digital innovation effort, two AI driven solutions were developed and deployed. The first concerns the Istat Contact Centre, where an AI based virtual assistant, through natural language interaction, was integrated to support respondents (providing guidance on survey participation) and data users (helping retrieve information). Trained on an internal, curated knowledge base and subjected to extensive pre release testing—including accuracy checks, UX assessments, and off topic controls—the system was authorized for deployment only after exceeding a 95% accuracy threshold. Since its introduction, it has managed more than 8,000 interactions, monitored through a dedicated dashboard that analyses conversation quality and supports continuous refinement through human in the loop supervision. The assistant is currently available on the Contact Centre webpage (for users and respondents) and on the Enterprises portal (for respondents), with planned extensions to additional entry points and to the telephone channel for automated call routing and classification. Continuous monitoring and iterative improvements ensure high reliability, reduce hallucinations, and strengthen compliance with institutional requirements related to accuracy, transparency, and user protection. The second initiative involved the integration of an AI assistant into IstatData, the Institute’s SDMX based data warehouse for official aggregate statistics. Instead of a generative conversational agent, Istat developed a “tool based AI assistant” comprising two components: a natural language search assistant and a copilot for customizing tables, charts, and maps. Operating exclusively on interface controls—never on the underlying data—the system transforms stringent constraints stemming from the EU AI Act, data integrity requirements, and institutional trust obligations into the backbone of a robust architecture. A multi stage validation process, including large scale simulated testing and a hierarchical retrieval algorithm, minimized semantic drift and virtually eliminated hallucinations, ensuring full methodological traceability and user control. The paper presents the methodological, infrastructural, and governance aspects of these initiatives, highlighting their regulatory and ethical implications. Together, they illustrate how AI can be responsibly integrated into public statistical services to improve usability, accessibility, and operational efficiency while preserving the rigor, transparency, and reliability that characterize official statistics.

Main author / Presenter
Salvatore Agrillo
The Italian National Institute of Statistics (Istat)

Read more Read less Salvatore Agrillo is a Computer Engineer, he works at ISTAT within the DCIT — the Directorate for Information Technology and Communication Systems. He is passionate about his profession and continuously deepen his expertise in artificial intelligence. His focus is on developing AI-driven solutions that simplify the work of both operators and end users. His job is also his greatest passion, so even in the free time he enjoys designing new tools, sometimes useful in everyday life.


CO-AUTHORS:

Carlo Boselli, The Italian National Institute of Statistics (Istat)
Alessio Cardacino, The Italian National Institute of Statistics (Istat)
Gabriella Fazzi, The Italian National Institute of Statistics (Istat)
Roberta Roncati, The Italian National Institute of Statistics (Istat)
Domenico Tucci, The Italian National Institute of Statistics (Istat)

Presentation title
Unlocking Legacy Statistics with Large Language Models
PPT PaperPRESENTATION PDF PaperPAPER

Within the framework of Essnet AIML4OS, Work Package 12, National Statistical Institutes (NSIs) are increasingly exploring the use of large language models (LLMs) to improve efficiency, accessibility, and quality across statistical production and dissemination processes.

Read more Read less In June 2025, a multi-country hackathon was organised, bringing together experts from six European countries: Portugal, the Netherlands, Sweden, Ireland, Norway, and France, to collaboratively explore practical, quality-focused applications of LLMs in official statistics. The initiative aimed to promote cross-country knowledge exchange, assess feasibility under real institutional and governance constraints, and develop testable prototypes that contribute to a shared methodological foundation for the responsible adoption of LLMs in official statistics.

The hackathon focused on key quality-related dimensions that are highly relevant for National Statistical Institutes (NSIs), including efficiency gains, reusability, data accessibility, on-premise compatibility, robustness of evaluation, feasibility, and long-term sustainability. The prototype presented in this paper, Dissemination Summary, was developed by the Portuguese team and addresses quality enhancement in statistical dissemination. It automatically generates concise, multilingual summaries of statistical reports published 20 to 30 years ago. Although these reports were carefully prepared, they are typically lengthy, with relevant information dispersed throughout the text, and the underlying data were not structured in database formats, as modern dissemination infrastructures did not yet exist at the time. As a result, retrieving and reusing this information today is time-consuming and costly.

The proposed approach leverages generates structured summaries enriched with keywords and thematic tags. Given the critical importance of validation, transparency, trustworthiness, and traceability in official statistics, every numerical value included in the generated summaries is accompanied by an explicit reference indicating its exact location in the original document, allowing users to directly verify the source of each figure.

Built on a retrieval-augmented generation (RAG) workflow, it ingests PDF reports authored by subject-matter experts, applies vector embeddings and LLM prompting, and outputs standardized summaries in machine-readable formats like JSON or Markdown. The design supports both local and external LLM deployments, aligning with institutional requirements for data confidentiality and on-premise compatibility. Evaluation across the predefined quality dimensions indicated high scores for reusability, feasibility, data accessibility, and lifespan, highlighting the need for further work on evaluation robustness.

This contribution demonstrates how collaborative, prototype-driven experimentation can translate emerging AI technologies into concrete quality improvements for official statistics. It offers practical insights into balancing innovation with institutional constraints and providing roadmap for integrating LLM-based tools into quality management frameworks within statistical organizations.

Main author / Presenter
Luis Ferreira
Statistics Portugal (INE)

Read more Read less Luis Ferreira holds a degree in Informatics. Has been working at Statistics Portugal (INE) since 2000, within the Methodological and Information Systems Department, where he focuses on the development and design of applications to support surveys and data collection.


CO-AUTHORS:

Paula Cruz, Statistics Portugal (INE)
Ivo Tavares, Statistics Portugal (INE)
Maria Ferreira, Statistics Portugal (INE)
Sónia Quaresma, Statistics Portugal (INE)

Presentation title
Leveraging AI and GIS for Official Statistics: Case Studies, Risks, and Considerations for Quality
PPT PaperPRESENTATION PDF PaperPAPER

Quality is the foundation of trust in official statistics, and the integration of AI and Geographic Information Systems (GIS) offers a practical path to increasing accuracy, insights, and the public's use of official statistics.

Read more Read less This presentation will demonstrate how deep learning tools can be applied to satellite imagery, flagging changes and updating building records- ultimately providing a new data source to augment field data collection and the use of administrative records. We will present use cases from NSOs around the world where AI and GIS have effectively improved coverage, timeliness, and consistency. At the same time, we will address the risks that accompany increased reliance on imagery, including scene misclassification, temporal drift, and model bias. We outline mitigation strategies such as rigorous ground-truthing, stratified sampling, uncertainty quantification, and continuous quality assurance to safeguard output integrity. Beyond imagery, geospatial machine learning methods help identify patterns and anomalies, verify the quality of collected data, and elevate the reliability of statistical outputs. Incorporating geospatial context in analysis reveals geographic dynamics that are otherwise hidden, sharpening decision-making and resource allocation. Finally, AI- and GIS-enabled dissemination tools —through interactive, map-based interfaces— make statistics more transparent and accessible, strengthening user understanding and confidence in official statistics. Taken together, these practices navigate opportunities and risks so that AI and GIS enhance quality and sustain trust in official statistics.

Main author / Presenter
Kate Hess
Environmental Systems Research Institute, Inc. (Esri), United States

Read more Read less Kate Hess is Esri's Business Development Manager for Official Statistics, based in New York City. She supports national statistics offices globally, promoting the use of GIS and remote sensing to modernize census operations, data analysis and dissemination. Previously, Kate spent 5 years as a Senior Solution Engineer at Esri working as a technical specialist with NSOs, UN agencies, and supporting efforts toward the UN Sustainable Development Goals. Before joining Esri, she worked with the NASA DEVELOP program, using satellite data to study climate change impacts, and supported the NASA Global Ecosystem Dynamics Investigation (GEDI) mission as a master's student at the University of Maryland. Her specialties include applications of GeoAI and extracting information from imagery, the integration of statistical and geospatial data, field data collection and operations, and GIS data dissemination/analysis applications.

Hall location

Cookies

This website uses cookies to ensure you get the best experience.

x