Publications of the Council

Large Language Models in the Oversight of Public Procurement: What They Detect and What Is Known about the Results. The Experience of Brazil, Korea, Chile, Colombia and the European Union

Download the report

Abstract

The public authorities of Brazil, the Republic of Korea, Chile, Colombia and the European Union have begun to apply large language models to the review of public procurement. The study asks which tools built on language models these authorities apply, what such tools check, which mechanisms of distortion of a competitive procedure they are able to detect, and what has been published about the results of their work.

Data and method. The study rests on the documents of the authorities that apply the tools: decisions and technical articles of the Federal Court of Accounts of Brazil and of the state courts of accounts, press releases and tender documentation of the Public Procurement Service of Korea, publications of ChileCompra and the award record in Mercado Publico, the directive of the Italian National Anti-Corruption Authority, and reports of the European Commission and the European Anti-Fraud Office. For each tool a card of six fields was completed; the functions of the tools were mapped to the 30 mechanisms of the CILC typology on three values (detects directly, detects indirectly, does not detect); performance was calculated from published figures with the denominator stated.

Principal findings. By October 2026 language models had not become the principal tool of review in any of the five jurisdictions. The inventory comprises six systems in which the use of a language model is confirmed by the authority's own documents and two announced projects whose technology or procurement function is unconfirmed; of the six confirmed systems two are in operation, and for only one has the authority published a decision with results. That one is Alice 360 of the Federal Court of Accounts of Brazil, where the model performs a preparatory task, recognising the subject matter of a procurement from imprecise descriptions, after which prices are compared statistically. The two operating tools cover five of the 30 mechanisms of the typology; counting the systems under construction in Korea, Chile and Seocho-gu, potential coverage reaches 10 mechanisms at five of the ten stages of the procurement cycle. No operating tool checks bid submission, bid evaluation, review, performance or payment, and the cluster of mechanisms “neutralisation of control” remains outside their field of view. The boundary of the language model's visibility coincides with the boundary of the external participant's visibility in the typology.

Conclusions. No authority publishes the share of well-founded alerts by individual function of the language model; the only published shares relate to tools without a language model and range from 3.4 to 23.5 per cent. The signs that the tools look for yield four candidates for new mechanisms of the typology. To monitor the quality of a language model and of its application, an authority must collect at least four quantities for each function: alerts issued, alerts reviewed by a person, alerts confirmed and alerts that led to measures. Without these data the effectiveness of the language model is indistinguishable from the effectiveness of the rules to which it has been added.

The record contains the report in English, Spanish and Russian.

Recommended citation

Limarau, D., Krukovskiy, I. (2026). Large language models in the oversight of public procurement: what they detect and what is known about the results. Report. Consejo Internacional para la Lucha contra la Corrupción. DOI 10.5281/zenodo.23241236. https://cilclegal.org/

All publications