Research
The technical work behind connecting and cleaning public spending data across jurisdictions, built on the Public Spending (PSNET) ontology.
Entity Resolution
Government payment records often list the same organization under different name variants. The team develops algorithms to unify these variants so that payments to a single company can be tracked consistently across datasets and jurisdictions, including links to Forbes Global 2000 company entries.
Classification Harmonization
Public spending categories are reported using different national schemes. The research addresses this by harmonizing variform payment category classifications, including CPV (Common Procurement Vocabulary) and NAICS (North American Industry Classification System), so spending can be compared across countries and administrative levels.
The research arm of Public Spending focuses on the technical engineering required to make sense of disparate government payment records from jurisdictions as varied as Greece, Australia, the United Kingdom, and cities across the United States. By developing algorithms that unify company name variants, the team ensures that payments to a single entity can be tracked consistently across datasets. This entity resolution work is foundational to the broader mission of following trillions in public money, allowing analysts, journalists, and citizens to see where funds actually flow rather than getting lost in inconsistent naming conventions.
Classification harmonization is another core research priority. Different national schemes report public spending categories in incompatible ways, from CPV codes in European contexts to NAICS classifications elsewhere. The research team works to map these parallel systems onto a unified framework built on the Public Spending ontology, so that comparable categories across countries can be examined side by side. This effort enriches expenditure records with meaningful interconnections, enabling economic and international comparisons that would otherwise be impossible given the variform ways governments label their spending.
The ontology itself is a living structure, continually refined as new datasets are ingested and new challenges emerge. The team focuses on developing and extending the model to accommodate the full range of payers and payees represented in public financial records, from small municipal vendors to global corporations listed among the world's largest companies. By linking payment records to entity classifications and standardized categories, the research creates a reusable foundation that supports everything from bulk downloads to interactive queries, making the underlying data accessible to anyone who wants to follow the money.
Ultimately, this technical work serves a public purpose: making government spending transparent, comparable, and analyzable on a global scale. The researchers prioritize practical outcomes, ensuring that the algorithms and classifications they build can be reused by other teams and organizations working with public financial data. Every refinement to entity resolution, every harmonized category, and every enriched record contributes to a growing datahub of interconnected spending information, giving citizens and oversight bodies alike the tools they need to watch how public money moves across borders, sectors, and time.
Ongoing Enrichment
The underlying dataset — built on the PSNET ontology — is continually enriched with public expenditure records gathered from government portals and related open data sources worldwide, supporting further interlinking and analysis.
PSNET Ontology
A data model for representing payers, payees, payments, and categories in a consistent, linked structure.
Name Unification
Algorithms that match differently written company names to a single underlying entity.
Category Mapping
Cross-referencing CPV and NAICS codes to align spending categories between countries.