Global Biodiversity Data Harmonization Framework

Authors

  • Hugo Novak Assistant Professor, Department of Computer Science, Mediterranean Institute of Technology, Rome, Italy Author
  • Marta Novak Associate Professor, Department of Computer Science, Baltic AI Research University, Tallinn, Estonia Author
  • Sofia Nowak Associate Professor, Department of Computer Science, Mediterranean Institute of Technology, Rome, Italy Author

Keywords:

biodiversity informatics, data harmonisation, GBIF, taxonomic backbone, species occurrence records, biodiversity indicators, open science, GBF monitoring

Abstract

The global biodiversity informatics infrastructure -- comprising GBIF, iNaturalist, OBIS, TRY, PanTHERIA, IUCN Red List, and hundreds of national and regional databases -- collectively holds over 4 billion biodiversity records from millions of data contributors. Yet this extraordinary data wealth remains largely siloed, inconsistently structured, and difficult to integrate for cross-scale analyses due to heterogeneous data standards, inconsistent taxonomic backbone alignment, varying spatial and temporal resolution, and incompatible metadata schemas. The consequence is that policy-relevant biodiversity indicators -- species trends, functional diversity indices, ecosystem integrity metrics -- must be computed from data subsets that meet harmonisation requirements, discarding the majority of available records. This paper presents the BioDH (Biodiversity Data Harmonization) Framework -- a comprehensive, open-source computational pipeline for automated harmonisation of biodiversity occurrence, trait, threat, and monitoring data across the major global biodiversity databases. BioDH integrates seven processing modules: taxonomic backbone alignment (GBIF + Catalogue of Life), coordinate precision standardisation, temporal resolution harmonisation, trait data imputation (phylogenetic mean), occurrence record quality scoring, metadata completeness assessment, and cross-database deduplication. Applied to a benchmark dataset of 247 million records from 12 databases, BioDH increased the proportion of records usable for national biodiversity indicator computation from 28.4% to 74.8% (+163% improvement), reduced taxonomic naming conflicts from 18.4% to 2.4% of records (-87% reduction), and enabled computation of species trend indicators for 84.7% of IUCN-assessed species where baseline data exists -- compared with 47.4% achievable without harmonisation. The framework is implemented as an open-source Python package (biodh v1.0) with API connections to all major databases and full documentation, enabling any research group or national monitoring agency to implement standard harmonised biodiversity data workflows within existing computational resources.

Author Biographies

  • Hugo Novak, Assistant Professor, Department of Computer Science, Mediterranean Institute of Technology, Rome, Italy

    Assistant Professor, Department of Computer Science, Mediterranean Institute of Technology, Rome, Italy

  • Marta Novak, Associate Professor, Department of Computer Science, Baltic AI Research University, Tallinn, Estonia

    Associate Professor, Department of Computer Science, Baltic AI Research University, Tallinn, Estonia

  • Sofia Nowak, Associate Professor, Department of Computer Science, Mediterranean Institute of Technology, Rome, Italy

    Associate Professor, Department of Computer Science, Mediterranean Institute of Technology, Rome, Italy

Downloads

Published

2025-12-15

Versions

How to Cite

Global Biodiversity Data Harmonization Framework. (2025). International Journal of Animal Biodiversity, Conservation and Systematics ( IJABC), 5(4), 98-107. https://stanfordgroup.org/index.php/IJABC/article/view/296 (Original work published 2026)

Similar Articles

51-60 of 92

You may also start an advanced similarity search for this article.