COMET datasets available on ORION-DBs

Explore metadata enrichments of arXiv and DataCite

news
COMET
metadata enrichment
Author

Bianca Kramer

Published

September 1, 2026

Abstract

The COMET initiative is developing collaborative metadata enrichment practices, including ways for provenanced enrichments to flow into open scholarly metadata sources.

COMET provides open datasets with enrichments to arXiv and DataCite on Zenodo and Huggingface. DataCite enrichments can also be retrieved through the DataCite API, but are not (yet) included in DataCite metadata files.

To facilitate broader use of these datasets, including combining them with arXiv and DataCite metadata, various COMET datasets are now provided via ORION-DBs.

COMET projects

Together with community partners, COMET is carrying out various projects aimed at developing methods for enriching various elements in existing open scholarly metadata collections. Completed projects include:

  • Matching preprints to published articles

  • Improving affiliations parsing of preprints

  • Extracting funding metadata from acknowlegdments

All three projects were piloted using arXiv preprints.

DataCite enrichments

Apart from standalone datasets as project results, COMET also provides enrichments to DataCite DOI metadata, using a standardized format compatible with DataCite metadata schema. These enrichments extend the methodology developed in COMET projects to a wider corpus of metadata.

Currently, three types of enrichments of DataCite DOI metadata are made available:

  • matching affiliations

  • identifying and matching funders

  • improving resource type classification

Supporting usage and interoperability

COMET makes these datasets available as JSON, CSV and/or Parquet files via Zenodo and Huggingface. The enriched DOI metadata can also be retrieved via the DataCite API, but are not (yet) included in DataCite and arXiv metadata files.

To promote usage of these enriched metadata,various COMET metadata files are now made available through ORION-DBs (provided by Sesame Open Science), where they can be queried using Google Big Query.

These datasets can also be directly combined with DataCite and arXiv at scale, using the data snapshots similarly provided through ORION-DBs.

Available datasets

The following datasets are available (follow links for documentation)

Share your ideas!

Do you have ideas on how to use these datasets to explore community-enriched metadata? We would love to know your use cases - let us know at info@orion-dbs.community, or Dione Mentis and Adam Buttrick at COMET!