COMET datasets available on ORION-DBs
Explore metadata enrichments of arXiv and DataCite
The COMET initiative is developing collaborative metadata enrichment practices, including ways for provenanced enrichments to flow into open scholarly metadata sources.
COMET provides open datasets with enrichments to arXiv and DataCite on Zenodo and Huggingface. DataCite enrichments can also be retrieved through the DataCite API, but are not (yet) included in DataCite metadata files.
To facilitate broader use of these datasets, including combining them with arXiv and DataCite metadata, various COMET datasets are now provided via ORION-DBs.

COMET projects
Together with community partners, COMET is carrying out various projects aimed at developing methods for enriching various elements in existing open scholarly metadata collections. Completed projects include:
Matching preprints to published articles
Improving affiliations parsing of preprints
Extracting funding metadata from acknowlegdments
All three projects were piloted using arXiv preprints.
DataCite enrichments
Apart from standalone datasets as project results, COMET also provides enrichments to DataCite DOI metadata, using a standardized format compatible with DataCite metadata schema. These enrichments extend the methodology developed in COMET projects to a wider corpus of metadata.
Currently, three types of enrichments of DataCite DOI metadata are made available:
matching affiliations
identifying and matching funders
improving resource type classification
Supporting usage and interoperability
COMET makes these datasets available as JSON, CSV and/or Parquet files via Zenodo and Huggingface. The enriched DOI metadata can also be retrieved via the DataCite API, but are not (yet) included in DataCite and arXiv metadata files.
To promote usage of these enriched metadata,various COMET metadata files are now made available through ORION-DBs (provided by Sesame Open Science), where they can be queried using Google Big Query.
These datasets can also be directly combined with DataCite and arXiv at scale, using the data snapshots similarly provided through ORION-DBs.
Available datasets
The following datasets are available (follow links for documentation)
- COMET project results
- arXiv preprint matching
- arXiv preprint author affiliation extraction
- arXiv funding entity extractions
- COMET DataCite enrichments
- affiliations
- funders
- resource types
- DataCite
- DataCite monthly data file
- arXiv
- arXiv metadata file (retrieved from Kaggle)