Team Ai
7 results

crossref

bluuebunny /crossref_metadata_2025 Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025.textsentence-similarity10M<n<100M0 likes864 downloads1y agoHugging Facecometadata /crossref-unique-affiliations Crossref unique affiliations Exact organization strings extracted from Crossref snapshot 2026-07. For the nine exact-string splits, no trimming, case folding, or Unicode normalization is performed. Repeated leaf occurrences are counted, and empty decoded strings alone are excluded. Normalized split The normalized split groups the exact strings in all after applying this normalization contract: Transliterate Unicode text to ASCII with Unidecode. Convert letters… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-unique-affiliations.text100M<n<1B0 likes863 downloads1mo agoHugging Facebluuebunny /crossref_metadata_2025_split Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.textsentence-similarity10M<n<100M0 likes763 downloads1y agoHugging Facecometadata /crossref-datacite-citations Crossref DataCite Citations A dataset of DataCite-registered works and the Crossref-registered works that cite them, extracted from Crossref reference metadata and confirmed against the DataCite monthly data file. Dataset Description Each record in the citation configurations is one DataCite DOI together with every confirmed citing work found in Crossref reference metadata. A reference is confirmed when it carries a DOI registered in DataCite or an arXiv… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-datacite-citations.feature-extraction1M<n<10M0 likes158 downloads23d agoHugging FaceRichardErkhov /April_2023_Public_Data_File_from_Crossref0 likes138 downloads2y agoHugging Facecometadata /crossref-arxiv-citations Crossref arXiv Citations A dataset of arXiv preprints and their citations extracted from Crossref metadata, validated against DataCite records. Dataset Description This dataset maps arXiv works to the works in Crossref that cite them. Each record represents an arXiv preprint with all known citations from Crossref-registered works. Built from the Crossref Metadata Plus monthly snapshot 2026-07 and the DataCite monthly data file 2026-07 with… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-arxiv-citations.tabulartext-classification1M<n<10M0 likes130 downloads23d agoHugging Face