Team Ai
Datasetpublic

LLMDH/OpenScience

Open Science Dataset Overview Open Science is a large-scale, permissively licensed text dataset derived from OpenAlex, containing over 100B (105,390,332,599) words. OpenAlex is an open database of scholarly publications, authors, institutions, and research outputs that serves as a comprehensive source for academic literature. Key Features Truly Open: Contains only permissively licensed data suitable for both commercial and non-commercial use… See the full description on the dataset page: https://huggingface.co/datasets/LLMDH/OpenScience.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes3.7kdownloads
Dataset Card

Open Science Dataset

Overview

Open Science is a large-scale, permissively licensed text dataset derived from OpenAlex, containing over 100B (105,390,332,599) words. OpenAlex is an open database of scholarly publications, authors, institutions, and research outputs that serves as a comprehensive source for academic literature.

Key Features

  • —Truly Open: Contains only permissively licensed data suitable for both commercial and non-commercial use
  • —Multilingual Coverage: Predominantly English with significant representation of Portuguese, German, Indonesian, Russian, and Spanish
  • —Academic Diversity: Encompasses journal articles, books, datasets, theses, and other document types

Dataset Statistics

License distribution:

cc-by      ███████████████████████████████ 93.09%
cc-by-sa   ██ 3.68%
public     ██ 3.23%

[image]

  • —Total number of documents: 11,534,164 ## Language Distribution
LanguageDocument CountWord Count% of Total Words
English9,692,48091,961,796,09887.28%
Portuguese381,7192,911,255,3702.76%
German167,1301,606,879,1281.53%
Indonesian297,6451,370,676,3661.30%
Russian199,3171,235,729,3511.17%
Spanish189,3051,548,199,2261.47%

Dataset Structure

Data Fields

Field
identifier
pdf_url
lang
error
title
source_name
publication_year
license
word_count
text

How to Use

Installation

python
from datasets import load_dataset
dataset = load_dataset('PleIAs/openscource')

Considerations and Limitations

  • —While permissively licensed, users should verify compatibility with their specific use case
  • —Language distribution may not be uniform
  • —Historical texts may contain dated terminology or perspectives

Acknowledgements

This dataset was developed with support from:

  • —Jean Zay (Eviden, Idris)
  • —Nebius AI
  • —Tracto AI