datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.NepTam-A-Nepali-Tamang-Parallel-Corpus
🧾 NepTam — A Nepali–Tamang Parallel Corpus
Dataset Summary
NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains:
20K gold-standard human-translated sentence pairs, and
80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus.
Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.streamlit_docsdigits
Dataset Card for digits dataset
Optical recognition of handwritten digits dataset
Note - How to load this dataset directly with the datasets library
from datasets import load_dataset
dataset = load_dataset("sklearn-docs/digits",header=None)
Dataset Summary
This is a copy of the test set of the UCI ML hand-written digits datasets https://archive.ics.uci.edu/ml/datasets/Optical+Recognition+of+Handwritten+Digits
The data set contains images of hand-written… See the full description on the dataset page: https://huggingface.co/datasets/sklearn-docs/digits.compile-benchmarksdoc-splits-1
[doc] file names and splits 1
This dataset contains a data.csv file at the root.
docs-python-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for creating python docs from methods. This is formatted from semeru/code-code-galeras-code-completion-from-docstring-3k-deduped
Dataset Description
Curated by: semeru/code-code-galeras-code-completion-from-docstring-3k-deduped
Language(s) (NLP): Python
License: [More Information Needed]
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ASHu2/docs-python-v1.doc-splits-3
[doc] file names and splits 3
This dataset contains three csv files at the root: my_train_file.csv, test-file.csv, validation1.csv.
doc-splits-6
[doc] file names and splits 6
This dataset contains six files at the root, four for the training split, and two for the test split.
doc-splits-2
[doc] file names and splits 2
This dataset contains three csv files at the root: train.csv, test.csv, validation.csv.
experts-backendsspring-docs
License
Spring Projects: Apache License 2.0. Copyright © 2024 Broadcom. All Rights Reserved.
doc-splits-4
[doc] file names and splits 4
This dataset contains three subdirectories, inside data/, called train, test and validation, with csv files in them.
qdrant_docs_qna_ragaswatsonx-docs-document-type
Watsonx Docs Document Type Classification
This dataset is a balanced binary document-level classification subset derived
from ibm-research/watsonxDocsQA.
Task
Classify IBM Watsonx documentation pages by their dominant user-facing purpose:
conceptual: documents primarily used to understand or look up information.
how-to: documents primarily used to complete a procedure or fix a problem.
Splits
Split
conceptual
how-to
Total
train
140
140
280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.laravel-docsdoc-splits-8
[doc] file names and splits 8
This dataset contains seven files under the data/ directory, three for the train split, one for the test split and three for the random split.
wasp-docsz-imagedoc-splits-5
[doc] file names and splits 5
This dataset contains three files inside data/, called training.csv, eval.csv and valid.csv.
code_docstringscode-docstring-datasetFDA_DocsLangChain_docs_usecases_integrationscodehawks-docs-qapybamm-docskernel-ltx-videolangchain-docs-csvdoc-splits-7
[doc] file names and splits 7
This dataset contains six files under the data/ directory, four in the train/ subdirectory, and two in the test/ subdirectory.
MNLP_rag_docs
