datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-dwlongharness
LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
LongHarness evaluates how language-model harnesses access and reason over long
contexts. It is designed to distinguish context-access strategies, including
direct reading, lexical and semantic retrieval, iterative agents, and recursive
language-model harnesses. The benchmark contains 200 evaluation instances
across four task suites.
Project website
Paper
GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/StringNLP/longharness.STRING
STRING v12.0
STRING is a protein association network database that integrates experimental, computational, text-mined, and curated evidence for functional and physical protein interactions.
Configs
Config
Raw source
Description
protein_links
protein.links.full.v12.0.txt.gz
Protein-protein association edges with all STRING evidence channels and combined_score.
protein_info
protein.info.v12.0.txt.gz
Protein identifiers, preferred names, sizes, and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/STRING.trossen_place_bead_on_string_10_gr00t_01bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.trossen_place_bead_on_string_10_gr00t_02_2trossen_place_bead_on_string_10_gr00t_02coco-detection-stringsProcessed the bounding boxes from coco to paligemma like.
Reference dataset -> detection-datasets/coco
DFFT-Video-Test-Stringcode-meta-reasoning-cleaned-final-string-idCVHQ-Video-StringQX-NEPHRIEL-Seventh-String
QX-NEPHRIEL / Seventh-String Closure
A quantitative rank-three Chollet inequality, with a full analytic argument and reproducible exact-arithmetic certificates.
Release v1.0.0, 3 October 2026. AI-generated derivation by Eve, prepared for Maciej Nowicki.
For complex Hermitian positive semidefinite (7\times7) matrices (A,B), each of rank at most three, the manuscript establishes
per(A∘B)≤2425per(A)per(B).
\operatorname{per}(A\circ B)
\le… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/QX-NEPHRIEL-Seventh-String.StringDBSeqsv12All the IDs and sequences in StringDB version 12
https://string-db.org/cgi/download
string-opsStringWars
StringKilla - Small Datasets for String Algorithms Benchmarking
The goal of this dataset is to provide a fairly diverse set of strings to evalute the performance of various string-processing algorithms in StringZilla and beyond.
English Texts
English Leipzig Corpora Collection
124 MB uncompressed
1'000'000 lines of ASCII
8'388'608 tokens of mean length 5
The dataset was originally pulled from Princeton's website:
wget --no-clobber -O leipzig1M.txt… See the full description on the dataset page: https://huggingface.co/datasets/ashvardanian/StringWars.genminiall_no_na_no_weird_stringtask079_conala_concat_strings
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task079_conala_concat_strings
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task079_conala_concat_strings.trossen_place_bead_on_string_10_gr00t_crop_02bacbench-ppi-stringdb-dna-small
Dataset for protein-protein interaction prediction across bacteria (DNA)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/).
Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores.
The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.bacbench-ppi-stringdb-protein-sequences-small
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.stringdbstrings
Music Demixing Benchmarks Strings
Unofficial, unpacked mirror of MVSep Quality Checker Strings dataset.
Strings Leaderboard (Strings)
Official zip file: 0.6 GB
Download
hf CLI and the huggingface_hub Python package are orders of magnitude faster than git.
Use
Hugging Face Command Line Interface (CLI)
git
huggingface_hub Python package
Choose
individual files
all files in the repository
subsets of files with filename pattern matching (fnmatch)
Read more… See the full description on the dataset page: https://huggingface.co/datasets/MusicDemixingBenchmarks/strings.trossen_place_bead_on_string_10_gr00t_01_2leandojo-lean4-formal-informal-stringsOpenScore-StringQuartets
OpenScore String Quartets (OMR Evaluation)
This dataset is derived from the OpenScore String Quartets corpus (Gotham et al., 2023), a collection of string quartets by "long 19th century" composers. It is designed for evaluating Optical Music Recognition (OMR) systems.
We extract a subset of the OpenScore String Quartets that contains both scanned images of real scores and the corresponding MusicXML ground truth. We also render clean images from the MusicXML files using… See the full description on the dataset page: https://huggingface.co/datasets/guangyangmusic/OpenScore-StringQuartets.task600_find_the_longest_common_substring_in_two_strings
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task600_find_the_longest_common_substring_in_two_strings
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task600_find_the_longest_common_substring_in_two_strings.trossen_place_bead_on_string_10_gr00t_clip_01task1189_check_char_in_string
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1189_check_char_in_string
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1189_check_char_in_string.trossen_ai_stationary_place_bead_on_string_10task1316_remove_duplicates_string
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1316_remove_duplicates_string
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1316_remove_duplicates_string.
