datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases!
Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills!
This dataset contains:
63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1.
Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.Origin-Sequence-Data
AL-GR/Origin-Sequence-Data: Raw User Behavior Sequences 📜
About the Dataset
Each row in this dataset (Origin-Sequence-Data) represents a step in a user's journey, consisting of a sequence of previously interacted items (user_history) and the next item they interacted with (target_item). All item IDs have been anonymized into short, unique strings.
This dataset is ideal for:
🧑🔬 Researchers who want to design their own data processing or prompting strategies for… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Origin-Sequence-Data.pdb_sequences
PDB Sequences
This dataset contains 780,163 protein sequences from the RCCB Protein Data Bank
Celestia3-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1 0528's science-reasoning skills!
This dataset contains:
90.9k synthetically generated science prompts, with all responses generated using DeepSeek R1 0528.
Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory.
All prompts are synthetic, taken… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528.Titanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.thermo-seq
Nanobody Thermal Stability Dataset
Dataset Overview
This dataset helps predict how stable nanobody sequences are at different temperatures. Thermal stability is important for nanobody engineering and applications, affecting how well they work in different environments.
The dataset includes two types of stability measurements:
Melting temperature (Tm): The temperature at which nanobodies start to unfold
Sequence stability: Stability scores based on sequence properties… See the full description on the dataset page: https://huggingface.co/datasets/ZYMScott/thermo-seq.Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.Titanium2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Titanium2-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills!
This dataset contains:
32.4k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2-DeepSeek-R1.Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks!
This dataset contains:
28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale.
Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Superpotion-DeepSeek-V3.2-Speciale.ro-offense-sequences
Dataset Card for "RO-Offense-Sequences"
Dataset Description
Homepage: https://github.com/readerbench/ro-offense-sequences
Repository: https://github.com/readerbench/ro-offense-sequences
Point of Contact: Teodora-Andreea Ion
Dataset Summary
a novel Romanian language dataset for offensive sequence detection with manually
annotated offensive sequences from a local Romanian sports news website (gsp.ro):
Resulting in 4800 annotated messages… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-offense-sequences.Titanium2.1-DeepSeek-R1Click here to support our open-source dataset and model releases!
Titanium2.1-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills!
This dataset contains:
31.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2.1-DeepSeek-R1.CelestiaCelestia is a dataset containing science-instruct data.
The 2024-10-30 version contains:
126k rows of synthetic science-instruct data, using synthetically generated prompts and responses generated using Llama 3.1 405b Instruct. Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory.
This dataset contains synthetically generated data and has not been subject to manual review.
vhh_affinity-seq
Nanobody (VHH) Affinity Prediction Dataset
Dataset Overview
This dataset helps predict the binding affinity between nanobodies (VHH, single-domain antibodies from camelids) and their target antigens. Affinity is a key parameter that measures how strongly an antibody binds to its antigen, usually expressed as dissociation constant (KD) or binding free energy.
High affinity is a critical property for therapeutic antibodies, so accurately predicting nanobody affinity is… See the full description on the dataset page: https://huggingface.co/datasets/ZYMScott/vhh_affinity-seq.protein_binding_sequences
Sequence Based Protein - Peptide Binding Dataset
Data sources:
Huang Laboratory
Propedia
YAPP-Cd
Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence
contains only the relevant chain.
Train / Val split: the dataset is split to 80% train 10% val and 10% test.
TitaniumTitanium is a dataset containing DevOps-instruct data.
The 2024-10-02 version contains:
26.6k rows of synthetic DevOps-instruct data, using synthetically generated prompts and responses generated using Llama 3.1 405b Instruct. Primary areas of expertise are AWS, Azure, GCP, Terraform, Dockerfiles, pipelines, and shell scripts.
This dataset contains synthetically generated data and has not been subject to manual review.
ro-offense-sequences
Dataset Card for "RO-Offense-Sequences"
Dataset Description
Homepage: https://github.com/readerbench/ro-offense-sequences
Repository: https://github.com/readerbench/ro-offense-sequences
Point of Contact: Teodora-Andreea Ion
Dataset Summary
a novel Romanian language dataset for offensive sequence detection with manually
annotated offensive sequences from a local Romanian sports news website (gsp.ro):
Resulting in 4800 annotated messages… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/ro-offense-sequences.TachibanaTachibana is a dataset containing code-instruct data.
The 2024-09-27 version contains:
104k rows of synthetic chat responses generated using Llama 3.1 405b Instruct.
60.6k Magicoder prompts from ise-uiuc/Magicoder-Evol-Instruct-110K
43.4k Glaive-code-assistant prompts from glaiveai/glaive-code-assistant
This dataset contains synthetically generated data and has not been subject to manual review.
Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.Tachibana4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Tachibana 4, containing the first 1.2k rows!
Tachibana 4 is an upcoming agentic coding dataset, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, systems programming, distributed systems… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro-PREVIEW.DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills!
This dataset contains:
4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528.
All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.Mitakihara-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
Mitakihara-DeepSeek-R1-0528 is a dataset focused on artificial intelligence, testing the limits of DeepSeek R1 0528's AI-reasoning skills!
This dataset contains:
16.9k synthetically generated prompts about AI, with all responses generated using DeepSeek R1 0528.
Subjects include computer science, artificial intelligence, MLOps, LLMs and diffusion models, math and CUDA, cutting-edge and future technologies, complex adaptive and… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara-DeepSeek-R1-0528.Tachibana2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Tachibana2-DeepSeek-R1 is a code-reasoning dataset, testing the limits of DeepSeek R1's coding skills!
This dataset contains:
27.2k synthetically generated code-reasoning prompts. All responses are generated using DeepSeek R1.
Synthetic prompts are generated using Llama 3.1 405b Instruct, based on the original sequelbox/Tachibana dataset with increased task complexity.
Responses demonstrate the code-reasoning capabilities of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana2-DeepSeek-R1.Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills!
This dataset contains:
a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale.
provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.Celestia3-DeepSeek-R1-0528-PREVIEWClick here to support our open-source dataset and model releases!
This is an early sneak preview of Celestia3-DeepSeek-R1-0528, containing the first 13.4k rows!
Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1's science-reasoning skills!
This early preview release contains:
13.4k synthetically generated science prompts. All responses are generated using DeepSeek R1 0528.
Primary subjects are physics, chemistry, biology, and computer science;… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528-PREVIEW.SupernovaSupernova is a dataset containing general synthetic chat data from the best available open-source models.
The 2024-09-27 version contains:
178.2k rows of synthetic chat responses generated using Llama 3.1 405b Instruct.
47k UltraChat prompts from HuggingFaceH4/ultrafeedback_binarized
131k SlimOrca prompts from Open-Orca/slimorca-deduped-cleaned-corrected
This dataset contains synthetically generated data and has not been subject to manual review.
Nanobody_Sequence_DatasetRepresentative sequence dataset extracted after clustering of Integrated NANOBODY® Database for Immunoinformatics (INDI) by MMseqs2 program.
For more information, please visit https://github.com/DynaX-C/EvoNB.
DES-Reasoning-DeepSeek-V3.1Click here to support our open-source dataset and model releases!
DES-Reasoning-DeepSeek-V3.1 is a dataset focused on analysis and reasoning, creating discrete event simulations testing the limits of DeepSeek V3.1's simulation, Python scripting, and analysis skills!
This dataset contains:
4.03k synthetically generated prompts to create discrete event simulations and analysis chat in response to user input, with all responses generated using DeepSeek V3.1.
All responses contain a multi-step… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DES-Reasoning-DeepSeek-V3.1.TCGA-Cancer-Variant-and-Clinical-Data
TCGA Cancer Variant and Clinical Data
Dataset Description
This dataset combines genetic variant information at the protein level with clinical data from The Cancer Genome Atlas (TCGA) project, curated by the International Cancer Genome Consortium (ICGC). It provides a comprehensive view of protein-altering mutations and clinical characteristics across various cancer types.
Dataset Summary
The dataset includes:
Protein sequence data for both mutated and… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/TCGA-Cancer-Variant-and-Clinical-Data.Tachibana3-Part2-DeepSeek-V3.2Click here to support our open-source dataset and model releases!
Tachibana3-Part2-DeepSeek-V3.2 is a dataset focused on high-difficulty code production tasks, testing the limits of DeepSeek V3.2's code-reasoning skills!
This dataset contains 9.3k high-difficulty code-production prompts:
Questions prioritize real-world, challenging coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, mobile, gamedev, cloud, QA, custom… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana3-Part2-DeepSeek-V3.2.
