datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens
Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-32B-Annotated-32768-Tokens
Overview
This dataset is a variant of marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated with an extended maximum sequence length. The responses in the generated_text column were generated with max output tokens = 32768 (instead of 7500 in the original dataset), allowing for longer and more complete chain-of-thought reasoning.
Generation Details
Model: Qwen/Qwen3-32B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens.the-stack-v2-dedup-sql-annotateThis dataset focuses the SQL corpus of the Stack v2 dedup dataset and provides annotations to each SQL file therein.
sqlparse is used to parse the SQL code, then count keywords and symbols.
Below are the annotation columns.
Column Name
Column Description
Keyword.DML
Data Manipulation Language commands for modifying database records, e.g. SELECT, INSERT, UPDATE, DELETE, COMMIT, MERGE, ROLLBACK.
Keyword.DDL
Data Definition Language commands for defining or altering database… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/the-stack-v2-dedup-sql-annotate.FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated
Federico García Lorca - Annotated Poetry Dataset
A curated and annotated dataset of 283 poems by Federico García Lorca, spanning 9 of his major works (1921--1940). Each poem is enriched with publication metadata and GPT-4-generated thematic and contextual annotations.
Use Case: LLM Generalization Evaluation
This dataset was created to evaluate how well large language models can generalize literary style from a small, domain-specific corpus. It has been used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated.OpenDebateEvidence-Annotated-Anonymized
OpenDebateEvidence-Annotated (Anonymized)
An LLM-annotated subset of OpenDebateEvidence debate evidence, with all
debater-identifying columns removed.
This is an anonymized, Parquet-converted redistribution of
Hellisotherpeople/OpenDebateEvidence-Annotated.
85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence
evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were
dropped. See Anonymization.
Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513
Qwen3.5 FP8 Annotations for 10K K2-Think OCR2 Coding Steps
This dataset contains Qwen3.5 FP8 step-level correctness annotations for K2-Think reasoning traces on a hard Codeforces subset of OpenCodeReasoning-2.
Summary
Source trace dataset: opencode_reasoning2_hard_codeforces2000_pr03_k2_thinking_extracted_pilot10
Source rows: 10 hard coding problem traces
Candidate step rule: claim with non-empty aligned_token_ids
Candidate steps: 15,267
Manifest-selected annotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513.
