Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-32B-Annotated-32768-Tokens Overview This dataset is a variant of marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated with an extended maximum sequence length. The responses in the generated_text column were generated with max output tokens = 32768 (instead of 7500 in the original dataset), allowing for longer and more complete chain-of-thought reasoning. Generation Details Model: Qwen/Qwen3-32B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens.tabulartext-generation10K<n<100K0 likes64 downloads9mo agoHugging Face02onekq-ai /the-stack-v2-dedup-sql-annotateThis dataset focuses the SQL corpus of the Stack v2 dedup dataset and provides annotations to each SQL file therein. sqlparse is used to parse the SQL code, then count keywords and symbols. Below are the annotation columns. Column Name Column Description Keyword.DML Data Manipulation Language commands for modifying database records, e.g. SELECT, INSERT, UPDATE, DELETE, COMMIT, MERGE, ROLLBACK. Keyword.DDL Data Definition Language commands for defining or altering database… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/the-stack-v2-dedup-sql-annotate.tabulartext-generation1M<n<10M0 likes43 downloads2y agoHugging Face03xaviviro /FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated Federico García Lorca - Annotated Poetry Dataset A curated and annotated dataset of 283 poems by Federico García Lorca, spanning 9 of his major works (1921--1940). Each poem is enriched with publication metadata and GPT-4-generated thematic and contextual annotations. Use Case: LLM Generalization Evaluation This dataset was created to evaluate how well large language models can generalize literary style from a small, domain-specific corpus. It has been used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated.tabulartext-generationn<1K0 likes29 downloads7mo agoHugging Face04Hellisotherpeople /OpenDebateEvidence-Annotated-Anonymized OpenDebateEvidence-Annotated (Anonymized) An LLM-annotated subset of OpenDebateEvidence debate evidence, with all debater-identifying columns removed. This is an anonymized, Parquet-converted redistribution of Hellisotherpeople/OpenDebateEvidence-Annotated. 85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were dropped. See Anonymization. Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.tabulartext-classification10K<n<100K0 likes28 downloads2mo agoHugging Face05JingweiNi /opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513 Qwen3.5 FP8 Annotations for 10K K2-Think OCR2 Coding Steps This dataset contains Qwen3.5 FP8 step-level correctness annotations for K2-Think reasoning traces on a hard Codeforces subset of OpenCodeReasoning-2. Summary Source trace dataset: opencode_reasoning2_hard_codeforces2000_pr03_k2_thinking_extracted_pilot10 Source rows: 10 hard coding problem traces Candidate step rule: claim with non-empty aligned_token_ids Candidate steps: 15,267 Manifest-selected annotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513.tabulartext-generationn<1K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.