datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
rebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.CodeReality
CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset
⚠️ Important Limitations
⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use.
Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.enpisi-coder-data
enpisi-coder RL dataset
Judge-rated npcsh agent traces and derived RL training data for the
enpisi-coder model family.
Produced by scripts/rate_traces.py (LLM-as-judge) and
scripts/analyze_ratings.py; built into SFT/DPO/GRPO/PPO splits by
scripts/train_from_csv.py.
Splits
Split
Rows
Description
rated_traces
900
Per-trace judge scores (correctness, tool_selection, efficiency, clarity, partial_credit, composite)
tasks
100
Benchmark task definitions… See the full description on the dataset page: https://huggingface.co/datasets/npc-worldwide/enpisi-coder-data.code-route-maroc-dataset
🚗 Code de la Route Marocain Dataset (Loi 52-05)
Ce jeu de données regroupe des questions, réponses et textes juridiques formalisés pour l'entraînement de modèles de langage (LLM) sur la réglementation routière au Maroc.
📊 Origine des données
Pipeline NLP Source : Récupéré depuis le projet Kaggle medaymanelkajdouhi/code-route-maroc-nlp.
Format d'export : Fichier export_final.csv converti en data.csv.
🎯 Utilisation
Ce dataset sert de support direct pour le… See the full description on the dataset page: https://huggingface.co/datasets/Zakariae-drabech/code-route-maroc-dataset.code_riskcoderag-eval
