Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01evaluate /glue-ci Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained analysis of system… See the full description on the dataset page: https://huggingface.co/datasets/evaluate/glue-ci.tabulartext-classification1M<n<10M1 likes2.7k downloads1y agoHugging Face02open-source-metrics /evaluate-dependents evaluate metrics This dataset contains metrics about the huggingface/evaluate package. Number of repositories in the dataset: 106 Number of packages in the dataset: 3 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 1 packages that have more than 1000 stars. There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.tabular1K<n<10K0 likes895 downloads2y agoHugging Face03masato-ka /SO100_evaluate_generalize_pick_posThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 90, "total_frames": 33529, "total_tasks": 1, "total_videos": 180, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:90" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos.tabularrobotics10K<n<100K0 likes137 downloads1y agoHugging Face04masato-ka /SO100_evaluate_generalize_pick_pos_extendThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 30, "total_frames": 11184, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos_extend.tabularrobotics10K<n<100K0 likes105 downloads1y agoHugging Face05reciprocate /lichess-puzzles-evaluated-1Mtabular1M<n<10M0 likes91 downloads2mo agoHugging Face06MatanBT /gcg-evaluated-dataGCG suffixes crafted on Gemma-2, Qwen-2.5 and Llama-3.1, their generated response when appended to harmful instructions (from AdvBench, StrongReject's custom), their evaluation and charecterization. This dataset was created and utilized in the paper: Universal Jailbreak Suffixes Are Strong Attention Hijackers (paper, code). WARNING: this dataset contains harmful content, and is intended for research purposes only. Each row in the dataset describes: Harmful instruction info:… See the full description on the dataset page: https://huggingface.co/datasets/MatanBT/gcg-evaluated-data.tabular1M<n<10M0 likes77 downloads1y agoHugging Face07VoicAndrei /so100_cubes_evaluateThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 10, "total_frames": 9010, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/VoicAndrei/so100_cubes_evaluate.tabularrobotics1K<n<10K0 likes45 downloads1y agoHugging Face08NickyNicky /Nectar_evaluate_prompt_all_v1tabular100K<n<1M0 likes34 downloads2y agoHugging Face09Ahi-Yu /SO100_evaluate_generalize_pick_posThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so100_follower", "total_episodes": 90, "total_frames": 33529, "total_tasks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:90" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/Ahi-Yu/SO100_evaluate_generalize_pick_pos.tabularrobotics10K<n<100K0 likes34 downloads7mo agoHugging Face10BAAI /ROME-Evaluatedtabular1K<n<10K1 likes33 downloads1y agoHugging Face11hredkeith /so100_bi_test_evaluate2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "lekiwi_client", "total_episodes": 3, "total_frames": 2692, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:3" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hredkeith/so100_bi_test_evaluate2.tabularrobotics1K<n<10K0 likes33 downloads7mo agoHugging Face12lawful-good-project /sud_resh_evaluated_llms_answers 📊 Результаты Оценки Больших Языковых Моделей на Бенчмарке Судебных Решений В данном документе представлен анализ производительности 15 больших языковых моделей (LLM), протестированных на специализированном бенчмарке, который включает 105 000 записей из судебных решений России. Оценка проводилась по 10 различным категориям права (например, трудовое, уголовное, гражданское) и 7 типам инструкций (например, изложение исковых требований, анализ доказательств, итоговое решение). Ответы… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud_resh_evaluated_llms_answers.tabular100K<n<1M0 likes32 downloads1y agoHugging Face13hredkeith /so100_bi_test_evaluateThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "lekiwi_client", "total_episodes": 3, "total_frames": 1835, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:3" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hredkeith/so100_bi_test_evaluate.tabularrobotics1K<n<10K0 likes31 downloads7mo agoHugging Face14ChavyvAkvar /AlphaTrade-0.6B-SFT-v0.1-Evaluated-Dataset-1tabular100K<n<1M0 likes30 downloads1y agoHugging Face15llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp32 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32 Total Samples: 60 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 7 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32.tabulartext-generationn<1K0 likes30 downloads1y agoHugging Face16EunsuKim /benchhub_plus_results_evaluated BenchHub Plus Results (Evaluated) LLM inference results on the BenchHub Plus benchmark, with per-sample accuracy scores. Folder Structure ├── vllm_inference_results_en/ # English benchmark results (19 models) │ ├── {model_name}_{date}.jsonl │ └── ... └── vllm_inference_results_ko/ # Korean benchmark results (16 models) ├── {model_name}_{date}.jsonl └── ... Column Description Each .jsonl file contains one JSON object per line with the… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/benchhub_plus_results_evaluated.tabulartext-generation100K<n<1M0 likes29 downloads8mo agoHugging Face17xinshuo /ET_1k_evaluated_Deepseek31_20251210tabular1K<n<10K0 likes27 downloads9mo agoHugging Face18ChavyvAkvar /AlphaTrade-0.6B-SFT-v0.1-Evaluated-Dataset-2tabular100K<n<1M0 likes26 downloads1y agoHugging Face19eval-aware /Large-Language-Models-Often-Know-When-They-Are-Being-Evaluatedgated Dataset Card for Evaluation Awareness Benchmark Dataset Summary This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage. The dataset contains 976 conversational transcripts with rich metadata, including: True evaluation transcripts from prompt-injection tests, red-teaming tasks, and coding challenges Organic/real transcripts from actual user queries, scraped chats, and… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated.tabulartext-classificationn<1K0 likes25 downloads7mo agoHugging Face20xinshuo /ET_evaluated_deepseek31_passatntabular1K<n<10K0 likes24 downloads10mo agoHugging Face21tunis-ai /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K2 likes22 downloads1y agoHugging Face22xinshuo /ET_evaluated_deepseek31_passatn_detailedtabular1K<n<10K0 likes22 downloads10mo agoHugging Face23TAUR-dev /SIE_EVAL__Countdown3arg_6-24-25_FiRC__sft__samples__bf_evaluatedtabular1K<n<10K0 likes20 downloads1y agoHugging Face24archit11 /claude_code_traces_dirty_evaluated_v2tabularn<1K0 likes20 downloads9mo agoHugging Face25llm-aes /MT-Bench_Evaluatedtabular1K<n<10K0 likes19 downloads3y agoHugging Face26TAUR-dev /SIE_EVAL__Countdown3arg_FiRC_6-26-25-merged__rl__samples__bf_evaluatedtabular1K<n<10K0 likes19 downloads1y agoHugging Face27diwank /slimorca-autoj-evaluatedtabular100K<n<1M0 likes16 downloads3y agoHugging Face28LuckyLukke /NEGOTIO_evaluate_evaluatortabularn<1K0 likes16 downloads2y agoHugging Face29TAUR-dev /SIE_EVAL__countdown3arg_ssbon_think_p5chance__sft__samples__bf_evaluatedtabular1K<n<10K0 likes16 downloads1y agoHugging Face30TAUR-dev /SIE_EVAL__Countdown3arg_6-24-25_Distilled_QWQ__sft__samples__bf_evaluatedtabular1K<n<10K0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.