Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes81 downloads2mo agoHugging Face02Sheida1 /excavision-curation-assets Excavision target-conditioned curation assets Downloadable corpus artifacts for the Curate for my site workflow in Excavision Explorer. Canonical files excavision_full_vitb14.npy: the original 882,728 × 768 full-frame DINOv2 ViT-B/14 embeddings in float32, used directly without PCA or dimensionality reduction (SHA-256: 1b4b235fbffea44399e347de0a16cb4b60eba2ed5a30e8d344c908a9fa1d6bb2). excavision_curation_pool.parquet: aligned sanitized filenames and seven… See the full description on the dataset page: https://huggingface.co/datasets/Sheida1/excavision-curation-assets.tabularn<1K0 likes62 downloads1mo agoHugging Face03nguyenkhanh87 /ViLegalQA-Synthetic-Curation ViLegalQA Synthetic Curation Dataset summary This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel. Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.tabularquestion-answering10K<n<100K0 likes52 downloads1mo agoHugging Face04argilla /ultrafeedback-binarized-curation Ultrafeedback binarized dataset using the mean of preference ratings Introduction This dataset contains the result of curation work performed by Argilla (using Argilla 😃). After visually browsing around 200 examples using the sort and filter feature of Argilla, we noticed a strong mismatch between the overall_score in the original UF dataset (and the Zephyr train_prefs dataset) and the quality of the chosen response. By adding the critique rationale to our Argilla… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-curation.tabular10K<n<100K3 likes29 downloads3y agoHugging Face05dvyomkesh /nemo-grpo-from083-full-edge-curation Nemotron 0.83 Edge-Prompt Curation This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset. Seed edge prompts: 134 New rollout rows after seed exclusion: 7668 New edge prompts: 1648 Full edge prompts, seed plus rollout: 1782 Full dataset rows: 7830 Edge rate over full dataset: 0.2276 The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl. The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.tabular1K<n<10K0 likes24 downloads5mo agoHugging Face06Team-RAMEN /ramen-g1-ikea-assembly-curation RAMEN G1 IKEA Assembly — segment curation labels Per-segment quality labels for all 533 episodes of BitRobot/G1_WBT_Dex1_Building-Children-Table (revision 4b5961c6f8b97ececa75ee73041b30bcbf463ae2), made by Team RAMEN for the IROS 2026 Humanoid IKEA Assembly Challenge (Unitree G1 + Dex1-1 assembling a children's table). Each label marks a frame range of one episode, says which skill the robot performs there, and grades how well it is done. We used the labels to cut the episodes… See the full description on the dataset page: https://huggingface.co/datasets/Team-RAMEN/ramen-g1-ikea-assembly-curation.tabular10K<n<100K0 likes16 downloads1d agoHugging Face07Reza2kn /curation-backup-review_common_voice25_validatedtabular10K<n<100K0 likes15 downloads4mo agoHugging Face08plaguss /curation-ultrafeedback-scorestabular1K<n<10K0 likes13 downloads3y agoHugging Face09Reza2kn /curation-backup-review_common_voice25_dev Dataset Card for curation-backup-review_common_voice25_dev This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Using this dataset with Argilla To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code: import argilla as rg ds… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/curation-backup-review_common_voice25_dev.tabularn<1K1 likes13 downloads4mo agoHugging Face10plaguss /curation-ultrafeedback-bad-ratedtabular1K<n<10K0 likes11 downloads3y agoHugging Face11Northstar124234 /6000Q_A3_curation_data_P3tabularn<1K0 likes11 downloads2y agoHugging Face12Reza2kn /curation-backup-review_semiclean31_awq_wertabular10K<n<100K0 likes8 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.