Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01patched-codes /static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub), where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks. OpenAI used the synth-vuln-fixes and fine-tuned a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo. More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.textn<1K20 likes879 downloads1y agoHugging Face02Dahoas /rm-static Dataset Card for "rm-static" Split of hh-static used for training reward models after supervised fine-tuning. text10K<n<100K124 likes861 downloads4y agoHugging Face03PolinAvA /openarm_statictabularn<1K0 likes718 downloads6mo agoHugging Face04Heiheihaha17 /DiSH-Bench-static-v2 DiSH-Bench static v2 — test release A simulated multi-receiver spatial speech benchmark, with aligned single-receiver and multi-receiver questions over the same physical scenes. This release contains 122,256 test questions, 2,000 physical scenes, 4,400 original four-channel FOA recordings, and 200 rooms. There are 58,829 single-input questions, 53,518 two-input questions and 9,909 three-input questions. Questions share recordings; they are not independent acoustic scenes. Source… See the full description on the dataset page: https://huggingface.co/datasets/Heiheihaha17/DiSH-Bench-static-v2.audio100K<n<1M0 likes428 downloads10d agoHugging Face05Purple69 /aera_semi_pnp_dr_16_06_2026_skip3_delta_no_go_home_no_static_smoothedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AR4_MK3", "total_episodes": 2855, "total_frames": 399698, "total_tasks": 156, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 23, "splits": { "train": "0:2855" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Purple69/aera_semi_pnp_dr_16_06_2026_skip3_delta_no_go_home_no_static_smoothed.imagerobotics100K<n<1M0 likes325 downloads4mo agoHugging Face06Purple69 /aera_semi_pnp_dr_02_05_2026_skip3_delta_no_go_home_no_static_smoothed__0005This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AR4_MK3", "total_episodes": 5349, "total_frames": 362999, "total_tasks": 72, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 25, "splits": { "train": "0:5349" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Purple69/aera_semi_pnp_dr_02_05_2026_skip3_delta_no_go_home_no_static_smoothed__0005.imagerobotics100K<n<1M0 likes318 downloads5mo agoHugging Face07yahyaabd /bps-statictable-qrels-fulltabular1M<n<10M0 likes316 downloads2y agoHugging Face08Dahoas /static-hhStatic split of Anthropic's Helpful Harmless dataset. Contains base-online and rejection sampled outputs. text100K<n<1M21 likes294 downloads4y agoHugging Face09wumiaoshou /Static-FineBenchimage1K<n<10K2 likes266 downloads1y agoHugging Face10msr-spare-1 /qwen3-30b-0718-gpt55-static-balanced-7872-spare-games-envs qwen3-30B-A3B-0718-gpt55-static-balanced-7872-400 — generated environments Environments generated by the SPARE proposer during training run h5wf1kei (qwen3-30B-A3B-0718-gpt55-static-balanced-7872-400), recovered from the spare-viz durable cache. The run's scratch directory no longer exists; this dataset is the surviving copy. Games 194 Steps covered 9 (step 0–115) With recovered skill 0 With hint 0 Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0718-gpt55-static-balanced-7872-spare-games-envs.textn<1K0 likes260 downloads2mo agoHugging Face11links-ads /wildfire-risk-static-datageospatial1M<n<10M0 likes247 downloads12h agoHugging Face12Tristan /static-rlhf-interface-datatabularn<1K0 likes229 downloads3y agoHugging Face13henry1477 /pcbslm-static-v2-unsloth-vlm PCBSLM static-v2 Unsloth VLM Portable multimodal Unsloth dataset for PCB layout/document-grounded training. The JSONL splits use Unsloth/Gemma-style chat messages: { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "assets/raw_docs/.../images/page.png"}, {"type": "text", "text": "instruction..."} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "{...json answer...}"} ]} ] } Files… See the full description on the dataset page: https://huggingface.co/datasets/henry1477/pcbslm-static-v2-unsloth-vlm.documentimage-text-to-text1K<n<10K0 likes227 downloads6mo agoHugging Face14EverMind-AI /EverMemBench-Static EverMemBench-S: Evaluating Evidence Access under Dense Semantic Interference 💻 Code: EverMind-AI/EverMemBench-Static Overview EverMemBench-S (EMB-S) is an adversarial Needle-in-a-Haystack benchmark built on a 326M-token MemoryBank with 160,280 documents across 8 domains. It evaluates long-context models and retrieval systems under dense semantic interference — where near-miss documents create realistic confusion that standard NIAH benchmarks cannot capture. 1,225… See the full description on the dataset page: https://huggingface.co/datasets/EverMind-AI/EverMemBench-Static.textquestion-answering1K<n<10K0 likes218 downloads7mo agoHugging Face15NDe2V /methods2test_small_static methods2test_small (cleaned) + static method-level context Static context for unit-test generation, computed with tree-sitter from each row's source only (no repository checkout; target is never read). Configs default: train / validation / test from NDe2V/methods2test_small_cleaned @ 3d6a2f2061736202832b7ac2bb4e84e1c70668e8 methods2test_runnable: test from andstor/methods2test_runnable (fm+fc+c+m+f+t+tc) @ ed8cded666b2f7d101ba04a7a8d121b6249d8c39… See the full description on the dataset page: https://huggingface.co/datasets/NDe2V/methods2test_small_static.text10K<n<100K0 likes202 downloads17d agoHugging Face16Purple69 /aera_semi_pnp_dr_16_06_2026_skip3_delta_no_go_home_no_static_smoothed_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AR4_MK3", "total_episodes": 2855, "total_frames": 428406, "total_tasks": 156, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 23, "splits": { "train": "0:2855" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Purple69/aera_semi_pnp_dr_16_06_2026_skip3_delta_no_go_home_no_static_smoothed_v2.imagerobotics100K<n<1M0 likes142 downloads3mo agoHugging Face17msr-spare-1 /spare-gpt55-static-corpus SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games This public dataset contains 7,872 validated Python game environments for actor-only SPARE training. Six cognitive skills, exactly 1,312 environments per skill Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0 Maximum 25 turns and 32K generation context Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/spare-gpt55-static-corpus.textreinforcement-learning1K<n<10K0 likes142 downloads3mo agoHugging Face18meerkat-ml /component-static-buildsimagen<1K0 likes113 downloads3y agoHugging Face19SHEC3R /HEC3R-ckpt-real_abl_static_priortextn<1K0 likes113 downloads17d agoHugging Face20NikolaosDionelis2023 /data_statics"Data Statics" data for PhilEO Bench The Data Statics for the PhilEO Bench evaluation framework Paper "Scaling-up the Pretraining of the Earth Observation Foundation Model PhilEO to the MajorTOM Dataset" - ESA Φ-lab http://github.com/ESA-PhiLab/PhilEO-MajorTOM Paper: http://arxiv.org/pdf/2506.14765 Also: Paper: http://arxiv.org/pdf/2506.14765v1 Scaling-Up the Pretraining of the Earth Observation Foundation Model PhilEO to the MajorTOM Dataset Also: PhilEO Bench:… See the full description on the dataset page: https://huggingface.co/datasets/NikolaosDionelis2023/data_statics.textn<1K1 likes107 downloads1y agoHugging Face21Purple69 /aera_semi_pnp_dr_02_05_2026_skip3_delta_no_go_home_no_static_smoothedimage100K<n<1M0 likes102 downloads5mo agoHugging Face22StaticExposure /benji-lora-datasetimagen<1K0 likes89 downloads2mo agoHugging Face23yahyaabd /bps-statictable-qrels-alltabular10M<n<100M0 likes87 downloads1y agoHugging Face24hyperdemocracy /usc-vecs-v1-chunks-v1-s4096-o512-sentence-transformers-static-retrieval-mrl-en-v1text100K<n<1M0 likes81 downloads1y agoHugging Face25hyperdemocracy /usc-vecs-v1-chunks-v1-s2048-o256-sentence-transformers-static-retrieval-mrl-en-v1text100K<n<1M0 likes69 downloads1y agoHugging Face26yahyaabd /query-hard-pos-neg-doc-pairs-statictabletext10K<n<100K0 likes44 downloads2y agoHugging Face27code2lora /code2lora-static Code2LoRA-Static — static track (single snapshot) The Static track of RepoPeftBench exactly as consumed by the paper's main-results table (Code2LoRA-Static, single anchor snapshot per repository). Contains only the paper-used QnAs (post quality-filter). For the out-of-distribution slice see code2lora/repopeftbench-ood. config split rows qna train 39,612 qna cr_val 6,213 qna cr_test 6,414 qna ir_val 4,833 qna ir_test 5,222 repos train / cr_val / cr_test 409 /… See the full description on the dataset page: https://huggingface.co/datasets/code2lora/code2lora-static.tabular10K<n<100K0 likes44 downloads3mo agoHugging Face28quill1112 /methods2test_small_static methods2test_small (cleaned) + static method-level context Static context for unit-test generation, computed with tree-sitter from each row's source only (no repository checkout; target is never read). Configs default: train / validation / test from NDe2V/methods2test_small_cleaned @ 3d6a2f2061736202832b7ac2bb4e84e1c70668e8 methods2test_runnable: test from andstor/methods2test_runnable (fm+fc+c+m+f+t+tc) @ ed8cded666b2f7d101ba04a7a8d121b6249d8c39… See the full description on the dataset page: https://huggingface.co/datasets/quill1112/methods2test_small_static.text10K<n<100K0 likes44 downloads17d agoHugging Face29hyperdemocracy /usc-vecs-v1-chunks-v1-s8192-o512-sentence-transformers-static-retrieval-mrl-en-v1text100K<n<1M0 likes39 downloads1y agoHugging Face30THULab /PickPlace_staticMarker PickPlace_staticMarker (TsFile) Apache TsFile version of BlankHead/PickPlace_staticMarker. Overview A LeRobot robot manipulation dataset. Each frame holds the commanded action and observed observation.state joint positions; camera views are stored as videos in the original dataset. Episodes: 76 Frames: 26911 Sampling rate: 30 fps Tasks: 1 Split: a single train split Robot: so_follower Cameras (not uploaded): wrist, up, side Schema (TsFile structure)… See the full description on the dataset page: https://huggingface.co/datasets/THULab/PickPlace_staticMarker.tabularroboticsn<1K0 likes38 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.