Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ByteDance-Seed /EdgeBench Overview EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.texttext-generationn<1K86 likes4.7k downloads3mo agoHugging Face02yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes4.2k downloads7mo agoHugging Face03Harvard-Edge /Wake-Vision-Train-Largeimage1M<n<10M2 likes1.6k downloads2y agoHugging Face04Harvard-Edge /Wake-Vision Dataset Card for Wake Vision Dataset Description "Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally, it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age, subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.imageimage-classification1M<n<10M11 likes1.4k downloads11mo agoHugging Face05MedOtter /Prostate-Anatomical-Edge-Cases Prostate-Anatomical-Edge-Cases Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases — a TCIA collection of pelvic radiotherapy planning CT with manually contoured organs at risk, curated so that most cases contain anatomy known to break autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023). Read before using — the name is misleading in two ways: This is CT, not MRI. Despite "Prostate" in the name it is not a prostate mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.imageimage-segmentationn<1K1 likes1.3k downloads3mo agoHugging Face06BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes610 downloads7mo agoHugging Face07UTSCybeR /Edge-Computing-JEV EdgeIntent v1 EdgeIntent v1 is a benchmark of natural-language requests to edge services, each paired with the typed intent contract it expresses. It was built for the paper Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration Delong Li, Xu Wang, Haochen Gong, Rui Lang, and Guangsheng Yu. University of Technology Sydney. arXiv: 2609.22753 Code, evaluation harness, and reproduction instructions:… See the full description on the dataset page: https://huggingface.co/datasets/UTSCybeR/Edge-Computing-JEV.texttext-classification10K<n<100K0 likes509 downloads9d agoHugging Face08Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes470 downloads7mo agoHugging Face09ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M1 likes367 downloads4mo agoHugging Face10chico-research /EdgeBench-Home EdgeBench-Home EdgeBench-Home is one bilingual smart-home agent benchmark with seven task families and two evaluation views. Dataset size 420 Chinese canonical cases 420 English canonical cases 840 canonical bilingual cases in total Each canonical case has both a State-Given (SG) and a State-Seeking (SS) view; these are derived views, not additional cases. Each language contains 60 cases in each task family. Chinese and English records are aligned by case ID and… See the full description on the dataset page: https://huggingface.co/datasets/chico-research/EdgeBench-Home.textother1K<n<10K1 likes345 downloads8d agoHugging Face11everycure /kg-edges Dataset Card for Every Cure Integrated Knowledge Graph (Edges) Dataset Summary The Every Cure KG is currently (as of February 2026) essentially an integrated, simplified and filtered merged KG comprising ROBOKOP and RTX-KG2. The "Edges" dataset contains the records for all edges in the graph, including metadata. See nodes dataset for the corresponding set of nodes. Source Data Attribution First-level knowledge sources Primary knowledge sources… See the full description on the dataset page: https://huggingface.co/datasets/everycure/kg-edges.tabular10M<n<100M1 likes304 downloads5mo agoHugging Face12DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes272 downloads7mo agoHugging Face13kozo2 /metabolomics-edges-expected-ge5 Cross-study metabolomics co-response edges (expected frequency ≥ 5) 620,265 edges over 34,378 nodes, drawn from pairwise metabolite co-response statistics across MetaboLights and Metabolomics Workbench studies, together with the node properties, two PyTorch Geometric graphs, and the full pipeline that produces them. A node is one differential comparison within one study assay — MTBLS1285_0001_00000028 is study MTBLS1285, assay 0001, feature 00000028; ST002832_AN004625_00002191… See the full description on the dataset page: https://huggingface.co/datasets/kozo2/metabolomics-edges-expected-ge5.tabular100K<n<1M0 likes269 downloads18d agoHugging Face14ahuang11 /tiger_layer_edgesAn unofficial re-packaged parquet files of TIGER/Line® Edges data provided by the US Census Bureau. See LICENSE.pdf for more details. tabular10M<n<100M1 likes229 downloads3y agoHugging Face15future-edge-group /ipulse-ai-historical-consensus-snapshots iPulse AI Historical Consensus Snapshots This dataset contains immutable, historical consensus outputs produced by iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is intended to make selected research outputs inspectable without disclosing raw prompts, individual advisor responses, proprietary price paths, licensed market-data payloads, or internal infrastructure. The first snapshot contains 746 asset-horizon records for 373 assets across 1-year and… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-historical-consensus-snapshots.tabularn<1K0 likes229 downloads6d agoHugging Face16JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes222 downloads5mo agoHugging Face17Edge0 /ark-asr-open-asr-leaderboard-results ARK-ASR Open ASR Leaderboard Results This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits. These files are intended for Open ASR Leaderboard maintainer verification. Scoring summary from normalizer.eval_utils.score_results: Split WER RTFx ami/test 10.02 352.12 earnings22/test 9.77 331.88 gigaspeech/test 8.00 217.72 librispeech/test.clean 1.53 412.12 librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.tabular10K<n<100K13 likes216 downloads4mo agoHugging Face18Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes202 downloads4mo agoHugging Face19mbrt /domain-resurrect-edges Dataset Card for domain-resurrect-edges This dataset contains link counts between domains on the Internet. The data is based on CommonCrawl. See mbrt/domain-resurrect for the companion dataset with scored domains based on Page Rank, and the Blog post on how this was computed. Dataset Details Dataset Description This dataset is a processed version of the CommonCrawl September crawl. Each row is the count of how many hyperlinks exist between the source… See the full description on the dataset page: https://huggingface.co/datasets/mbrt/domain-resurrect-edges.text1B<n<10B0 likes168 downloads1y agoHugging Face20ArkhAngelLifeJiggy /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes160 downloads13d agoHugging Face21istLab /edgelet-imu-gestures edgelet IMU Gestures 9-axis IMU recordings of six hand gestures, made with an Arduino Nano 33 BLE. This is the training data for the MLIoT course labs (edgelet); it is the same data as edgelet's built-in base dataset. Classes (6): idle, circle, leftright, updown, snake, twist Channels (9): accX accY accZ (m/s²), gyrX gyrY gyrZ (deg/s), magX magY magZ (µT) Sampling rate: 100 Hz. Each recording is about 10 s. How the data was split Test set withheld. The full… See the full description on the dataset page: https://huggingface.co/datasets/istLab/edgelet-imu-gestures.tabularn<1K0 likes153 downloads12d agoHugging Face22xtro-edge /v1imagen<1K0 likes147 downloads14d agoHugging Face23edgevane /edgevane-trainset-1Dataset with wikipedia fitst paragraph and public literature texttext-generation1M<n<10M0 likes143 downloads24d agoHugging Face24huggan /edges2shoes Citation @article{pix2pix2017, title={Image-to-Image Translation with Conditional Adversarial Networks}, author={Isola, Phillip and Zhu, Jun-Yan and Zhou, Tinghui and Efros, Alexei A}, journal={CVPR}, year={2017} } text10K<n<100K2 likes142 downloads4y agoHugging Face25Cheva123 /jamjuri-edge-v4-stage2-datasets Jamjuri-Edge V4 — Stage 2 Datasets (E1 / E2 / E3) Cleaned, deduplicated, non-thinking SFT datasets for the Jamjuri-Edge V4 Stage-2 experts. Series: part of JamjuriEDGE (4B) — collection · series card ✦ English Overview Three SFT datasets used to train the Stage-2 experts of Cheva123/Jamjuri-EDGE-Preview-100, plus the Stage-1 curriculum mixture (data/stage1/train.parquet) that trained the shared Stage-1 parent. Every row is pre-rendered with… See the full description on the dataset page: https://huggingface.co/datasets/Cheva123/jamjuri-edge-v4-stage2-datasets.tabulartext-generation100K<n<1M0 likes141 downloads18d agoHugging Face26PureOne /friendship-graph-modular-edge-irregularity-proof Defect Conservation and Exact Modular Edge-Irregularity Strength of Friendship Graphs Public AI-friendly research release · candidate proof · independently verifiable artifacts This repository contains a complete candidate resolution of Open Problem 3.3 from Koam, Ahmad, Bača, and Semaničová-Feňovčíková, AIMS Mathematics 8(1), 2023, concerning the modular edge irregularity strength of friendship graphs. Main candidate theorem For the friendship graph (F_n=K_1\vee… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/friendship-graph-modular-edge-irregularity-proof.text0 likes136 downloads29d agoHugging Face27svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes133 downloads5mo agoHugging Face28NeroSeungSan /synthengine-cot-edge-case-v1 SynthEngine CoT Edge Case Dataset v1.0 Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases. 🔗 Full dataset (1000 records) available on Gumroad This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0. 🎯 Why This Dataset? In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.text10K<n<100K0 likes122 downloads4mo agoHugging Face29hookprobe /edge-ids-threats HookProbe Edge IDS Threat Telemetry Real-world, anonymised threat verdicts from the HookProbe production edge intrusion-detection system. Unlike synthetic lab datasets (CICIDS2017, UNSW-NB15, Kitsune) this is what an actual edge sensor mesh observes on the open internet, labelled by the SENTINEL ensemble (isolation forest + calibrated naive-Bayes) that ships with HookProbe. Sensor: Raspberry Pi edge node + NAPSE AI-native flow classifier Enrichment: RDAP country + ASN lookups… See the full description on the dataset page: https://huggingface.co/datasets/hookprobe/edge-ids-threats.tabulartabular-classification1M<n<10M0 likes113 downloads10d agoHugging Face30Inkwell-Software /screenplay-format-edge-cases Screenplay Format Edge Cases 48 original Fountain specimens in 24 contrast pairs — two near-identical inputs per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs the one difference changes how the lines are classified. In the other 7 it changes the surface and the labels hold: a lowercase scene prefix, a cue extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline note and centered-text markers. Version: 1.0.0 · Maintainer:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-format-edge-cases.texttext-classificationn<1K1 likes85 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.