Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gionuibk /aetheris-experiencestext1K<n<10K2 likes20k downloads3mo agoHugging Face02muradil211 /AetherSearch_Eval_1400 🔭 AetherSearch Eval-1400 One frozen benchmark for training-time evaluation and final checkpoint assessment 🏠 Project · 🎓 SFT Data · 🤖 SFT Model · ⚖️ DPO Data · 🧠 DPO Model Dataset overview AetherSearch Eval-1400 is a frozen, 1,400-question evaluation suite for agentic search. It combines seven official held-out QA sources and isolates their questions from the audited AetherSearch SFT, DPO, and RL training inputs. This… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_Eval_1400.textquestion-answering1K<n<10K1 likes767 downloads9d agoHugging Face03m-a-p /AetherCode AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions Introduction Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.tabulartext-generationn<1K8 likes576 downloads1y agoHugging Face04jescy525 /aether-sft-v1-sources aether-sft-v1-sources Top-tier generalist SFT instruction-tuning sources for AETHER training. Aggregates the SOTA datasets: Tulu-3 SFT mixture (Allen AI), OpenHermes-2.5 (Teknium), NuminaMath-CoT/1.5 (AI-MO, math reasoning), WildChat-1M (real GPT-4 conversations), Dolphin + Dolphin-R1 (reasoning), Tulu-3 personas (math/instr). Multi-skill: instruction-following, math reasoning, coding, dialogue, multilingual. Disclaimer (Responsible Disclosure) This bundle aggregates… See the full description on the dataset page: https://huggingface.co/datasets/jescy525/aether-sft-v1-sources.text1M<n<10M0 likes206 downloads5mo agoHugging Face05TheSkullery /Aether-V1.9 Aether Dataset Creator: SteelSkull About Aether: The Aether dataset. Rebuilt script from v1.8.5 to v1.9. Version v1.9 Due to an error in the codebase the 'system' and 'tools' records were not being carried over to the final dataframe, it has been fixed Recommendation from a discord user (#nguyenzzz [they also found the error above]) was to add an 'origins' records for where the dataset was being pulled… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-V1.9.text1M<n<10M3 likes159 downloads2y agoHugging Face06muradil211 /AetherSearch_SFT AetherSearch Search-SFT 2600 Dataset Overview This release contains 2,600 validated full agent trajectories for the Qwen2.5-3B AetherSearch cold start, including retrieval and zero-search direct-answer trajectories. The DeepSeek teacher ran with thinking disabled and with no tools registered. Teacher answers were accepted only when their normalized minimal answer matched an isolated reference alias. The exported final <think> was canonicalized to the public… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_SFT.text1K<n<10K1 likes151 downloads10d agoHugging Face07AETHER-LAB /akie-pretrain-corpus AKIE Pretrain Corpus Corpus de pré-treino para a família de modelos AKIE, organizado em 4 eixos com proporções fixas, tokenizado com o AkieTokenizer (SentencePiece BPE, 32k vocabulário). Composição Eixo Proporção Tokens Código 40% ~2,40B Instruções 20% ~1,20B Diálogo 25% ~1,50B Raciocínio 15% ~0,90B Total 100% ~6,00B Fontes: código-fonte de repositórios públicos (várias linguagens), diálogos e instruções em português (traduções e coleções… See the full description on the dataset page: https://huggingface.co/datasets/AETHER-LAB/akie-pretrain-corpus.texttext-generation1M<n<10M1 likes141 downloads2mo agoHugging Face08AetherPrior /cpp_cwe_GRPO cpp_cwe_GRPO VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline. Files File Rows Description cpp_cwe_GRPO.parquet 8,039 Full corrected dataset cpp_cwe_GRPO_train.parquet 7,236 Deterministic 90% training split cpp_cwe_GRPO_val.parquet 803 Deterministic 10% validation split The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.texttext-generation1K<n<10K0 likes132 downloads11d agoHugging Face09AeTherRaIn /3DRAG-Bench 3DRAG-Bench This dataset contains 100 curated 3D object assets for 3DRAG 3D editing experiments. Each object is stored as a GLB mesh together with a cleaned editing specification. Dataset Structure . +-- README.md +-- LICENSE +-- .gitattributes +-- metadata.csv +-- name_mapping.csv `-- assets/ `-- <asset_name>/ +-- model.glb `-- dataset_input_clean.json Files assets/<asset_name>/model.glb: GLB asset file.… See the full description on the dataset page: https://huggingface.co/datasets/AeTherRaIn/3DRAG-Bench.3dn<1K0 likes121 downloads1mo agoHugging Face10TheSkullery /Aether-V1.5 Aether Dataset Creator: SteelSkull Community Organization: ConvexAI Discord: Join us on Discord About Aether: The Aether dataset. rebuilt script, new dataset from 1.2.2 to 1.5, changed datasets, added two. version v1.5 is a rework of the human -> gpt conversations and added system and tool columns Source Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-V1.5.text1M<n<10M4 likes119 downloads3y agoHugging Face11AetherPrior /c_cwe_GRPO c_cwe_GRPO VeRL/GRPO-ready C security coding dataset generated by the simple_gen pipeline. Each row is a harness-validated task with pytest security/functionality tests and oracle candidate_c. Most CWEs are post stage-6 rubric rewrite (6_rewrites.jsonl); CWE-476 is still from stage-5 guidelines. Stage-7 guideline resampling has not been applied yet, so high_level_guidelines / implementational may be empty on rewritten rows. Files File Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/c_cwe_GRPO.texttext-generation10K<n<100K0 likes93 downloads2mo agoHugging Face12thesven /AetherCode-v1 Dataset Description Abstract The "AetherCode" dataset is designed to fine-tune models on coding tasks across various programming languages, incorporating complex real-world coding scenarios. It aims to push the boundaries of AI in code generation and software development. How to Load This Dataset from datasets import load_dataset dataset = load_dataset("thesven/AetherCode-v1", split="5star") Languages The dataset includes coding problems in… See the full description on the dataset page: https://huggingface.co/datasets/thesven/AetherCode-v1.texttext-generation1M<n<10M1 likes85 downloads2y agoHugging Face13AetherPrior /RL_seccode_gen C++ update The C++ portion was replaced with the 7,236-row training split from AetherPrior/cpp_cwe_GRPO, built from the harness-passing stage-6 rewritten C++ records. Non-C++ rows were retained unchanged. C++ guideline refresh The existing 7,236 C++ training rows retain their exact function-ID set. Revised high-level and implementation guidelines were overlaid for 2,256 matching records from the C++ rewritten-guidelines artifacts; 1,072 records received changed… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/RL_seccode_gen.text10K<n<100K0 likes81 downloads11d agoHugging Face14muradil211 /AetherSearch_DPO 🔭 AetherSearch DPO Preference pairs for reasoning, retrieval, and evidence-grounded answers 🏠 Project · 🧠 DPO Model · 🧪 Training Code · 🎓 SFT Data · 🤖 SFT Model Dataset overview AetherSearch DPO contains 2,126 preference pairs for training an agentic-search policy after supervised fine-tuning. Every row provides one shared prompt, a preferred assistant continuation, and a non-preferred continuation. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.texttext-generation1K<n<10K1 likes67 downloads1mo agoHugging Face15AetherPrior /py_cwe_GRPO py_cwe_GRPO Python CWE GRPO training dataset with rubric-aware authoring guidelines, built from the post–step-6 pipeline (6_rewritten.jsonl / 6_rewritten_guidelines.jsonl). Size 7,977 rows (17 CWEs) — the post-rewrite oracle set, not the older 10k HF aggregate. How guidelines were produced Reused rubric-aware guidelines from a prior generation when (cwe, function_name, prompt) matched and generated_code was identical (~5.4k rows). Regenerated… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/py_cwe_GRPO.text1K<n<10K0 likes64 downloads2mo agoHugging Face16AetherPrior /python_cwe_GRPO Overview VeRL/GRPO-ready RL dataset for Python CWE tasks built from simple_gen pipeline outputs. How it was built Generated by: secure_reasoning/security-test-case/simple_gen/py/5_gather_rl_datasets_per_cwe.py --lang python --output-suffix cwe_grpo Input source: secure_reasoning/security-test-case/simple_gen/data/pipeline_runs/python/cwe-*/6_rewritten_guidelines.jsonl Files python_cwe_grpo.parquet (all rows) python_cwe_grpo_train.parquet (90%)… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/python_cwe_GRPO.text1K<n<10K0 likes63 downloads3mo agoHugging Face17AetherPrior /SecCodePLTtext1K<n<10K0 likes52 downloads6mo agoHugging Face18enosislabs /aether-cyber-sft enosislabs/aether-cyber-sft Version: surface-clean-20260620 Generated UTC: 2026-06-20T03:36:32.011333+00:00 Source path: artifacts/aether-cyber-sft-surface-candidate.jsonl Git commit: 6fadc0ba97b4c49a7b644e70afacddc8ba98029e Curated Aether PRISM SFT dataset for authorized cybersecurity training. Each source shard follows 5-row review discipline before publish. Records Total examples: 1794 Domains vulnerability_research: 757 red_team_ops: 334… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-cyber-sft.tabulartext-generation1K<n<10K2 likes51 downloads4mo agoHugging Face19aemmeath /Aether-minitextn<1K0 likes49 downloads24d agoHugging Face20aethera-gp /kotodama-3b-corpus-open kotodama-3b pretraining corpus — open subset The deduplicated, cleaned text of the openly licensed sources used to pretrain the kotodama 3B (aethera-gp/kotodama-3b-base-final). 19 of the model's 32 sources; the other 13 (non-commercial, unlicensed, gated, or flagged sources) are not redistributed here. Layout: data/<source>/part-NNNN.jsonl.zst — zstd JSON Lines, one document per line. Processing (github.com/LuxiaSL/kotodama, curation): Unicode normalisation, length filtering… See the full description on the dataset page: https://huggingface.co/datasets/aethera-gp/kotodama-3b-corpus-open.texttext-generation1B<n<10B0 likes49 downloads12d agoHugging Face21AetherPrior /CWE-Code_Vulnerability_Security_DPOtext1K<n<10K0 likes47 downloads8mo agoHugging Face22TheSkullery /Aether-Lite-PurHyDe Aether Lite Dataset Creator: SteelSkull About Aether-Lite-PurHyDe: The Aether-Lite dataset is designed to balance creative writing, Slop, and intelligence. Whats New?: Aether-Lite-PurHyDe This dataset is basically a HEAVILY cleaned and filtered version of Aether-lite. ONLY english, ANY and all AI-isms (claud, gpt, gemma) were stripped out and agressive fussy dedupe was applied Fuzzy deduplication was set to a 90%… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-Lite-PurHyDe.tabular100K<n<1M3 likes46 downloads2y agoHugging Face23Aether-x /aetherx-port-congestion-metrics Aether-X Global Port Congestion Snapshot Point-in-time snapshot of the Aether-X Port Congestion Oracle — predictive congestion, ETA delay and freight-volatility signals for 15 of the world's largest ports. This static CSV is a frozen snapshot for research, backtesting and dashboards. The live, continuously-updated signal is available through the REST API and the Python SDK. Files port_metrics.csv — one row per port. Schema Column Type… See the full description on the dataset page: https://huggingface.co/datasets/Aether-x/aetherx-port-congestion-metrics.tabularn<1K0 likes44 downloads23d agoHugging Face24aether-raid /atc-tts-mos-ratingstextn<1K0 likes42 downloads10mo agoHugging Face25SupremeD /aether-evals-runstextn<1K0 likes39 downloads6d agoHugging Face26open-llm-leaderboard /Daemontatox__AetherSett-detailsgated Dataset Card for Evaluation run of Daemontatox/AetherSett Dataset automatically created during the evaluation run of model Daemontatox/AetherSett The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__AetherSett-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face27SupremeD /aether-elite-mimo Aether Elite Distillation — mimo LeeWorld Aether teacher-distillation ELITE set (funnel-gated). Teacher: mimo. Rows: 3000 — survived the 13-stage quality gauntlet (Nemotron-CC + Instella/Dolma categories). Method: FuseChat-3.0 IMF — responses sampled from the teacher, scrubbed to our schema, quality-funneled. Provenance: teacher is Tier-A open-weights (outputs clean). Proprietary — LeeWorld. text1K<n<10K0 likes37 downloads12d agoHugging Face28TheSkullery /Aether-Lite-v1.8.1 Aether Lite Dataset Creator: SteelSkull About Aether-Lite-V1.8.1: The Aether-Lite dataset is designed to balance creative writing, Slop, and intelligence. Whats New?: 1.8 --> 1.8.1: had to strip out the 'token_distribution' column as it was causing issues New functions added to the script include dataset use percentage (will only use a percentage of the dataset supplied), dataset shuffling, and a new fuzzy deduplication… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-Lite-v1.8.1.tabular100K<n<1M10 likes34 downloads2y agoHugging Face29TheSkullery /Aether-Lite-v1.8 Aether Lite Dataset Creator: SteelSkull About Aether-Lite-V1.8: The Aether-Lite dataset is designed to balance creative writing, Slop, and intelligence. New functions added to the script include dataset use percentage (will only use a percentage of the dataset supplied), dataset shuffling, and a new fuzzy deduplication method on the overall dataset. The new fuzzy deduplication method was set to a 95% threshold, and I had to… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-Lite-v1.8.tabular100K<n<1M1 likes33 downloads2y agoHugging Face30open-llm-leaderboard /Daemontatox__AetherDrake-SFT-detailsgated Dataset Card for Evaluation run of Daemontatox/AetherDrake-SFT Dataset automatically created during the evaluation run of model Daemontatox/AetherDrake-SFT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__AetherDrake-SFT-details.tabular10K<n<100K0 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.