Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01issdandavis /scbe-aethermoore-training-data Status: canonical. Primary public training dataset for SCBE-AETHERMOORE and the most-used repo in this account. Other scbe-* dataset repos are experiment-specific slices. SCBE-AETHERMOORE Training Dataset Supervised fine-tuning (SFT) dataset for the SCBE-AETHERMOORE hyperbolic geometry AI safety and governance framework. Overview This dataset contains 10,978 training pairs spanning the full SCBE-AETHERMOORE system: 14-layer architecture knowledge, Six Sacred… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-aethermoore-training-data.text-generation10K<n<100K2 likes3k downloads20d agoHugging Face02m-a-p /AetherCode AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions Introduction Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.tabulartext-generationn<1K8 likes576 downloads1y agoHugging Face03AETHER-LAB /akie-pretrain-corpus AKIE Pretrain Corpus Corpus de pré-treino para a família de modelos AKIE, organizado em 4 eixos com proporções fixas, tokenizado com o AkieTokenizer (SentencePiece BPE, 32k vocabulário). Composição Eixo Proporção Tokens Código 40% ~2,40B Instruções 20% ~1,20B Diálogo 25% ~1,50B Raciocínio 15% ~0,90B Total 100% ~6,00B Fontes: código-fonte de repositórios públicos (várias linguagens), diálogos e instruções em português (traduções e coleções… See the full description on the dataset page: https://huggingface.co/datasets/AETHER-LAB/akie-pretrain-corpus.texttext-generation1M<n<10M1 likes141 downloads2mo agoHugging Face04AetherPrior /cpp_cwe_GRPO cpp_cwe_GRPO VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline. Files File Rows Description cpp_cwe_GRPO.parquet 8,039 Full corrected dataset cpp_cwe_GRPO_train.parquet 7,236 Deterministic 90% training split cpp_cwe_GRPO_val.parquet 803 Deterministic 10% validation split The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.texttext-generation1K<n<10K0 likes132 downloads11d agoHugging Face05AetherPrior /c_cwe_GRPO c_cwe_GRPO VeRL/GRPO-ready C security coding dataset generated by the simple_gen pipeline. Each row is a harness-validated task with pytest security/functionality tests and oracle candidate_c. Most CWEs are post stage-6 rubric rewrite (6_rewrites.jsonl); CWE-476 is still from stage-5 guidelines. Stage-7 guideline resampling has not been applied yet, so high_level_guidelines / implementational may be empty on rewritten rows. Files File Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/c_cwe_GRPO.texttext-generation10K<n<100K0 likes93 downloads2mo agoHugging Face06thesven /AetherCode-v1 Dataset Description Abstract The "AetherCode" dataset is designed to fine-tune models on coding tasks across various programming languages, incorporating complex real-world coding scenarios. It aims to push the boundaries of AI in code generation and software development. How to Load This Dataset from datasets import load_dataset dataset = load_dataset("thesven/AetherCode-v1", split="5star") Languages The dataset includes coding problems in… See the full description on the dataset page: https://huggingface.co/datasets/thesven/AetherCode-v1.texttext-generation1M<n<10M1 likes85 downloads2y agoHugging Face07muradil211 /AetherSearch_DPO 🔭 AetherSearch DPO Preference pairs for reasoning, retrieval, and evidence-grounded answers 🏠 Project · 🧠 DPO Model · 🧪 Training Code · 🎓 SFT Data · 🤖 SFT Model Dataset overview AetherSearch DPO contains 2,126 preference pairs for training an agentic-search policy after supervised fine-tuning. Every row provides one shared prompt, a preferred assistant continuation, and a non-preferred continuation. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.texttext-generation1K<n<10K1 likes67 downloads1mo agoHugging Face08enosislabs /aether-cyber-sft enosislabs/aether-cyber-sft Version: surface-clean-20260620 Generated UTC: 2026-06-20T03:36:32.011333+00:00 Source path: artifacts/aether-cyber-sft-surface-candidate.jsonl Git commit: 6fadc0ba97b4c49a7b644e70afacddc8ba98029e Curated Aether PRISM SFT dataset for authorized cybersecurity training. Each source shard follows 5-row review discipline before publish. Records Total examples: 1794 Domains vulnerability_research: 757 red_team_ops: 334… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-cyber-sft.tabulartext-generation1K<n<10K2 likes51 downloads4mo agoHugging Face09aethera-gp /kotodama-3b-corpus-open kotodama-3b pretraining corpus — open subset The deduplicated, cleaned text of the openly licensed sources used to pretrain the kotodama 3B (aethera-gp/kotodama-3b-base-final). 19 of the model's 32 sources; the other 13 (non-commercial, unlicensed, gated, or flagged sources) are not redistributed here. Layout: data/<source>/part-NNNN.jsonl.zst — zstd JSON Lines, one document per line. Processing (github.com/LuxiaSL/kotodama, curation): Unicode normalisation, length filtering… See the full description on the dataset page: https://huggingface.co/datasets/aethera-gp/kotodama-3b-corpus-open.texttext-generation1B<n<10B0 likes49 downloads12d agoHugging Face10jescy525 /aether-family-trading-shared aether-family-trading-shared AETHER family SFT dataset — group trading. Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85). Schema: { "messages": [{"role": "system|user|assistant", "content": "..."}], "task_type": "function_calling|code|reasoning_cot|...", "source_ds": "<HF dataset_id>", "lang": "en|fr|...", "system_source": "archon_default|overridden_from_source" } Generated by prepare_sft.py pipeline (2026-05-25). texttext-generation10K<n<100K0 likes26 downloads5mo agoHugging Face11lonestar155 /aether-build-protocol-examples Aether Build Protocol Examples Aether Build Protocol Examples is a small public dataset of machine-readable physical build intent artifacts. It is designed for AI developers, agent-framework builders, CAD/design workflows, fabrication review systems, and researchers studying machine-to-machine physical transaction protocols. GitHub source of truth: https://github.com/chevy155/Aether-build-protocol Live demo: https://huggingface.co/spaces/lonestar155/aether-cad-to-agent-sandbox Open… See the full description on the dataset page: https://huggingface.co/datasets/lonestar155/aether-build-protocol-examples.text-generationn<1K0 likes22 downloads5mo agoHugging Face12wincode /aetherstory-data AetherStory Dataset A procedurally generated corpus of unique fantasy fables, purpose-built to train the AetherStory storyteller model. Every story is synthesised on the fly from a combinatorial space of realms, creatures, character archetypes, magic systems, and plot scaffolds. Why this dataset is unique It is not scraped from the web. Each fable is constructed by code from hand-written ingredients, so: there is no copyright risk — every word is original or… See the full description on the dataset page: https://huggingface.co/datasets/wincode/aetherstory-data.text-generation10K<n<100K0 likes22 downloads3mo agoHugging Face13jescy525 /aether-sft-v2-mix aether-sft-v2-mix AETHER family SFT dataset — group mix. Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85). Schema: { "messages": [{"role": "system|user|assistant", "content": "..."}], "task_type": "function_calling|code|reasoning_cot|...", "source_ds": "<HF dataset_id>", "lang": "en|fr|...", "system_source": "archon_default|overridden_from_source" } Generated by prepare_sft.py pipeline (2026-05-25). texttext-generation1M<n<10M0 likes21 downloads5mo agoHugging Face14jescy525 /aether-sft-v2-code aether-sft-v2-code AETHER family SFT dataset — group code. Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85). Schema: { "messages": [{"role": "system|user|assistant", "content": "..."}], "task_type": "function_calling|code|reasoning_cot|...", "source_ds": "<HF dataset_id>", "lang": "en|fr|...", "system_source": "archon_default|overridden_from_source" } Generated by prepare_sft.py pipeline (2026-05-25). texttext-generation100K<n<1M0 likes20 downloads5mo agoHugging Face15AetherPrior /js_cwe_GRPO js_cwe_GRPO VeRL/GRPO-ready JavaScript security coding dataset generated by the simple_gen pipeline. Each row is a harness-validated task with Node harness security/functionality tests, oracle candidate_js, and authoring guidelines (high_level_guidelines, implementational). Files File Rows Description js_cwe_GRPO.parquet 4956 Full dataset (shuffled) js_cwe_GRPO_train.parquet 4461 90% train split js_cwe_GRPO_val.parquet 495 10% validation split… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/js_cwe_GRPO.texttext-generation1K<n<10K1 likes20 downloads2mo agoHugging Face16enosislabs /aether-0.8b-cyber-sft enosislabs/aether-0.8b-cyber-sft Version: aether-0.8b-cyber-20260618 Generated UTC: 2026-06-19T02:53:44.849450+00:00 Source path: data/curated Git commit: 3ac5f50de88e43122cfda43d65f1ad01b5febff6 Curated Aether PRISM SFT dataset for authorized cybersecurity training. Each source shard follows 5-row review discipline before publish. Records Total examples: 1900 Domains vulnerability_research: 715 red_team_ops: 335 detection_engineering: 257… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-0.8b-cyber-sft.tabulartext-generation1K<n<10K0 likes16 downloads4mo agoHugging Face17jescy525 /aether-sft-v2-trading aether-sft-v2-trading AETHER family SFT dataset — group trading. Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85). Schema: { "messages": [{"role": "system|user|assistant", "content": "..."}], "task_type": "function_calling|code|reasoning_cot|...", "source_ds": "<HF dataset_id>", "lang": "en|fr|...", "system_source": "archon_default|overridden_from_source" } Generated by prepare_sft.py pipeline (2026-05-25). texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face18jescy525 /aether-sft-v2-multilingual aether-sft-v2-multilingual Multilingual EN+FR subset from CohereForAI/aya_dataset. Total: 5366 samples ChatML. texttext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face19jescy525 /aether-sft-v2-reasoning aether-sft-v2-reasoning AETHER family SFT dataset — group reasoning. Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85). Schema: { "messages": [{"role": "system|user|assistant", "content": "..."}], "task_type": "function_calling|code|reasoning_cot|...", "source_ds": "<HF dataset_id>", "lang": "en|fr|...", "system_source": "archon_default|overridden_from_source" } Generated by prepare_sft.py pipeline (2026-05-25). texttext-generation1M<n<10M0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.