Team Ai
20 results

gem

Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M216 likes18k downloads1mo agoHugging FaceGEM /wiki_linguaWikiLingua is a large-scale multilingual dataset for the evaluation of crosslingual abstractive summarization systems. The dataset includes ~770k article and summary pairs in 18 languages from WikiHow. The gold-standard article-summary alignments across languages was done by aligning the images that are used to describe each how-to step in an article.summarization50 likes17k downloads4y agoHugging FaceGEM /wiki_auto_asset_turk Dataset Card for GEM/wiki_auto_asset_turk Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.text100K<n<1M8 likes12k downloads2y agoHugging FaceGEM /xlsumWe present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.summarization6 likes10k downloads2y agoHugging Faceabotresol /emotion-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes5.3k downloads2mo agoHugging FaceHarland /AudioMCQ-StrongAC-GeminiCoT [ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly. Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis. 🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.audio10K<n<100K7 likes4.8k downloads3mo agoHugging Face