datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
survivaleaglellama8b-eagle-sharegpteagle3-hidden-states-gemma4eagle-grpo-iter19-q4k-uniform-noimatrix-enc
eagle-grpo-iter19 — I-Quality pack, uniform Q4_K, NO imatrix
READ THIS BEFORE COMPARING THIS PACK TO ANY iter_267 NUMBER.
What this is
An I-Quality .iqpt pack of the Oaica V4-Flash iter19 checkpoint (SFT + GRPO final,
DeepSeek-V4-Flash 284B, 43 layers, 256 routed experts/layer), produced by
pipeline/iquality/pack_cbalanced_proposed.py --preset c-balanced-proposed.
Files are AES-256-CTR encrypted (one random IV per file). The manifest mapping
original_path ->… See the full description on the dataset page: https://huggingface.co/datasets/sprapp/eagle-grpo-iter19-q4k-uniform-noimatrix-enc.sbucaptionssec-13f-holdings
SEC Form 13F Hedge Fund Holdings
Quarterly US-listed equity holdings for 9 institutional managers,
reconstructed from their own Form 13F-HR filings with the SEC.
Built for the trackers at y-yin.io/research and
published here because the filings are public domain and the parsing is fiddly
enough to be worth sharing. Read straight from each filing's infotable.xml —
the structured document EDGAR renders its own filing pages from — so no HTML is
scraped.
Coverage: 9 funds, 452… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/sec-13f-holdings.eagle-grpo-iter19-fp4-encysa-web-scrape-dataset-qa-formatted-small-versioneagle-sftr2-iter49-fp4-enceagle-iter267-sra-phase1-enceagle-sft-hfnative-routefix-encKorean_Wikipedia_Dataset_for_GPT2_August_2022
Dataset Card for korean_wikipedia_dataset_for_GPT2
Dataset Description
Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022.
email: oscar.eaglewatch@gmail.com
Dataset Summary
This is to make a pre-trained GPT-2 Korean model
Languages
Korean
Dataset Structure
Data Instances
Train wikipedia article count: 334420
validation wikipedia article count: 83605
Data Fields
'text'
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.eagle2-llamagen2-training-datallama4-sglang-eagle3Eagle-1.8M
Dataset Card for Eagle-1.8M
Dataset Sources
Dataset Name
Sample Number
Note
LLaVA v1.5
665k
Multi-modal conversation
DocVQA
39k
Document understanding
synDog-EN
50k
OCR
ChartQA
28k
Chart understanding
DVQA
25k
Chart understanding
AI2D
15k
Open-Hermes 2.5
ShareGPT-4V
100k
Detailed caption generated by GPT-4V
laion-GPT4V
11k
Detailed caption generated by GPT-4V
LVIS-Instruct4V
220k
Multi-modal conversation
LRV-Instruct
150k
Multi-modal… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/Eagle-1.8M.eagleOsworld_eagle_datasets
welcomer
A small Node.js CLI utility that prints personalised greetings to standard
output. It is used by our team onboarding scripts to welcome new joiners.
Usage
npm start
The script reads a hardcoded list of names and prints one greeting per line.
Status
The current source files (src/greet.js, src/app.js) were written against
ES5 syntax. We are planning a syntax-only modernisation pass to ES6+ before
the v1.1 release; runtime behaviour should… See the full description on the dataset page: https://huggingface.co/datasets/MohanGupta-turing/Osworld_eagle_datasets.qwen3_8b_eagle3-parquet
qwen3_8b_eagle3 (Parquet)
Sharded Parquet conversion of Tengyunw/qwen3_8b_eagle3.
Original distribution is a single ~13 GB JSON file; this repo splits it into
61 Parquet shards of ~10,000 rows each for streaming-friendly access
via the datasets library.
Schema
id: string
conversations: list<struct<from: string, value: string>> (ShareGPT format)
Stats
Rows: 607,865
Shards: 61 (data/train-NNNNN-of-00061.parquet)
Compression: zstd
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/shadowpa0327/qwen3_8b_eagle3-parquet.eagle
Eagle 🦅: Ethical Dataset Given from Real Interactions
Introduction
This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs).
If you use the Eagle dataset in your research, please cite the following:
@inproceedings{Eagle:arxiv:2024,
title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.eagle
EAGLE
Dataset Description
The EAGLE dataset is sourced from an ICLR 2023 paper and is designed for predicting two-dimensional unsteady turbulent flows on unstructured dynamic meshes. The data describes the velocity and pressure fields produced by interactions between a moving flow source and various ground geometries. It contains three types of geometric configurations: Cre, Spl, and Tri.
Paper: EAGLE: Large-scale Learning of Turbulent Fluid Dynamics with Mesh… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/eagle.eagle3-speculative-decoding-energy-sweep
EAGLE3 Speculative Decoding Energy Sweep
Per-config energy/throughput/latency measurements for EAGLE3 speculative decoding
(speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens)
served with sglang, across batch sizes. Collected for an RL project that learns to
pick speculative-decoding parameters to hold GPU energy utilization in a target band.
Model: unsloth/Llama-3.2-1B-Instruct + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft head.
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/eagle3-speculative-decoding-energy-sweep.EagleSFT
Dataset Card for 🦅 EagleSFT
Dataset Summary
This dataset contains 536,231 pairs of human questions and machine-generated responses intended for supervised fine-tuning (SFT) of large language models. The dataset includes both Russian and English content, with linked IDs allowing for cross-lingual analysis. It was created by processing an initial collection of 739,732 human questions posed to LLMs, predominantly in Russian (about 99%) with a small portion in English (about… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/EagleSFT.openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1
OpenAI GSM8K Enhanced with DeepSeek API
🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI.
🔗 Access the Dataset: OpenAI GSM8K Enhanced
What’s Cooking? 🍳
Dataset Specifications
Total Samples: ~10K, with about 8K training and 1K testing entries.
Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.details_cookinai__Bald-Eagle-7B
Dataset Card for Evaluation run of cookinai/Bald-Eagle-7B
Dataset automatically created during the evaluation run of model cookinai/Bald-Eagle-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cookinai__Bald-Eagle-7B.pothole-eagle-datasetEagleX-WorldContinued
Dataset Card for EagleX v2 Dataset
This dataset was used to train RWKV Eagle 7B for continued pretrain of 1.1T tokens (approximately) (boosting it to 2.25T) with the final model being released as RWKV EagleX v2.
Dataset Details
Dataset Description
EagleX-WorldContinued is a pretraining dataset built from many of our datasets over at Recursal AI + a few others.
Curated by: M8than, KaraKaraWitch, Darok
Funded by [optional]: Recursal.ai
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/EagleX-WorldContinued.FinSen
Enhancing Financial Market Predictions: Causality-Driven Feature Selection
Note:[Please help give a Like ❤️ if you think this FinSen dataset is good for you, Thanks:)]
This paper introduces FinSen dataset that revolutionizes financial market analysis by integrating economic and financial news articles from 197 countries with stock market data. The dataset’s extensive coverage spans 15 years from 2007 to 2023 with temporal information, offering a rich, global perspective 160,000… See the full description on the dataset page: https://huggingface.co/datasets/EagleWHLiang/FinSen.Effi-Eagle3-Train-Dataset
EAGLE3 离线特征生成说明与必要文件清单
TL;DR
本仓库提供 DRAgent 2k、ShareGPT 20k、ShareGPT 2k 三份 EAGLE3 离线特征的生成指南、两份确定输入 JSONL,以及固定参考版本的完整 SpecForge 源码。以约 1.04 GB 输入和小型源码归档记录生成方法,替代继续上传约 1.70 TB 的特征;本次没有重新运行 GPU capture。原特征仍保存在本地 /data/yilin/specforge-trim/hidden_states。
复现使用下文 data/ 中的输入和 source/ 中的源码;材料的校验清单是 SHA256SUMS.inputs.json。
已找回 2026-08-21 三份成功生成日志、原生成脚本及两份输入 JSONL。输入文件合计 1,041,157,920 bytes,约 1.04 GB;teacher 是 **Qwen/Qwen3-8B**。生成流程是:先取 N 条可用对话 → seed 42 打乱 → qwen 模板分词与 loss… See the full description on the dataset page: https://huggingface.co/datasets/julyanghar/Effi-Eagle3-Train-Dataset.eagle-mix
Eagle-Mix Dataset
Dataset Description
Eagle-Mix is a comprehensive mixed dataset created for training Eagle models. It combines high-quality conversational data from multiple sources to provide diverse training examples.
Dataset Composition
The dataset is composed of the following sources:
Dataset
Count
Mean Length
Median Length
Max Length
ShareGPT
68,623
6,128
6,445
93,262
UltraChat
207,865
5,686
5,230
53,213
OpenThoughts2-1M
1,143,205
16,175
10… See the full description on the dataset page: https://huggingface.co/datasets/Qinghao/eagle-mix.
