apollo
Datasets
All datasets matching “apollo”rl-dynamics-gradient-sketches-toy-adam5e5
Gradient sketches of a toy RL run (Adam, lr 5e-5, rep 3)
Random-projection sketches of per-rollout policy gradients from a small GRPO run in Apollo Research's RL-dynamics
project (base model Qwen/Qwen3.5-4B with LoRA adapters, 129,859,584 trainable parameters, 20 training steps,
1,024 rollouts per step). Every gradient was projected with the same subsampled randomised Hadamard transform (SRHT)
to k = 262,144 coordinates, so inner products between sketches approximate inner… See the full description on the dataset page: https://huggingface.co/datasets/apollo-research/rl-dynamics-gradient-sketches-toy-adam5e5.Skylion007-openwebtext-tokenizer-gpt2apollo-openweb-tokenizer-gpt2-to-llama2-2024monology-pile-uncopyrighted-tokenizer-EleutherAI-gpt-neox-20bLongTimeScopeIf you use TimeScope please cite the following:
@misc{zohar2025apollo2,
title = {Apollo2: Exploring the Long-Video Frontier of Large Multimodal Models},
author = {Zohar, Orr and Wang, Xiaohan and Li, Rui and Marafioti, Andrés and Farré, Miquel and Noyan, Merve and von Werra, Leandro and Yeung-Levy, Serena and Wolf, Thomas},
year = {2025},
}
ApolloCorpus
Multilingual Medicine: Model, Dataset, Benchmark, Code
Covering English, Chinese, French, Hindi, Spanish, Hindi, Arabic So far
👨🏻💻Github •📃 Paper • 🌐 Demo • 🤗 ApolloCorpus • 🤗 XMedBench
中文 | English
🌈 Update
[2024.03.07] Paper released.
[2024.02.12] ApolloCorpus and XMedBench is published!🎉
[2024.01.23] Apollo repo is published!🎉
Results
Apollo-0.5B • 🤗 Apollo-1.8B • 🤗 Apollo-2B • 🤗 Apollo-6B • 🤗 Apollo-7B… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloCorpus.
