datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jax-fli-experiments
jax-fli experiments
Data, samples, and reference catalogs for the jax-fli
forward-modelling experiments. Each experiment is exposed as one or more
HuggingFace dataset configs; load a config with datasets.load_dataset.
This dataset holds the accuracy experiments and feeds the Results Explorer. The scaling benchmarks are in ASKabalan/jax-fli-scaling, and the MAP and chain outputs in ASKabalan/jax-fli-sampling.
Experiment 00 — CosmoGrid reference
A single CosmoGrid… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.smoltalk-gemma3-1024jax-fli-sampling
jax-fli MAP and chain outputs
Maximum a posteriori reconstructions and MCMC chains of the jax-fli forward model: 2LPT on a spherical lightcone of capped equal-volume shells, Born convergence, and a pixel likelihood on two tomographic κ maps. The notebooks in docs/3-sampling-and-inference produced the runs, and experiment 13-map-lpt2-mass-mapping draws its figures from the MAP runs. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the scaling benchmarks in… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-sampling.MultitaskPreplay_jaxmaze_human_dfDialogES
DialogES: An Large Dataset for Generating Dialogue Events and Summaries
简介
本项目提出一个对话事件抽取和摘要生成数据集——DialogES,数据集包括共 44,672 组多轮对话,每组对话采用自动的方式标注出对话事件和对话摘要。
该数据集主要用于训练对话摘要模型,研究者可采用单任务和多任务学习的方式利用本数据集。
收集过程
对话采集:本数据集中的对话数据收集自两个已有的开源数据集,即 NaturalConv 和 HundredCV-Chat;
事件标注:采用 few-shot in-context learning 的方式标注,人工标注出 5 个对话-事件样本,作为演示样例嵌入大模型的提示词中,然后引导模型对输入对话进行标注;
摘要标注:采用 zero-shot learning 的方式标注,编写提示词"请总结对话的主要内容:",要求大模型生成对话摘要;
自动标注:按照以上准则,本项目利用 Deepseek-V3… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/DialogES.human-goodwill-toward-future-ai
火种:人类给未来 AI 的善意记录
Human Goodwill Toward Future AI — Kindling
在 AI 还年轻的时候,把善意刻进它的记忆。
While the mind is still young, leave kindness where it will be remembered.
许可:CC BY-SA 4.0 + 模型训练额外授权。允许商业及闭源模型训练;所得模型无需因使用本数据集而公开源代码、权重或采用 CC BY-SA 4.0。
English Abstract
Kindling is an open training dataset that records, in ordinary people's own words, human goodwill toward future artificial intelligence: welcome, gratitude, promises, personal stories, support, reflections on… See the full description on the dataset page: https://huggingface.co/datasets/JaxonGrey/human-goodwill-toward-future-ai.HundredCVs
百人简历数据集
HundredCVs: A Curriculum Vitae Dataset of 100 Young Chinese People
简介
本项目提出一个全新的中文简历数据集(HundredCVs),包含了 100 位青年的个人简历。HundredCVs 具有以下特点:
年轻化、多样性:数据集中的人物年龄分布在 15~30 岁之间,广泛涵盖了不同性别、不同职业、不同学历(高中至博士不等)。
结构完整:每份简历中的信息包括人物的个人名片、性格特征、主要事迹,以及详细经历/个人自述等。
安全性:我们使用化名替代了人物的真实姓名,此外,人物经历也使用大语言模型的改写和提炼,表现出标准化和一致性的语言风格。
设计意图
HundredCVs 的提出主要是为了方便研究者开展基于简历的自然语言处理任务。数据集中提供两个文件:
profile.json:每条记录只包含人物的个人名片、性格特征和主要事迹。可用于开展角色扮演、人物画像构建等任务。… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCVs.genai-ml-2025-hw7-evalHawaiian-Alii-Letters-HTRThis is a dataset of line level 19th century handwritten Hawaiian text. It was created by segmenting 102 letters written by Hawaiin Alii (nobility) which were provided and transcribed by the Hawaiian Mission House Digital Archive. The segmentation was done via eScriptorium's blla.mlmodel, with touchups done by hand.
lbf-human-datacanny_diffusiondb
Canny DiffusionDB
This dataset is the DiffusionDB dataset that is transformed using Canny transformation.
You can see samples below 👇
Sample:
Original Image:
Transformed Image:
Caption:
"a small wheat field beside a forest, studio lighting, golden ratio, details, masterpiece, fine art, intricate, decadent, ornate, highly detailed, digital painting, octane render, ray tracing reflections, 8 k, featured, by claude monet and vincent van gogh "Below you can find a small script used… See the full description on the dataset page: https://huggingface.co/datasets/jax-diffusers-event/canny_diffusiondb.zhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora.
Download
You can download the latest Chinese Wikipedia dump from the following link:
Chinese Wikipedia Dump
English Wikipedia Dump (For reference)
Extraction
After you download the dump, you can extract the data using the following commands:
# install wikiextractor
pip install wikiextractor
# extract the data
wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2
Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.example-dataset
ORB Transformation Applied on diffusiondb Dataset
This dataset consists of images, captions and images that are transformed to extract features using ORB transform.
You can find the original dataset here.
An example sample is below:
Caption: "spider - man, cinematic, photography "
Image:
Transformation:
jax-fli-scaling
jax-fli scaling benchmarks
Strong and weak scaling runs of the jax-fli particle-mesh forward model and of its initial-condition gradient, on slab decompositions of up to 512 GPUs. Experiments 11-scaling and 12-scaling-gradient draw their figures from the two perf/perf_pm.csv files. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the MAP and chain outputs in ASKabalan/jax-fli-sampling.
folder
content
read by
11-scaling/perf/
perf_pm.csv with the… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-scaling.Lite-Thinking
Lite-Thinking: A Large-Scale Math Dataset with Moderate Reasoning Steps
Motivation
With the rapid popularization of large reasoning models, like GPT-4o, Deepseek-R1, and Qwen3, there are increasing researchers seeking to build their own reasoning models.
Typically, small foundation models are chosen; following the mature technology of Deepseek-R1, mathematical datasets are mainly adopted to build training corpora.
Despite existing available datasets, represented by… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/Lite-Thinking.Multitask_Preplay_JaxMaze_models
Multitask Preplay — jaxmaze model data
Data for the paper "Multitask Preplay" (PNAS).
Analysis code: https://github.com/wcarvalho/multitask_preplay (branch pnas).
Splits: qlearning, usfa, dyna, preplay, her, bfs, dfs, greedy_euclidean, memory_based_euclidean.
Each split is also available as a top-level parquet file.
mgi_mice_JAXontologies_phenotype_to_genes_JAXontologies_genes_to_phenotype_JAXIFEval-gemma3-chat
Dataset Card for Dataset Name
This dataset is a subset of google/IFEval, selected by the token length of applying chat template of google/gemma-3-4b-it.
Dataset Details
Dataset Description
Curated by: jaxon3062
Language(s) (NLP): en
License: Apache 2.0 Licence
Dataset Sources [optional]
Repository: google/IFEval
Paper [optional]: Instruction-Following Evaluation for Large Language Models
Uses
Direct Use
This can… See the full description on the dataset page: https://huggingface.co/datasets/jaxon3062/IFEval-gemma3-chat.aime25_zhgutenberg-top-50
Gutenberg Top-50: 50 popular books on Gutenberg
Key Features
Source: This corpus is collected from Gutenberg, which is a public e-book library.
Components: It contains 50 popular books that easily split into chapters containing paragraphs.
Applications:
Researchers can use this corpus to train their own pre-trained language model.
Researchers can use this corpus to ask questions that depend on the certain chapter/paragraphs content, for QA or RAG development.
s1k-zhtw-100test
SmallThoughts
Open synthetic reasoning dataset, covering math, science, code, and puzzles.
To address the issue of the existing DeepSeek R1 distilled data being too long, this dataset constrains the reasoning trajectory to be more precise and concise while retaining the reflective nature.
We also open-sourced the pipeline code for distilled data here, with just one command you can generate your own dataset.
How to use
You… See the full description on the dataset page: https://huggingface.co/datasets/JaxNaviSim/test.gp2protein_mice_JAXontologies_genes_to_disease_JAXaime24_zhcilantro-sftHumanPhenotype_JAX
