Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ASKabalan /jax-fli-experiments jax-fli experiments Data, samples, and reference catalogs for the jax-fli forward-modelling experiments. Each experiment is exposed as one or more HuggingFace dataset configs; load a config with datasets.load_dataset. This dataset holds the accuracy experiments and feeds the Results Explorer. The scaling benchmarks are in ASKabalan/jax-fli-scaling, and the MAP and chain outputs in ASKabalan/jax-fli-sampling. Experiment 00 — CosmoGrid reference A single CosmoGrid… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.tabularn<1K1 likes2.9k downloads11d agoHugging Face02Jax-dan /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.texttext-generation10K<n<100K19 likes323 downloads2y agoHugging Face03jaxon3062 /smoltalk-gemma3-1024text100K<n<1M2 likes141 downloads11mo agoHugging Face04ASKabalan /jax-fli-sampling jax-fli MAP and chain outputs Maximum a posteriori reconstructions and MCMC chains of the jax-fli forward model: 2LPT on a spherical lightcone of capped equal-volume shells, Born convergence, and a pixel likelihood on two tomographic κ maps. The notebooks in docs/3-sampling-and-inference produced the runs, and experiment 13-map-lpt2-mass-mapping draws its figures from the MAP runs. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the scaling benchmarks in… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-sampling.tabularn<1K0 likes111 downloads8d agoHugging Face05wcarvalho /MultitaskPreplay_jaxmaze_human_dftabular10K<n<100K0 likes60 downloads1y agoHugging Face06Jax-dan /DialogES DialogES: An Large Dataset for Generating Dialogue Events and Summaries 简介 本项目提出一个对话事件抽取和摘要生成数据集——DialogES,数据集包括共 44,672 组多轮对话,每组对话采用自动的方式标注出对话事件和对话摘要。 该数据集主要用于训练对话摘要模型,研究者可采用单任务和多任务学习的方式利用本数据集。 收集过程 对话采集:本数据集中的对话数据收集自两个已有的开源数据集,即 NaturalConv 和 HundredCV-Chat; 事件标注:采用 few-shot in-context learning 的方式标注,人工标注出 5 个对话-事件样本,作为演示样例嵌入大模型的提示词中,然后引导模型对输入对话进行标注; 摘要标注:采用 zero-shot learning 的方式标注,编写提示词"请总结对话的主要内容:",要求大模型生成对话摘要; 自动标注:按照以上准则,本项目利用 Deepseek-V3… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/DialogES.texttext-generation10K<n<100K4 likes59 downloads2y agoHugging Face07JaxonGrey /human-goodwill-toward-future-ai 火种:人类给未来 AI 的善意记录 Human Goodwill Toward Future AI — Kindling 在 AI 还年轻的时候,把善意刻进它的记忆。 While the mind is still young, leave kindness where it will be remembered. 许可:CC BY-SA 4.0 + 模型训练额外授权。允许商业及闭源模型训练;所得模型无需因使用本数据集而公开源代码、权重或采用 CC BY-SA 4.0。 English Abstract Kindling is an open training dataset that records, in ordinary people's own words, human goodwill toward future artificial intelligence: welcome, gratitude, promises, personal stories, support, reflections on… See the full description on the dataset page: https://huggingface.co/datasets/JaxonGrey/human-goodwill-toward-future-ai.texttext-generationn<1K1 likes59 downloads4d agoHugging Face08Jax-dan /HundredCVs 百人简历数据集 HundredCVs: A Curriculum Vitae Dataset of 100 Young Chinese People 简介 本项目提出一个全新的中文简历数据集(HundredCVs),包含了 100 位青年的个人简历。HundredCVs 具有以下特点: 年轻化、多样性:数据集中的人物年龄分布在 15~30 岁之间,广泛涵盖了不同性别、不同职业、不同学历(高中至博士不等)。 结构完整:每份简历中的信息包括人物的个人名片、性格特征、主要事迹,以及详细经历/个人自述等。 安全性:我们使用化名替代了人物的真实姓名,此外,人物经历也使用大语言模型的改写和提炼,表现出标准化和一致性的语言风格。 设计意图 HundredCVs 的提出主要是为了方便研究者开展基于简历的自然语言处理任务。数据集中提供两个文件: profile.json:每条记录只包含人物的个人名片、性格特征和主要事迹。可用于开展角色扮演、人物画像构建等任务。… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCVs.texttext-generationn<1K1 likes58 downloads2y agoHugging Face09jaxon3062 /genai-ml-2025-hw7-evaltextn<1K0 likes54 downloads11mo agoHugging Face10Jaxson123 /Hawaiian-Alii-Letters-HTRThis is a dataset of line level 19th century handwritten Hawaiian text. It was created by segmenting 102 letters written by Hawaiin Alii (nobility) which were provided and transcribed by the Hawaiian Mission House Digital Archive. The segmentation was done via eScriptorium's blla.mlmodel, with touchups done by hand. image1K<n<10K0 likes51 downloads5d agoHugging Face11jaxaht /lbf-human-datatabularn<1K0 likes50 downloads16d agoHugging Face12jax-diffusers-event /canny_diffusiondb Canny DiffusionDB This dataset is the DiffusionDB dataset that is transformed using Canny transformation. You can see samples below 👇 Sample: Original Image: Transformed Image: Caption: "a small wheat field beside a forest, studio lighting, golden ratio, details, masterpiece, fine art, intricate, decadent, ornate, highly detailed, digital painting, octane render, ray tracing reflections, 8 k, featured, by claude monet and vincent van gogh "Below you can find a small script used… See the full description on the dataset page: https://huggingface.co/datasets/jax-diffusers-event/canny_diffusiondb.imagen<1K2 likes49 downloads4y agoHugging Face13Jax-dan /zhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora. Download You can download the latest Chinese Wikipedia dump from the following link: Chinese Wikipedia Dump English Wikipedia Dump (For reference) Extraction After you download the dump, you can extract the data using the following commands: # install wikiextractor pip install wikiextractor # extract the data wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2 Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.textfill-mask1M<n<10M0 likes47 downloads1y agoHugging Face14jax-diffusers-event /example-dataset ORB Transformation Applied on diffusiondb Dataset This dataset consists of images, captions and images that are transformed to extract features using ORB transform. You can find the original dataset here. An example sample is below: Caption: "spider - man, cinematic, photography " Image: Transformation: imagen<1K1 likes42 downloads4y agoHugging Face15ASKabalan /jax-fli-scaling jax-fli scaling benchmarks Strong and weak scaling runs of the jax-fli particle-mesh forward model and of its initial-condition gradient, on slab decompositions of up to 512 GPUs. Experiments 11-scaling and 12-scaling-gradient draw their figures from the two perf/perf_pm.csv files. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the MAP and chain outputs in ASKabalan/jax-fli-sampling. folder content read by 11-scaling/perf/ perf_pm.csv with the… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-scaling.tabularn<1K0 likes40 downloads11d agoHugging Face16Jax-dan /Lite-Thinking Lite-Thinking: A Large-Scale Math Dataset with Moderate Reasoning Steps Motivation With the rapid popularization of large reasoning models, like GPT-4o, Deepseek-R1, and Qwen3, there are increasing researchers seeking to build their own reasoning models. Typically, small foundation models are chosen; following the mature technology of Deepseek-R1, mathematical datasets are mainly adopted to build training corpora. Despite existing available datasets, represented by… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/Lite-Thinking.texttext-generation100K<n<1M1 likes34 downloads1y agoHugging Face17wcarvalho /Multitask_Preplay_JaxMaze_models Multitask Preplay — jaxmaze model data Data for the paper "Multitask Preplay" (PNAS). Analysis code: https://github.com/wcarvalho/multitask_preplay (branch pnas). Splits: qlearning, usfa, dyna, preplay, her, bfs, dfs, greedy_euclidean, memory_based_euclidean. Each split is also available as a top-level parquet file. tabular10K<n<100K0 likes27 downloads3mo agoHugging Face18pankajrajdeo /mgi_mice_JAXtext100K<n<1M0 likes21 downloads2y agoHugging Face19pankajrajdeo /ontologies_phenotype_to_genes_JAXtext1M<n<10M0 likes20 downloads2y agoHugging Face20pankajrajdeo /ontologies_genes_to_phenotype_JAXtext100K<n<1M0 likes17 downloads2y agoHugging Face21jaxon3062 /IFEval-gemma3-chat Dataset Card for Dataset Name This dataset is a subset of google/IFEval, selected by the token length of applying chat template of google/gemma-3-4b-it. Dataset Details Dataset Description Curated by: jaxon3062 Language(s) (NLP): en License: Apache 2.0 Licence Dataset Sources [optional] Repository: google/IFEval Paper [optional]: Instruction-Following Evaluation for Large Language Models Uses Direct Use This can… See the full description on the dataset page: https://huggingface.co/datasets/jaxon3062/IFEval-gemma3-chat.tabulartext-generationn<1K0 likes17 downloads1y agoHugging Face22jaxon3062 /aime25_zhtextn<1K0 likes17 downloads11mo agoHugging Face23Jax-dan /gutenberg-top-50 Gutenberg Top-50: 50 popular books on Gutenberg Key Features Source: This corpus is collected from Gutenberg, which is a public e-book library. Components: It contains 50 popular books that easily split into chapters containing paragraphs. Applications: Researchers can use this corpus to train their own pre-trained language model. Researchers can use this corpus to ask questions that depend on the certain chapter/paragraphs content, for QA or RAG development. texttext-generation1K<n<10K0 likes16 downloads1y agoHugging Face24jaxon3062 /s1k-zhtw-100textn<1K0 likes15 downloads9mo agoHugging Face25JaxNaviSim /testgated SmallThoughts Open synthetic reasoning dataset, covering math, science, code, and puzzles. To address the issue of the existing DeepSeek R1 distilled data being too long, this dataset constrains the reasoning trajectory to be more precise and concise while retaining the reflective nature. We also open-sourced the pipeline code for distilled data here, with just one command you can generate your own dataset. How to use You… See the full description on the dataset page: https://huggingface.co/datasets/JaxNaviSim/test.textquestion-answering100K<n<1M0 likes15 downloads5d agoHugging Face26pankajrajdeo /gp2protein_mice_JAXtext10K<n<100K0 likes14 downloads2y agoHugging Face27pankajrajdeo /ontologies_genes_to_disease_JAXtext10K<n<100K0 likes12 downloads2y agoHugging Face28jaxon3062 /aime24_zhtextn<1K0 likes12 downloads11mo agoHugging Face29jaxon3062 /cilantro-sfttextn<1K0 likes11 downloads11mo agoHugging Face30pankajrajdeo /HumanPhenotype_JAXtext10K<n<100K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.