Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes24k downloads3h agoHugging Face02Gramscii-IT /european-open-data-catalogue Open Data catalogue This repository publishes independently versioned metadata and licensed source snapshots: A discovery catalogue with 69000 dataset entries, covering the providers listed in the discovery table below. 3 independently pinned availability indexes with 911,795 joint combinations across 35 datasets, built from complete source responses within the explicitly declared scope. Licensed Cruscotto source snapshots, stored separately from the metadata, preserve the… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.text100K<n<1M1 likes5.8k downloads14h agoHugging Face03opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes4.4k downloads1y agoHugging Face04Logics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes2.1k downloads9mo agoHugging Face05SkillCorner /opendata-bodypose SkillCorner Open Data — Body Pose 3D body-pose data derived from broadcast video, released alongside the SkillCorner Open Data repository as a joint initiative between SkillCorner and PySport. Initial testing release. Two matches, published so the community can work with the format and tell us what is useful before we consider a wider release. Feedback is genuinely wanted — open an issue on the opendata repo or reply in the Community tab here. What is in here… See the full description on the dataset page: https://huggingface.co/datasets/SkillCorner/opendata-bodypose.tabular100K<n<1M0 likes521 downloads1mo agoHugging Face06mlfoundations /open_lm_test_datatextn<1K0 likes417 downloads3y agoHugging Face07opendatalab /CiteVQA CiteVQA English | 简体中文 CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs. The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.textvisual-question-answering1K<n<10K11 likes357 downloads5mo agoHugging Face08Logics-MLLM /Logics-STEM-SFT-Dataset-Open-5.3Mtext1M<n<10M4 likes347 downloads9mo agoHugging Face09OpenChristianDataOrg /open-christian-data Open Christian Data Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use. Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.tabulartext-generation100K<n<1M1 likes252 downloads3mo agoHugging Face10opendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes239 downloads1y agoHugging Face11gohumanize /gohumanize-open-humanizer-dataset GoHumanize Open Humanizer Dataset 2,957 training pairs and 300 test pairs for teaching a language model to rewrite AI-styled English prose into natural human writing. Each pair is: input: a passage rewritten by a large language model in the register typical of LLM output (formal, smooth, hedged, connective phrases, no contractions); output: the original human-written passage, from a public-domain book or, since version 2, from a US federal government publication. The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.tabulartext-generation1K<n<10K0 likes217 downloads17d agoHugging Face12opendatalab /K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究 texttext-generationn<1K3 likes127 downloads1y agoHugging Face13open-misconceptions /miscon-data Open Misconceptions A public catalogue of misconceptions with stable IDs. Each row is one record: a belief a learner could hold, its kind, the evidence pattern that reveals it (with a concrete example), discriminators against slips and neighbouring misconceptions, alignments to external schemes, and provenance. This dataset mirrors dist/miscon.jsonl from the tagged release of https://github.com/open-misconceptions/miscon-data. The canonical form of a record is its stable URI… See the full description on the dataset page: https://huggingface.co/datasets/open-misconceptions/miscon-data.textn<1K0 likes109 downloads1mo agoHugging Face14Open-COT-Data /COT-Dataset-Mathtext1K<n<10K0 likes100 downloads2y agoHugging Face15opendatalab /SlimPajama-Meta-rater-Reasoning-30B Top 30B token SlimPajama Subset selected by the Reasoning rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.tabulartext-generation1M<n<10M1 likes90 downloads1y agoHugging Face16Defetya /ru-open-llama-training-datasetstext1M<n<10M1 likes87 downloads3y agoHugging Face17opendatalab /SlimPajama-Meta-rater-Professionalism-30B Top 30B token SlimPajama Subset selected by the Professionalism rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.tabulartext-generation1M<n<10M0 likes84 downloads1y agoHugging Face18open-llm-leaderboard /databricks__dolly-v2-7b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.tabular10K<n<100K0 likes78 downloads2y agoHugging Face19opendatalab /SlimPajama-Meta-rater-Cleanliness-30B Top 30B token SlimPajama Subset selected by the Cleanliness rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.tabulartext-generation1M<n<10M0 likes68 downloads1y agoHugging Face20open-nhe /Elysium-X-150-FR-dataset Technical paper: Elysium X 150 FR: A Small LoRA Adapter for Sparse, Per-Speaker Emotion and Appraisal Labeling over a 150-Coordinate Schema (Zenodo preprint, DOI 10.5281/zenodo.23155240), Pratham Prateek Mohanty. Paper page with in-browser demo: open-nhe/Elysium-X-150-FR-Paper. Elysium X 150 FR Emotion Dataset 4,000 labeled target turns from 1,000 four-turn English dialogues, annotated on the 150-coordinate emotion schema used to train Elysium X 150 FR (OpenNHE / Pratham… See the full description on the dataset page: https://huggingface.co/datasets/open-nhe/Elysium-X-150-FR-dataset.tabulartext-classification1K<n<10K0 likes68 downloads5d agoHugging Face21opendatalab /Meta-rater-PRRC-Rater-dataset PRRC Rater Training and Evaluation Dataset Dataset Description This dataset contains the full training and evaluation data for the PRRC rater models described in Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. It is designed for training and benchmarking models that score text along four key quality dimensions: Professionalism, Readability, Reasoning, and Cleanliness. Source: Subset of SlimPajama-627B, annotated for PRRC dimensions… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Meta-rater-PRRC-Rater-dataset.tabulartext-classification100K<n<1M1 likes67 downloads1y agoHugging Face22serialhex /Open-Biblical-Dataset Open Biblical Dataset 11,105 examples. Synthetic instruction-tuning data for a Christian theology / pastoral-counseling fine-tune: a real situated question, a reasoning trace, and a grounded pastoral answer — every answer required to quote and stay inside the bounds of a real source text, never an isolated verse or an appeal to unnamed authority ("many theologians have held..."). Built with data-forge — a custom generation pipeline (question-gen → answer-gen → offline lint pass)… See the full description on the dataset page: https://huggingface.co/datasets/serialhex/Open-Biblical-Dataset.texttext-generation10K<n<100K0 likes62 downloads8d agoHugging Face23OpenDataMoroccanLaw /morocco-cassation-court-decisions Morocco Cassation Court Decisions 29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0 Why this dataset exists In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.texttext-generation10K<n<100K0 likes60 downloads2mo agoHugging Face24open-llm-leaderboard /databricks__dolly-v2-12b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face25open-llm-leaderboard /godlikehhd__alpaca_data_score_max_0.1_2600-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.1_2600 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.1_2600 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.1_2600-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face26open-nhe /elysium-x-500-fr-dataset 🌌 Elysium X 500 FR Dataset (HumanMonitored v0.5) 12,000 short conversations, each with one "target line" labelled with one or two of 500 emotion coordinates (with a strength of 0.25, 0.5, 0.75 or 1.0). It was made by the OpenNHE Technologies team to train and test the Elysium X 500 FR emotion model. Licence: CC BY-NC-SA 4.0. Non-commercial use only. Please do not call this "open source" in the strict sense. 📦 What is in it Item Value Rows 12… See the full description on the dataset page: https://huggingface.co/datasets/open-nhe/elysium-x-500-fr-dataset.tabulartext-classification10K<n<100K1 likes56 downloads5d agoHugging Face27Lines /Open-Domain-Oral-Disease-QA-Dataset Open-Domain-Oral-Disease-QA-Dataset Dataset Details Dataset Description This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in the domain of oral disease. We currently offer a suite of evaluation datasets encompassing models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B. More data is under reviewed. This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/Lines/Open-Domain-Oral-Disease-QA-Dataset.textn<1K5 likes47 downloads2y agoHugging Face28SuperbEmphasis /Open-ert-small-datasetThis is a subset of: https://huggingface.co/datasets/openerotica/long-roleplay-v0.1 I am using mistral's new DEVSTRAL model to take the entire conversation in JSON format and rate it. I chose DEVSTRAL due to the mistral models being very consistent and well rounded. The Devstral model I was hoping could understand the JSON format a bit better. I ask the mode to rate each RP based on many different factors including grammar, prose, length (And a few others I will keep to myself :D). I then… See the full description on the dataset page: https://huggingface.co/datasets/SuperbEmphasis/Open-ert-small-dataset.textn<1K0 likes43 downloads1y agoHugging Face29open-llm-leaderboard /godlikehhd__alpaca_data_ins_max_5200-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_ins_max_5200 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ins_max_5200 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ins_max_5200-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face30open-llm-leaderboard /databricks__dolly-v2-3b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-3b Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-3b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.