Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01group2sealion /uet_iai_nlp_data_for_llmsData sources come from the following categories: 1.Web crawler dataset: Website UET (ĐH Công nghệ): tuyensinh.uet.vnu.edu.vn; new.uet.vnu.edu.vn Website HUS (ĐH KHTN): hus.vnu.edu.vn Website EUB (ĐH Kinh tế): ueb.vnu.edu.vn Website IS (ĐH Quốc tế): is.vnu.edu.vn Website Eduacation (ĐH Giáo dục): education.vnu.edu.vn Website NXB ĐHQG: press.vnu.edu.vnList domain web crawler CC100:link to CC100 vi Vietnews: link to bk vietnews dataset C4_vi: link to C4_vi Folder Toxic store files demo… See the full description on the dataset page: https://huggingface.co/datasets/group2sealion/uet_iai_nlp_data_for_llms.0 likes9.5k downloads2y agoHugging Face02NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Package version recorded in run metadata: 0.2.0 (some runs lack version metadata). Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes3.2k downloads12d agoHugging Face03elmoghany /Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text Dataset Overview A collection of 27 domains (“topics”) and 3100 question-answer pair. Each topic comes with average 117 QA pairs.Every QA entry comes with: references: one or more source files the answer is extracted from time with each reference comes the starting and ending time the answer is extracted from the reference video_files: the video files where the answer can be found (future) video title & description from metadata.csv File structure You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.question-answering1K<n<10K3 likes2.1k downloads1y agoHugging Face04nnheui /llm-srbenchgated LLM-SRBench: Benchmark for Scientific Equation Discovery with LLMs We introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorization… See the full description on the dataset page: https://huggingface.co/datasets/nnheui/llm-srbench.textn<1K15 likes1.5k downloads1y agoHugging Face05FPEvalRepoPublic /LLMsGeneratedCode0 likes1.1k downloads5mo agoHugging Face06Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 64 runs, 2,782,350 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes977 downloads13d agoHugging Face07zkolter /llm_speedrun LLM Speedrun token streams Pre-tokenized training artifacts for the LLM speedrun exercises. File Description Tokens tokenizer_50M.bpe JSON-serialized BPE tokenizer — fineweb-edu-10BT.shuffle.bin Shuffled FineWeb-Edu sample/10BT token stream 9,440,023,113 smoltalk.shuffle.bin Shuffled SmolTalk data/all token stream 875,269,408 The .bin files are headerless, little-endian unsigned 16-bit token IDs and can be memory-mapped with NumPy: from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/zkolter/llm_speedrun.text-generation1 likes975 downloads19d agoHugging Face08bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes925 downloads2y agoHugging Face09subbuvincent /dec1-jme-student-sourcing-training-and-llms0 likes749 downloads7mo agoHugging Face10llm-jp /scaling-data-constrained-llms Scaling Data-Constrained Language Models with Synthetic Data This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026). Overview This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting. Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.texttext-generation100M<n<1B5 likes737 downloads7mo agoHugging Face11pkuHaowei /llm-srbench LLM-SRBench: Benchmark for Scientific Equation Discovery with LLMs This dataset contains LLM-SRBench, a comprehensive benchmark for evaluating Large Language Models (LLMs) on scientific equation discovery (symbolic regression) tasks. Paper: LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models (ICML 2025 Oral) Original Repository: deep-symbolic-mathematics/llm-srbench Original Dataset: nnheui/llm-srbench 📊 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/pkuHaowei/llm-srbench.textn<1K0 likes674 downloads7mo agoHugging Face12Multi-Agent-LLMs /DEBATE DEBATE: Diverse Multi-Agent Debates This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework". Citation comming soon. tabulartext-generation10K<n<100K2 likes552 downloads1y agoHugging Face13FYQ12138 /llm_sam_audio_data0 likes495 downloads5mo agoHugging Face14subbuvincent /mcae-llms-benchmark-report0 likes480 downloads7mo agoHugging Face15bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes475 downloads2y agoHugging Face16Stereotypes-in-LLMs /hiring-bias-mitigation-synthetic-data Hiring-bias mitigation — synthetic training data Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a protected attribute (military status, gender, religion), in English and Ukrainian. Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4. Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.tabulartext-generation100K<n<1M0 likes358 downloads17d agoHugging Face17megrisdal /llms-txt Context & Motivation https://llmstxt.org/ is a project from Answer.AI which proposes to "standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time." I've noticed many tool providers begin to offer /llms.txt files for their websites and documentation. This includes developer tools and platforms like Perplexity, Anthropic, Hugging Face, Vercel, and others. I've also come across https://directory.llmstxt.cloud/, a directory of… See the full description on the dataset page: https://huggingface.co/datasets/megrisdal/llms-txt.textn<1K13 likes336 downloads5d agoHugging Face18ufca-llms /juris-tcuJurisTCU is a Brazilian Portuguese legal IR resource built from the curated jurisprudence collection of the Brazilian Federal Court of Accounts (TCU), from which we derive the benchmark subset used here. In its original form, the dataset contains 16,045 jurisprudence documents organized into more than 20 fields (metadata and textual fields). The most relevant are ENUNCIADO and EXCERTO, which correspond, respectively, to a summary of the ruling and to the excerpt from the decision that supports… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/juris-tcu.texttext-retrieval10K<n<100K2 likes332 downloads6mo agoHugging Face19Wutaghost /LLMscore-ICLR-OpenReview LLMscore-ICLR-OpenReview This dataset is the released original dataset for the paper Position: Peer Review Should Be Calibrated via LLM Scoring by Zijin Chen, Lesui Yu, Xiaofei Liao, Hai Jin, and Qinbin Li. The paper has been accepted to the ICML 2026 Position Track. Its concrete purpose is peer review analysis: the dataset is meant for studying how paper-review rationales, numeric ratings, LLM-derived anchor scores, and review-score residuals interact in scientific peer… See the full description on the dataset page: https://huggingface.co/datasets/Wutaghost/LLMscore-ICLR-OpenReview.text-classification100K<n<1M1 likes332 downloads4mo agoHugging Face20ufca-llms /BR-TaxQAThis dataset corresponds to BR-TaxQA-R, a collection derived from materials of the Brazilian Federal Revenue Service (Receita Federal) on personal income tax (IRPF). It contains 715 questions with their corresponding answers and introduces user-oriented, FAQ-style query formulations in a legal-tax domain. Some questions are explicitly linked to other related questions. In our relevance design, the immediate answer to the queried question is treated as the primary positive with score = 2, while… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/BR-TaxQA.texttext-retrieval1K<n<10K0 likes313 downloads6mo agoHugging Face21ufca-llms /normas-tcuNormasTCU is a Brazilian Portuguese legal IR test collection composed of normative documents from the Brazilian Federal Court of Accounts (TCU). These normative acts may have internal effects (e.g., rules governing internal procedures) or external effects (e.g., rules regulating how the court interacts with other public institutions) and differ from jurisprudential documents in both purpose and structure. Jurisprudential documents typically describe specific cases and present the legal… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/normas-tcu.texttext-retrieval10K<n<100K0 likes306 downloads6mo agoHugging Face22neemiasbsilva /multimodal-LLMs-See-Sentiment MLLMsent — datasets and experiment results Every input and every output of "Multimodal LLMs See Sentiment" (arXiv:2508.16873): the image descriptions generated by six multimodal LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete per-fold results of all 141 experiments. Paper: arXiv:2508.16873 Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.texttext-classification10K<n<100K1 likes305 downloads2mo agoHugging Face23ufca-llms /juaJUÁ-Juris is centered on jurisprudence drawn from the curated jurisprudence collection of the Brazilian Federal Court of Accounts (TCU). In this collection, each instance contains an enunciado and an excerto: the enunciado is an abstractive summary of the ruling, while the excerto is the passage from the ruling that supports that summary. In our retrieval setup, the enunciado serves as the query, and the corresponding excerto is treated as the ground-truth positive passage. Within the… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/jua.texttext-retrieval10K<n<100K0 likes304 downloads6mo agoHugging Face24disco-eth /Reasoning-Boosts-Opinion-Alignment-in-LLMs Reasoning Boosts Opinion Alignment in LLMs Download and load with DataDict.load_from_disk. Dataset fields id: Identifier for the respondent / party / candidate (see lists below) question: Identifier for the question question_text: The actual question answer: The answer to the respondent / party / candidate gave to the question. A = Yes, B = No, C = Neutral. answer_comment: The argument / comment used for SFT. political_position: The ideological group / party for this… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/Reasoning-Boosts-Opinion-Alignment-in-LLMs.0 likes293 downloads1y agoHugging Face25commoncrawl /llms.txt llms.txt files extracted from the Common Crawl corpus This dataset contains llms.txt files extracted from the Common Crawl corpus. Specifically, we extracted response records matching */llms.txt and */llms-full.txt with status 200 and MIME-type text/plain or text/markdown. What is llms.txt? The /llms.txt file: A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time. Our own analysis of this dataset is available… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/llms.txt.text100K<n<1M2 likes290 downloads1mo agoHugging Face26cutycat2000x /Global-LLMs-Replies Global LLMs Replies GPT-4o -> 74,644 rows mixtral-8x22b -> 13,129 rows claude-3-haiku -> 3,871 rows textquestion-answering10K<n<100K3 likes284 downloads2y agoHugging Face27stacklok /llm-security-leaderboard-contentstabularn<1K0 likes278 downloads1y agoHugging Face28sqres /llmsplit_deepseektext1M<n<10M0 likes272 downloads1y agoHugging Face29hoangphu7122002ai /llm-serving-bench-runs llm-serving-bench runs Heavy outputs of the GPU sessions of https://github.com/hoangphu7122002/llm-serving-bench (answers, canary, server logs, per-level parquets, quality-pack raw files). Paths mirror results/ in the repo; the git repo keeps the reports and an HF.txt pointer. Pull one session: uv run python scripts/data_hub.py pull-run --session runs/<name> --dest results. Sessions runs/20261009-1254_session4_quant-cache · 2026-10-09 · 15a370143e68… See the full description on the dataset page: https://huggingface.co/datasets/hoangphu7122002ai/llm-serving-bench-runs.0 likes267 downloads6h agoHugging Face30hsiung /llm-similarity-risktext1M<n<10M0 likes257 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.