Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /BertaQA Dataset Card for BertaQA BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.tabularquestion-answering10K<n<100K1 likes463 downloads2y agoHugging Face02open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face03open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details.tabular10K<n<100K0 likes25 downloads2y agoHugging Face04open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face05open-llm-leaderboard /Q-bert__MetaMath-1B-detailsgated Dataset Card for Evaluation run of Q-bert/MetaMath-1B Dataset automatically created during the evaluation run of model Q-bert/MetaMath-1B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Q-bert__MetaMath-1B-details.tabular10K<n<100K0 likes20 downloads2y agoHugging Face06toksuitebackup /bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes20 downloads11mo agoHugging Face07Linkhero2 /bertopic-conflictos-chile-v21-masterpiece 🏆 BERTopic v21 - THE MASTERPIECE Mejoras sobre v20: SmartDeduplicator (0.7): Más permisivo, conserva "Argentina", "transfronterizo" Safety Net: Garantiza mínimo 6 keywords por topic Stopwords SOTA: Sin meses, años, nombres personales Visualizaciones: Barcharts y mapas HTML automáticos Métricas: K Natural: 50 Total docs: 3,268 Timestamp: 2026-01-13 17:07:41.278922 Archivos: bertopic_topic_explanations_v21.xlsx: Keywords por K… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v21-masterpiece.tabularn<1K0 likes14 downloads9mo agoHugging Face08ssurface /hallucination-bert-spans Hallucination BERT Span Dataset Flat, one-row-per-span dataset intended for span/token-classification (BIO-tagging style) hallucination detection over agent tool-calling traces, derived from the same judging pipeline as the reasoning-distillation set in this collection. File ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join needed. Each row is one hallucinated span: span (verbatim text), type (taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.tabular10K<n<100K0 likes13 downloads2mo agoHugging Face09ICKD /sst5-bert-scaledtabular10K<n<100K0 likes9 downloads2y agoHugging Face10WhiteRavenPlus /McAuleyLabsAmazonReviewTextEmbeddings_BERTtabular100K<n<1M1 likes5 downloads5mo agoHugging Face11Linkhero2 /bertopic-conflictos-chile-v16-refined 🏆 BERTopic v16.1 - The Refined MMR Diversity 0.8 para eliminar keywords redundantes. N-Grams (1,4) para capturar conceptos complejos. K Natural: 51 topics Columnas separadas para fácil lectura. tabularn<1K0 likes3 downloads9mo agoHugging Face12Linkhero2 /bertopic-conflictos-chile-v19-complete 🏆 BERTopic v19 - "The Complete" Correcciones Críticas desde v17 y v18: Problema Causa Solución v19 "hidroeléctrica" perdida reduce_frequent_words=True reduce_frequent_words=FALSE "Luis", "Rafael" como keywords Sin KeyBERTInspired KeyBERTInspired RESTAURADO Meses del año como keywords No en stopwords Meses en stopwords Nombres personales No filtrados Nombres en stopwords "Enel Green Power" fragmentado Sin fusiones Fusiones SELECTIVAS… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v19-complete.tabularn<1K0 likes3 downloads9mo agoHugging Face13ICKD /sst5-berttabular10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.