datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb2-ro-bert
FineWeb2-Ro-BERT
FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here.
Key Features
Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders.
Usage
You… See the full description on the dataset page: https://huggingface.co/datasets/surogate/fineweb2-ro-bert.BertaQA
Dataset Card for BertaQA
BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.fineweb2-ro-bert
FineWeb2-Ro-BERT
FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here.
Key Features
Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders.
Usage
You can load… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/fineweb2-ro-bert.wm_benchmark_test_uploadmagic-bert-dataset
Magic-BERT Binary File Classification Dataset
This dataset contains tokenized binary file samples for training MIME type classification models.
Each sample is a 64KB chunk from the beginning of a file, tokenized using a byte-level BPE tokenizer.
Splits
Split
Samples
Train
37,111
Validation
4,591
Test
4,748
Features
Each sample contains:
Feature
Type
Description
blake2b
string
Content hash (unique sample ID)
mime_type
string
MIME… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/magic-bert-dataset.aio-passages-bpr-bert-base-japanese-v3
Dataset Card for llm-book/aio-passages-bert-base-japanese-v3-bpr
書籍『大規模言語モデル入門』で使用する、「AI王」コンペティションのパッセージデータセットに BPR によるパッセージの埋め込みを適用したデータセットです。
llm-book/aio-passages のデータセットに対して、llm-book/bert-base-japanese-v3-bpr-passage-encoder によるパッセージのバイナリベクトルが embeddings フィールドに追加されています。
Licence
本データセットで利用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。
jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3
Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3"
More Information needed
job_listing_german_cleaned_bert
Dataset Card for "job_listing_german_cleaned_bert"
More Information needed
bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
Elite-Chess-Games
Elite Chess Data
Processed and added version of this dataset
marc
MARC: Metaphor Abstraction and Reasoning Corpus
What This Is
MARC identifies puzzles where figurative language and visual examples are genuinely complementary: the model fails given examples alone, fails given the metaphor alone, but succeeds when both are presented together. We call this the MARC property. The corpus provides 78 MARC-verified puzzles with 1,230 domain-diverse figurative descriptions and complete behavioral trial data for three language models.
Suppose… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc.marc2
MARC2: Metaphor Abstraction and Reasoning Corpus v2
MARC2 extends the MARC-from-LARC methodology to the ARC-AGI2 dataset. It provides a corpus of figurative language puzzles where metaphorical descriptions help AI models solve abstract reasoning tasks they cannot solve from examples alone.
The MARC Property
A task has the MARC property (for a given model) when:
Examples alone fail — the model cannot solve the task from input/output examples
Figurative description… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc2.bertopic-conflictos-chile-v22-complete
🏆 BERTopic v22 - THE COMPLETE MASTERPIECE
Mejoras sobre v21:
MODELO GUARDADO: SafeTensors para reutilización
TODAS LAS VISUALIZACIONES: DataMapPlot, Hierarchy, Heatmap, etc.
MÉTRICAS MATEMÁTICAS: Coherence (UMass, NPMI), Diversity, Silhouette
TOPICS OVER TIME: Análisis temporal
HIERARCHICAL TOPICS: Árbol jerárquico
GET_DOCUMENT_INFO: Metadata completa
REDUCE_OUTLIERS: Reducción inteligente
Métricas Matemáticas:
Topic Diversity: 0.6995305164319249… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v22-complete.test_split_with_embeddings_bert_base_portuguese
Dataset Card for "test_split_with_embeddings_bert_base_portuguese"
More Information needed
Bertmarket-positivity-bert-tokenizedvirus_bert_chunk_2kbpvirus_bert_chunk_2kbp_tokenizedft_bert_benchmark1clincal_bert_large_code_mappingtrec-bert-scaledyelp-bert-scaledft_bert_benchmarksp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173)
Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details.Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_p-beta10-gamma0.3-lr1.0e-6-scale-log-details.SSC-BERT-above-thresholdagnews-bert-scaledJimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.punctuation-mec-bert
Dataset Card for "mec-punctuation-v2"
More Information needed
