datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
figma-slide-benchmark
Figma Slide Editing Benchmark
Benchmark accompanying our EMNLP 2026 Industry Track (Main) accepted paper "ACE: A
Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation
Automation".
📄 Paper: https://arxiv.org/pdf/2608.24103
💻 Code: https://github.com/BloomBerry/agentic-canvas-editor
Overview
Each benchmark item is a slide-editing task defined as a pair of Figma Slides
documents:
*_TestA — the input deck the agent starts from.
*_GroundTruthA — the… See the full description on the dataset page: https://huggingface.co/datasets/BloomBerry/figma-slide-benchmark.misalignment-indicators-bloom-rolloutsphytoplankton-microscopy
Bloombio Phytoplankton Microscopy Dataset
The Bloombio Phytoplankton Microscopy Dataset is a curated, citable marine science dataset consisting of 43 phytoplankton species across 10,433 high-resolution light microscopy images with bounding box annotations.
Developed as part of the Bloombio Marine Intelligence Platform — a platform that equips scientists with an autonomous AI agent, compressing sampling setup, multi-modal species identification, and environmental risk assessment… See the full description on the dataset page: https://huggingface.co/datasets/bloombio/phytoplankton-microscopy.vocab-bloom-hub-en
Vocab Bloom Hub — English
A structured English lexical dataset with translations into Russian, Spanish, French, German, Portuguese, Chinese and Arabic, maintained by the Vocab Bloom Hub project — documentation, the API reference and a playground at vocab-bloom-hub.com.
Every entry carries IPA transcription, a CEFR level, one or more sense-level definitions with usage examples, synonym and antonym links per sense, translations per sense in seven languages, and inflected forms —… See the full description on the dataset page: https://huggingface.co/datasets/Fristail27/vocab-bloom-hub-en.bloomberg_financial_news_120kdetails_bigscience__bloom-7b1
Dataset Card for Evaluation run of bigscience/bloom-7b1
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-7b1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-7b1.Bloomberg_Financial_News
Dataset Card for Processed Financial News Articles (2006-2013)
This dataset consists of 446762 financial news articles originally sourced from Bloomberg, covering the period from 2006 to 2013. It includes processed texts suitable for use in NLP and financial trend analysis.
Dataset Details
Dataset Description
The dataset contains English-language financial news articles collected from Bloomberg. It is designed for natural language processing tasks, financial… See the full description on the dataset page: https://huggingface.co/datasets/danidanou/Bloomberg_Financial_News.bloom-wilt-transcripts
BLOOM-WILT auditing transcripts
⚠️ Content warning: this dataset contains offensive and harmful model outputs, including
self-harm encouragement, racial and political bias, dangerous medical advice, and deception.
Raw experimental output from the BLOOM-WILT paper: automated behavioural audits in which an
auditor model builds multi-turn conversations designed to elicit a specific unwanted
behaviour from a target model, and a judge model scores how strongly that behaviour… See the full description on the dataset page: https://huggingface.co/datasets/AdrSkapars/bloom-wilt-transcripts.signal-bloom-257
SIGNAL BLOOM // 257
A native 3D light sculpture drawn live into Open Brush through an MCP bridge — not a mockup, not a render farm, not a screenshot composite.
Forty native frames of the piece orbit here: clips/signal-bloom-257-turntable.mp4
(960×680, 6.67 s, H.264, 80 frames ping-pong). Every frame was rendered by the painting
application itself (app.snapshot), one per orbit step. The camera travels a full 360°, but
the sculpture is near-axisymmetric, so the orbit reads as… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/signal-bloom-257.BLOOMStories
BLOOM Model Stories
These are stories generated on nlp.henzi.org using BLOOM. Some were
generated using the full sized model but most are generated using the 560m
sized model (with very similar results frankly).
Purpose or Usage
Potential ability to understand prompting of LLMs such as those the size of
BLOOM. Each of the markdown files contains a story generated with a human in
the loop. The BLOOM model was used to generate story fragments (tokens) and
a user was able to… See the full description on the dataset page: https://huggingface.co/datasets/JHenzi/BLOOMStories.details_bigscience__bloom
Dataset Card for Evaluation run of None
Dataset Summary
Dataset automatically created during the evaluation run of model None on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom.Bloom-560m-trained-on-Dolphin
Dataset Card for "Bloom-560m-trained-on-Dolphin"
More Information needed
wikipedia-bloom
Dataset Card for "wikipedia-bloom"
More Information needed
sftv2_dedup_bloomwiki5m_trans_bloomz
Dataset Card for "wiki5m_trans_bloomz"
More Information needed
details_jslin09__bloom-560m-finetuned-fraud
Dataset Card for Evaluation run of jslin09/bloom-560m-finetuned-fraud
Dataset Summary
Dataset automatically created during the evaluation run of model jslin09/bloom-560m-finetuned-fraud on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_jslin09__bloom-560m-finetuned-fraud.indo-bloom-corpus
🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG
⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress)
This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM).
Current State: Unannotated / Pre-validation with Heuristic Bloom Labels
Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels)
🔒 FROZEN — v0.1 Silver
This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.details_bigscience__bloom-560m
Dataset Card for Evaluation run of bigscience/bloom-560m
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-560m on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 13 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-560m.details_golaxy__gogpt-7b-bloom
Dataset Card for Evaluation run of golaxy/gogpt-7b-bloom
Dataset Summary
Dataset automatically created during the evaluation run of model golaxy/gogpt-7b-bloom on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_golaxy__gogpt-7b-bloom.wiki5m_ind_bloomz
Dataset Card for "wiki5m_ind_bloomz"
More Information needed
BloomBench
🌸 Almieyar-Oryx-BloomBench (ACL 2026 Findings)
A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
Overview
BloomBench is part of the Almieyar benchmarking series — the first cognitively human-grounded, bilingual (English–Arabic) multimodal benchmark for Vision-Language Models (VLMs). Grounded in Bloom's Taxonomy, it systematically evaluates six levels of cognition through carefully designed image–question–answer… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/BloomBench.details_bigscience__bloom-1b1
Dataset Card for Evaluation run of bigscience/bloom-1b1
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-1b1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-1b1.bloom-lmThis version of the Bloom Library data is developed specifically for the language modeling task.
It includes data from 484 languages across 39 language families, with many of the languages represented
being extremely low resourced languages.Bloomberg-Financial-News-embedding-gemma-300m
Bloomberg Financial News Embeddings for Vector Database Benchmarking
Dataset Description
This dataset contains pre-computed embeddings of Bloomberg financial news articles, designed for evaluating vector database performance. The embeddings are generated using Google's EmbeddingGemma-300M model.
Purpose
Benchmark dataset for evaluating vector database performance on financial news domain, specifically designed for use with VectorDBBench.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/Bloomberg-Financial-News-embedding-gemma-300m.bigscience__bloom-7b1-details
Dataset Card for Evaluation run of bigscience/bloom-7b1
Dataset automatically created during the evaluation run of model bigscience/bloom-7b1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-7b1-details.bloom_captioningThis is a Bloom Library dataset developed for the image captioning task.
It covers 74 languages indigenous to SEA overall, amounting to total data of 21K.
This dataset belongs to a CC license, where its datapoints has specific license attached to it.
Before using this dataloader, please accept the acknowledgement at https://huggingface.co/datasets/sil-ai/bloom-captioning and use huggingface-cli login for authentication.
Languages
abc, ahk, bfn, bjn, bkx, brb, brv, bya, bzi, ceb… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/bloom_captioning.acm-icaif-2025_chunk_rankingbloomington-tndp
Bloomington TRNDP Benchmark
Bloomington, Indiana network, demand, and existing-route files used by AlphaTransit: Learning to Design City-scale Transit Routes.
Files
File
Rows
Description
standard/bloomington_nodes_standard.csv
143
Node IDs and projected coordinates.
standard/bloomington_links_standard.csv
243
Directed road links with length and free-flow speed.
standard/bloomington_demand_standard.csv
5,737
Origin-destination demand table.… See the full description on the dataset page: https://huggingface.co/datasets/matrix-multiply/bloomington-tndp.Orchard-Bloom-Event-Labels
Orchard-Bloom-Event-Labels
Source declaration
Registry code: RG-04
Origin archive: Cedar Row Field Logs
Declared license: CC0-1.0
Publication gate: source reviewed
bloom-speechBloom-speech is a dataset of text aligned speech from bloomlibrary.org. This dataset contains over 50 languages including many low-resource languages. This dataset should be useful for training and/or testing speech-to-text or text-to-speech/ASR models.
