datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bitcoin-mining-pool-templates
Bitcoin mining pool templates
Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag.
Contents
Table
Record
bitcoin_mining_pool_jobs
A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag
Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.tatoeba-bitext-mining
Tatoeba
An MTEB dataset
Massive Text Embedding Benchmark
1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus
Task category
t2t
Domains
Written
Reference
https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Tatoeba"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.open-subtitles-bitext-miningopen-subtitles-256s-bitext-mininginstructions-pair-miningmteb-bitext-mining-aggregated
MTEB BitextMining Aggregated Dataset (Full)
This dataset aggregates ALL configs from 10 BitextMining datasets in the MTEB (Massive Text Embedding Benchmark) Multilingual v2 benchmark into a single, unified dataset for comprehensive bitext mining evaluation.
Dataset Summary
Total Examples: 448,229 sentence pairs
Source Datasets (Configs): 10 MTEB BitextMining tasks
Total Splits: 332 language pairs/configurations
Languages: 300+ unique language codes across all datasets… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/mteb-bitext-mining-aggregated.bucc-bitext-mining
BUCC.v2
An MTEB dataset
Massive Text Embedding Benchmark
BUCC bitext mining dataset
Task category
t2t
Domains
Written
Reference
https://comparable.limsi.fr/bucc2018/bucc2018-task.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["BUCC.v2"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/bucc-bitext-mining.asia-energy-world-bank-energy-and-mining-indicators
India - Energy and Mining
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
The world economy needs ever-increasing amounts of energy to sustain economic growth, raise living standards, and reduce poverty. But today's trends in energy use are not sustainable. As the world's population grows and economies become more industrialized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-energy-world-bank-energy-and-mining-indicators.africa-world-bank-energy-mining-time-series
Africa World Bank Energy and Mining Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal energy indicators for African countries… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-energy-mining-time-series.sparkproof-miningtatoeba-bitext-miningopen-subtitles-500-bitext-miningMining-Analysis
Mining-Analysis Dataset
A research dataset for error analysis of satellite land-use classifiers, with model predictions, ground truth, RGB imagery, raw spectral bands, and per-month cloud masks. Used in a study of how cloud occlusion affects model predictions of West African land cover (mining, oil palm, rubber).
Background
We use satellite images to detect land-use changes in West Africa — specifically mining sites, oil palm plantations, and rubber plantations.… See the full description on the dataset page: https://huggingface.co/datasets/mqraitem/Mining-Analysis.german_argument_mining
Dataset Card for Annotated German Legal Decision Corpus
Dataset Summary
This dataset consists of 200 randomly chosen judgments. In these judgments a legal expert annotated the components
conclusion, definition and subsumption of the German legal writing style Urteilsstil.
"Overall 25,075 sentences are annotated. 5% (1,202) of these sentences are marked as conclusion, 21% (5,328) as
definition, 53% (13,322) are marked as subsumption and the remaining 21% (6,481) as other.… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/german_argument_mining.Arguement_Mining_CL2017tokens along with chunk id. IOB1 format Begining of arguement denoted by B-ARG,inside arguement
denoted by I-ARG, other chunks are O
Orginial train,test split as used by the paper is providedus-oil-gas-energy-mining-utility-layoffs-warn-act-notices-daily
US oil and gas, energy, mining and utility layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-24. 1,281 layoff and closure notices filed by
oil and gas producers and oilfield-service contractors, coal and hard-rock mines, refineries and pipelines, electric and gas utilities, power plants, solar and wind manufacturers and installers, and waste, recycling and environmental-services operators with US state labor departments — 138,884 workers,
634… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-oil-gas-energy-mining-utility-layoffs-warn-act-notices-daily.mining-legal-arguments-us-corporate-case-law
Mining Legal Arguments in U.S. Corporate Case Law
This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.IndustryCorpus2_mining
IndustryCorpus2: Mining
This repository contains the IndustryCorpus2: Mining domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year = {2024},
publisher… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_mining.quora-duplicates-mining
Dataset Card for Quora Duplicate Questions
This dataset contains the Quora Question Pairs dataset in a format that is easily used with the ParaphraseMiningEvaluator evaluator in Sentence Transformers. The data was originally created by Quora for this Kaggle Competition.
Usage
from datasets import load_dataset
from sentence_transformers.SentenceTransformer import SentenceTransformer
from sentence_transformers.evaluation import ParaphraseMiningEvaluator
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/quora-duplicates-mining.cwv-strategy-mining-dataset
CWV Strategy Mining Dataset
Real, merged GitHub PRs mined for measured Core Web Vitals (CWV)
techniques, plus every intermediate pipeline artifact through final
generated playbook candidates. Produced by
cwv-playbook-miner,
scanning the full public GH Archive event stream. Every decision in the
pipeline runs on real PR text (title, body, diff, comments, reviews) — never
a short phrase, never a bare similarity threshold. See the repo's
SESSION_NOTES.md for the full chronological… See the full description on the dataset page: https://huggingface.co/datasets/Ayush-Singh/cwv-strategy-mining-dataset.data-mining-cweurope-worldbank-energy-mining
Energy & Mining — Europe (World Bank WDI)
🇪🇺 55,013 observations · 44 Europe countries · 1962–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 55,013 observations of Energy & Mining data across 44 Europe countries, spanning 1962–2025, covering 41 distinct indicators.
About the source
The World Bank's World Development Indicators (WDI) is the world's most-cited reference for global development data. It compiles officially-recognized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-worldbank-energy-mining.global-mining-areas
Quick Links
Website
GitHub
Hugging Face
LinkedIn
X
Global Mining Areas, Modernized
21,060 mining polygons covering 57,278 km² worldwide, representing
the physical land footprint of mining activity — open-pit/open-cut areas,
tailings dams, waste-rock dumps, water ponds, processing infrastructure, and
other directly associated mining land use. Modernized from Maus et al.
(2020)'s PANGAEA release into NORA's standardized schema. No new… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/global-mining-areas.bitcoin-phase-mining-breakthrough
Bitcoin Phase-Space Mining Breakthrough
🔥 First Empirical Proof: Quantum Phase Geometry Optimizes Bitcoin Mining
This dataset contains the complete results of a groundbreaking experiment applying Theory of Everything (ToE) phase geometry to Bitcoin mining optimization.
Key Discovery
Harmonic Metallic Oscillation (combining metallic ratios δ₂, δ₄, δ₆) finds hashes 1,078× better than random search over 1 billion attempts.
Results Summary… See the full description on the dataset page: https://huggingface.co/datasets/beanapologist/bitcoin-phase-mining-breakthrough.latent-mining
Latent Mining
Latent Mining is a benchmark-construction method for scientific-agent tasks where useful public evidence diverges from a withheld verifier-backed outcome. The resulting tasks test whether agents can make calibrated scientific triage decisions under incomplete information.
This dataset contains the first public biology subset: 165 cross-locus regulatory-edit triage tasks. Each task asks an agent to choose among candidate noncoding edits for a specified assay and… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/latent-mining.africa-synth-mining-safety-incidents-all
African Mining Safety Incidents Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/Aasishshr7/africa-synth-mining-safety-incidents-all.MiningEmissions
COINjecture Mining Emissions Dataset
Overview
This dataset contains comprehensive mining and emission data from all nodes in the COINjecture Network B. This is a non-synthetic, empirical dataset extracted directly from live production nodes.
Dataset Structure
The dataset contains JSONL files with mining emission records. Each record includes:
Core Mining Data
Block Height: Block number in the chain
Block Hash: Cryptographic hash of the block… See the full description on the dataset page: https://huggingface.co/datasets/COINjecture/MiningEmissions.square-01a-ours-mining-shape-beta05-r3-diagnostic-rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 18378,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-01a-ours-mining-shape-beta05-r3-diagnostic-rollouts.sentinel-lfm-mining-patches
sentinel-lfm — illegal-mining single-frame patches
128px RGB patches cropped from the Roboflow illegal-mining dataset, labelled
mine (1) / no-mine (0). Split by source image (no leakage) into
train/val/test. Provided as PNGs + vlm_sft-format JSONL (one image + prompt
-> JSON answer) so it drops straight into VLM fine-tuning.
split
pos
neg
total
train
1410
555
1965
val
303
66
369
test
303
116
419
RGB only (no multispectral). Each JSONL row is a single-turn VLM… See the full description on the dataset page: https://huggingface.co/datasets/ASTRALK/sentinel-lfm-mining-patches.SGLang-vLLM-PR-Mining
Engineering Signals of Human-AI Collaboration Dataset
Dataset accompanying the paper:
Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Development
Authors: Jiada Li, Xuesong Ye, Olamide Olowoniyi
Paper: arXiv:2608.13884
Citation
@article{li2026engineering,
title={Engineering Signals of Human-AI… See the full description on the dataset page: https://huggingface.co/datasets/jiadali2026/SGLang-vLLM-PR-Mining.
