Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /tatoeba-bitext-mining Tatoeba An MTEB dataset Massive Text Embedding Benchmark 1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus Task category t2t Domains Written Reference https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["Tatoeba"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.texttranslation100K<n<1M9 likes1.4k downloads8mo agoHugging Face02mesolitica /instructions-pair-miningtext100K<n<1M2 likes730 downloads3y agoHugging Face03mteb /bucc-bitext-mining BUCC.v2 An MTEB dataset Massive Text Embedding Benchmark BUCC bitext mining dataset Task category t2t Domains Written Reference https://comparable.limsi.fr/bucc2018/bucc2018-task.html How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["BUCC.v2"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/bucc-bitext-mining.texttranslation10K<n<100K4 likes634 downloads8mo agoHugging Face04SaylorTwift /mteb-bitext-mining-aggregated MTEB BitextMining Aggregated Dataset (Full) This dataset aggregates ALL configs from 10 BitextMining datasets in the MTEB (Massive Text Embedding Benchmark) Multilingual v2 benchmark into a single, unified dataset for comprehensive bitext mining evaluation. Dataset Summary Total Examples: 448,229 sentence pairs Source Datasets (Configs): 10 MTEB BitextMining tasks Total Splits: 332 language pairs/configurations Languages: 300+ unique language codes across all datasets… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/mteb-bitext-mining-aggregated.textsentence-similarity100K<n<1M0 likes452 downloads6mo agoHugging Face05africatic /africa-world-bank-energy-mining-time-series Africa World Bank Energy and Mining Labeled Time Series Data This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries. ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers. Sector Scope Temporal energy indicators for African countries… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-energy-mining-time-series.textn<1K0 likes425 downloads3mo agoHugging Face06electricsheepasia /asia-energy-world-bank-energy-and-mining-indicators India - Energy and Mining Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28 Abstract Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX. The world economy needs ever-increasing amounts of energy to sustain economic growth, raise living standards, and reduce poverty. But today's trends in energy use are not sustainable. As the world's population grows and economies become more industrialized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-energy-world-bank-energy-and-mining-indicators.tabulartabular-regression1K<n<10K0 likes374 downloads5mo agoHugging Face07gittensor-model-hub /sparkproof-miningtext1K<n<10K0 likes370 downloads3mo agoHugging Face08loicmagne /open-subtitles-500-bitext-miningtext100K<n<1M0 likes331 downloads3y agoHugging Face09loicmagne /open-subtitles-bitext-miningtext1M<n<10M1 likes317 downloads2y agoHugging Face10joelniklaus /german_argument_mining Dataset Card for Annotated German Legal Decision Corpus Dataset Summary This dataset consists of 200 randomly chosen judgments. In these judgments a legal expert annotated the components conclusion, definition and subsumption of the German legal writing style Urteilsstil. "Overall 25,075 sentences are annotated. 5% (1,202) of these sentences are marked as conclusion, 21% (5,328) as definition, 53% (13,322) are marked as subsumption and the remaining 21% (6,481) as other.… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/german_argument_mining.texttext-classification10K<n<100K4 likes203 downloads4y agoHugging Face11loicmagne /tatoeba-bitext-miningtext100K<n<1M0 likes191 downloads2y agoHugging Face12APProjects /us-oil-gas-energy-mining-utility-layoffs-warn-act-notices-daily US oil and gas, energy, mining and utility layoffs — the actual WARN Act filings, rebuilt every day Last rebuilt: 2026-09-24. 1,281 layoff and closure notices filed by oil and gas producers and oilfield-service contractors, coal and hard-rock mines, refineries and pipelines, electric and gas utilities, power plants, solar and wind manufacturers and installers, and waste, recycling and environmental-services operators with US state labor departments — 138,884 workers, 634… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-oil-gas-energy-mining-utility-layoffs-warn-act-notices-daily.texttabular-classification1K<n<10K0 likes169 downloads15d agoHugging Face13lbrenap1 /mining-legal-arguments-us-corporate-case-law Mining Legal Arguments in U.S. Corporate Case Law This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.tabulartext-classification10K<n<100K0 likes157 downloads1mo agoHugging Face14Sam2021 /Arguement_Mining_CL2017tokens along with chunk id. IOB1 format Begining of arguement denoted by B-ARG,inside arguement denoted by I-ARG, other chunks are O Orginial train,test split as used by the paper is providedtextn<1K1 likes156 downloads5y agoHugging Face15loicmagne /open-subtitles-256s-bitext-miningtext100K<n<1M0 likes141 downloads2y agoHugging Face16juzharii /text-mining-ce-dataset-v3tabular100K<n<1M0 likes114 downloads2mo agoHugging Face17BAAI /IndustryCorpus2_mining IndustryCorpus2: Mining This repository contains the IndustryCorpus2: Mining domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year = {2024}, publisher… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_mining.tabular1M<n<10M2 likes108 downloads2mo agoHugging Face18sentence-transformers /quora-duplicates-mining Dataset Card for Quora Duplicate Questions This dataset contains the Quora Question Pairs dataset in a format that is easily used with the ParaphraseMiningEvaluator evaluator in Sentence Transformers. The data was originally created by Quora for this Kaggle Competition. Usage from datasets import load_dataset from sentence_transformers.SentenceTransformer import SentenceTransformer from sentence_transformers.evaluation import ParaphraseMiningEvaluator # Load the… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/quora-duplicates-mining.textfeature-extraction100K<n<1M0 likes98 downloads2y agoHugging Face19jiadali2026 /SGLang-vLLM-PR-Mining Engineering Signals of Human-AI Collaboration Dataset Dataset accompanying the paper: Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Development Authors: Jiada Li, Xuesong Ye, Olamide Olowoniyi Paper: arXiv:2608.13884 Citation @article{li2026engineering, title={Engineering Signals of Human-AI… See the full description on the dataset page: https://huggingface.co/datasets/jiadali2026/SGLang-vLLM-PR-Mining.tabularn<1K1 likes98 downloads10d agoHugging Face20Jarrodbarnes /latent-mining Latent Mining Latent Mining is a benchmark-construction method for scientific-agent tasks where useful public evidence diverges from a withheld verifier-backed outcome. The resulting tasks test whether agents can make calibrated scientific triage decisions under incomplete information. This dataset contains the first public biology subset: 165 cross-locus regulatory-edit triage tasks. Each task asks an agent to choose among candidate noncoding edits for a specified assay and… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/latent-mining.textquestion-answeringn<1K0 likes87 downloads4mo agoHugging Face21electricsheepeurope /europe-worldbank-energy-mining Energy & Mining — Europe (World Bank WDI) 🇪🇺 55,013 observations · 44 Europe countries · 1962–2025 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 55,013 observations of Energy & Mining data across 44 Europe countries, spanning 1962–2025, covering 41 distinct indicators. About the source The World Bank's World Development Indicators (WDI) is the world's most-cited reference for global development data. It compiles officially-recognized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-worldbank-energy-mining.tabulartabular-classification10K<n<100K0 likes72 downloads5mo agoHugging Face22Aasishshr7 /africa-synth-mining-safety-incidents-all African Mining Safety Incidents Dataset | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/Aasishshr7/africa-synth-mining-safety-incidents-all.tabulartabular-classification1K<n<10K0 likes71 downloads1mo agoHugging Face23juzharii /text-mining-ce-dataset Vietnamese Legal Cross-Encoder Dataset Training data for a cross-encoder reranker on Vietnamese legal documents. Source Built from YuITC/Vietnamese-Legal-Documents. Schema Column Type Description qid int64 Query ID cid int64 Document (context) ID query string Legal question document string Candidate document label int64 1 = positive, 0 = negative split string train or test negative_type string random, same_topic_wrong_article… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/text-mining-ce-dataset.tabulartext-classification100K<n<1M0 likes68 downloads3mo agoHugging Face24electricsheepafrica /africa-world-bank-energy-and-mining-indicators-for-nigeria Nigeria - Energy and Mining | Africa (original) Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-energy-and-mining-indicators-for-nigeria.tabulartabular-classification1K<n<10K0 likes64 downloads2mo agoHugging Face25electricsheepafrica /africa-synth-mining-equipment-failure-all African Mining Equipment Failure Dataset | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mining-equipment-failure-all.tabulartabular-classification1K<n<10K0 likes60 downloads2mo agoHugging Face26DCAgent2 /terminal_bench_2_a1_pr_mining_20260325_012831textn<1K0 likes56 downloads7mo agoHugging Face27NoraResearchLab /global-mining-areas Quick Links Website GitHub Hugging Face LinkedIn X Global Mining Areas, Modernized 21,060 mining polygons covering 57,278 km² worldwide, representing the physical land footprint of mining activity — open-pit/open-cut areas, tailings dams, waste-rock dumps, water ponds, processing infrastructure, and other directly associated mining land use. Modernized from Maus et al. (2020)'s PANGAEA release into NORA's standardized schema. No new… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/global-mining-areas.geospatialother10K<n<100K0 likes55 downloads1mo agoHugging Face28opentensor /openvalidators-mining DEPRECATION NOTICE As of August 1, 2023, the OpenValidators Mining dataset has been officially deprecated and discontinued. We are no longer updating or maintaining this dataset. If this data has any relevance to your current or future projects, or you have any questions or concerns related to the deprecation, please do not hesitate to contact us. We are more than willing to assist you and provide guidance for alternative solutions where possible. Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/opentensor/openvalidators-mining.text1M<n<10M5 likes52 downloads3y agoHugging Face29laion /terminal_bench_2_a1_pr_mining_20260805_114333text1K<n<10K0 likes51 downloads2mo agoHugging Face30Lyntas /mining_domain_specific_terminologytextn<1K2 likes47 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.