datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Uzaib52/indian-tech-career-intelligence-2026.automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.kanchana1990_global-clinical-trial-intelligence-20242026
Global Clinical Trial Intelligence 2024–2026
Mirror of the Kaggle dataset kanchana1990/global-clinical-trial-intelligence-20242026 by Kanchana1990, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data.
5,000 registry-verified trials. NLP-ready. ML-ready
Original description (from Kaggle)
Dataset Overview
Clinical trial data sits at the intersection of regulatory… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/kanchana1990_global-clinical-trial-intelligence-20242026.Artificial-intelligence-dataset-for-IR-systems
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
information-retrieval
semantic-search
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.specialty-clinic-synthetic-claims
Specialty Clinic Synthetic Claims Sample
This sample contains 1000 generated-from-scratch synthetic healthcare claim rows.
Use case: Model specialty clinic procedure authorization, documentation, and payer coverage friction across multi-specialty outpatient settings.
Synthetic-Only Boundary
No PHI.
No customer data.
No patient-level source records.
Generated from public healthcare references, explicit modeling assumptions, and deterministic
synthesis code.… See the full description on the dataset page: https://huggingface.co/datasets/Upstream-Intelligence/specialty-clinic-synthetic-claims.ADeLe_battery_v1dot0
Dataset Card for ADeLe
Dataset Summary
ADeLe (Annotated-Demand-Levels) battery is a single, unified test set whose every item is labelled with the level (0-5+) it demands on 18 general ability dimensions (e.g. attention and scan, logical reasoning, various knowledge areas) plus an “unguessability” dimension. It is produced by applying the DeLeAn rubrics, via GPT-4o annotators, to AI benchmarks.
Version 1.0 contains 16 108 items drawn from 63 tasks spread across a diverse… See the full description on the dataset page: https://huggingface.co/datasets/CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0.liquidity-intelligence-benchmarks
VOIDTRACE AI Liquidity Intelligence Benchmarks
Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks.
Built by VOIDTRACE AI.
Dataset Description
This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.us-industrial-facility-intelligence-sample
US Industrial Facility Intelligence — Free Sample
This is a free 100-record sample. It is a subset of the full 1,464-record commercial dataset, provided so you can evaluate the data before deciding whether the full toolkit is useful to you.
An independent, unofficial dataset by NeuroLab Works. Not affiliated with, sponsored by, or endorsed by the U.S. EPA.
What this is
100 real, deduplicated US industrial facilities regulated under EPA's Toxics Release Inventory… See the full description on the dataset page: https://huggingface.co/datasets/NeuroLabWorks/us-industrial-facility-intelligence-sample.algozee_ai-driven-global-market-intelligence-dataset
AI-Driven Global Market Intelligence Dataset
Global Financial Market Data for Risk, Trend, and Investment Analysis
Dataset Info
Source: Kaggle
Original Size: 9.04 MB
Kaggle Downloads: 166
Files: 1
Files
global_market_ai_dataset.csv
Mirrored from Kaggle
career-intelligence-benchmarks
Psychometric Career Intelligence Benchmarks
Benchmark dataset of 20 student psychometric assessment cases with individual scores for personality trait, cognitive ability, interest alignment, motivation clarity, strength discovery, and career readiness.
Built by Psychometric.fyi.
Dataset Description
This dataset contains benchmark data for an educational resource exploring the science of psychometric assessments and career decision making — helping students… See the full description on the dataset page: https://huggingface.co/datasets/psychometric-fyi/career-intelligence-benchmarks.shipwreck-records-sample
Wreck Intelligence: sample of shipwreck records (UK and Ireland), v1.0
Canonical, citable version: Zenodo, DOI 10.5281/zenodo.23187882. This Hugging Face copy is a mirror; please cite the Zenodo DOI.
Wreck Intelligence is a global interactive shipwreck map with detailed, sourced information on every wreck: history, position, and depth.
Map: https://wreckintelligence.com
Contact: hello@wreckintelligence.com
For interest and initial research only. Not a dive plan or dive… See the full description on the dataset page: https://huggingface.co/datasets/wreck-intelligence/shipwreck-records-sample.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Jidnesh298/indian-tech-career-intelligence-2026.review-intelligence-benchmarks
Review Removal Intelligence Benchmarks
Benchmark dataset of 20 review management cases with individual scores for review risk, authenticity, policy compliance, issue detection, platform coverage, and workflow efficiency.
Built by ReviewRemoval.Services.
Dataset Description
This dataset contains benchmark data for an automated review management system helping businesses and reputation management teams monitor reviews, identify potential violations, and manage… See the full description on the dataset page: https://huggingface.co/datasets/review-removal-services/review-intelligence-benchmarks.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/ShaikFayaz042/indian-tech-career-intelligence-2026.serp-intelligence-benchmarks
SERP Intelligence Benchmarks
Benchmark dataset of 20 SERP intelligence cases with individual scores for SERP visibility, search intent, ranking pattern, competitor visibility, SERP feature, and content opportunity signals.
Built by SERPChecker.fyi.
Dataset Description
This dataset contains benchmark data for a research focused SERP intelligence framework helping SEO researchers, content strategists, and digital marketers analyze search results, understand ranking… See the full description on the dataset page: https://huggingface.co/datasets/serpchecker-fyi/serp-intelligence-benchmarks.advanced_cyber_threat_intelligence_telemetryalgozee_business-intelligence-chatbot-interaction-dataset
Business Intelligence Chatbot Interaction Dataset
Real-World Conversational Data for Analytics-Driven Business Insights
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 159
Files: 1
Files
BI_Chatbot_Interactions.csv
Mirrored from Kaggle
Capability_Intelligence
Workforce Capability Intelligence Dataset
Structured workforce capability data for 4.2 million companies. 55M+ records across 30 dimensions covering headcount, growth signals, role composition, and skill concentrations.
What This Is
A company-year dataset mapping what organizations are actually built to do — tracked through workforce skill concentrations, role structures, and growth dynamics. Each company-year observation expands into multiple capability rows, ranked by… See the full description on the dataset page: https://huggingface.co/datasets/Vivameda/Capability_Intelligence.Cyberthreat_intelligence
