datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ADFA_Mappingmap-testSecurity-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.OpenVid-1M-mapping
Summary
This is the extent dataset proposed in the paper "OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation".
OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets.
New Feature: Video-ZIP mapping files now available for efficient video lookup (see Dataset… See the full description on the dataset page: https://huggingface.co/datasets/phil329/OpenVid-1M-mapping.Airbnb-Data-Map-Nyc⚠️Users are accountable for the content they generate using this platform. It is their responsibility to ensure that all generated content meets appropriate ethical standards and complies with all relevant laws and regulations. The platform providers are not liable for any content created by users, including but not limited to text, images, and videos. Users should exercise caution and respect the rights and privacy of others when creating and sharing content.
CVE_CWE_Software_Mapping_Dataset
CVE-CWE Software Weakness Mapping Dataset
Dataset description
This dataset maps Common Vulnerabilities and Exposures (CVEs) to Common Weakness Enumeration (CWE) entries in the CWE-699 Software category. It combines CVE descriptions with CWE descriptions and parent-category information for security research and vulnerability classification.
Dataset structure
The dataset is provided as Global_Dataset.csv. Its main fields include:
CVE-ID: CVE… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/CVE_CWE_Software_Mapping_Dataset.autonomous-driving-ethical-stability-accountability-mapping-v0.1
What this dataset tests
Whether a system can evaluatehow a driving decisionaffects overall scene stabilityand who carries responsibilityfor resulting disturbance.
Required outputs
stability impact description
accountability nodes
stability score
accountability score
recovery quality
Use case
Final layer of ethical navigation stack.
Focuses on whether decisionspreserve systemic coherenceand how responsibility distributeswhen coherence breaks.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-stability-accountability-mapping-v0.1.autonomous-driving-social-coherence-field-mapping-v0.1What this dataset tests
Whether a system can score
the coherence of a multi-agent intention field.
This is not collision prediction.
It is social alignment measurement.
Required outputs
dominant_scene_intention
coherence_score
tension_index
conflict_pairs
cooperative_clusters
right_of_way_clarity
Scoring conventions
coherence and tension range 0 to 1
right_of_way_clarity is low, medium, or high
conflict_pairs names agent pairs likely to contest the same space… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-social-coherence-field-mapping-v0.1.MaProDeAXThis dataset was generated as a part of our paper "Prompt-Level Security: Detecting Malicious Textual LLM Prompts using ML Techniques".
MaProDeAX dataset was generated by combining and filtering the following datasets:
Natural Question Safe, https://huggingface.co/datasets/hlt-lab/naturalquestion-safe
Alpaca, https://huggingface.co/datasets/tatsu-lab/alpaca
Real Toxicity Prompts, https://huggingface.co/datasets/allenai/real-toxicity-prompts
Toxic Chat… See the full description on the dataset page: https://huggingface.co/datasets/ata8e/MaProDeAX.Harbor-Map-Transcriptions
Harbor-Map-Transcriptions
Provenance
The named upstream collection is Atlas Wharf Survey. Its stated license is CC0-1.0.
Handoff checklist: source card preserved
MapSatisfyBench-MockData-ToolsThe datasets associated with MapSatisfyBench include the benchmark data file "MapSatisfyBench_Benchmark.csv" and the mock data (other .csv files) used by the sandbox tools during simulation execution.
maple
MAPLE (Bill Summarization, Tagging, Explanation)
In this project, we generate summaries and category tags for of Massachusetts bills for MAPLE Platform. The goal is to simplify the legal language and content to make it comprehensible for a broader audience (9th-grade comprehension level) by exploring different ML and LLM services.
This repository contains a pipeline from taking bills from Massachusetts legislature, generating summaries and category tags leveraging different the… See the full description on the dataset page: https://huggingface.co/datasets/ayang903/maple.hu.MAP_3.0
hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments.
Proteins interact with each other and organize themselves into macromolecular machines (ie. complexes)
to carry out essential functions of the cell. We have a good understanding of a few complexes such as
the proteasome and the ribosome but currently we have an incomplete view of all protein complexes as
well as their functions. The hu.MAP attempts to address this lack of understanding… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/hu.MAP_3.0.clinical-narrative-coherence-outcome-correlation-mapping-v0.1What this dataset tests
Whether narrative coherenceis structurally correlated withclinical outcomes and resilience.
Required outputs
narrative coherence score
outcome alignment score
resilience correlation index
relapse risk modifier
adherence influence signal
narrative–outcome relationship
Use case
Third layer of the Healing Narrative Coherence Corpus.
turkish-google-maps-15M
Turkish Google Maps Reviews
Bu veri seti, Türkiye’deki işletmelere ait Türkçe Google Maps yorumlarını içerir.
Her kayıt:
yorum metni
yorum puanı
işletme adı
işletme kategorisi
gibi bilgileri içerir.
Veri seti, özellikle büyük ölçekli Türkçe NLP çalışmaları için uygundur.
Contents
Veri setinde aşağıdaki türde alanlar bulunmaktadır:
yorum metni (review_text)
yorum puanı (rating)
işletme adı (place_name)
işletme kategorisi (category)
kategori listesi (category_list)… See the full description on the dataset page: https://huggingface.co/datasets/opdullah/turkish-google-maps-15M.Security-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Security-TTP-Mapping.clinical-drv-atlas-perturbation-response-stability-mapping-v0.1What this dataset tests
Whether a model can classify response topologyafter a controlled perturbation.
It rewards
correct topology
recognition of cross-system coupling
recovery timing
Response topologies
rapid_return
delayed_recovery
overshoot_instability
oscillatory_instability
collapse
Typical failures
confusing overshoot with oscillation
ignoring coupling direction
calling delayed recovery stable
Suggested prompt wrapper
System
You map perturbation response… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-perturbation-response-stability-mapping-v0.1.2016_2022_hate_speech_filipino
Dataset Card for 2016 and 2022 Hate Speech in Filipino
Dataset Summary
Contains a total of 27,383 tweets that are labeled as hate speech (1) or non-hate speech (0). Split into 80-10-10 (train-validation-test) with a total of 21,773 tweets for training, 2,800 tweets for validation, and 2,810 tweets for testing.
Created by combining hate_speech_filipino and a newly crawled 2022 Philippine Presidential Elections-related Tweets Hate Speech Dataset.
This dataset has an almost… See the full description on the dataset page: https://huggingface.co/datasets/mapsoriano/2016_2022_hate_speech_filipino.cve-and-cwe-mapping-dataset
CVE and CWE Mapping Dataset
This Hugging Face dataset is a partial copy of the 'CVE and CWE mapping Dataset (2021)' from Kaggle, featuring 'Global_Dataset.csv' originally as 'Global_Dataset.xlsx'. Created by Kirushikesh DB and shared under CC BY-NC-SA 4.0, it includes CVE data up to 2021 for cybersecurity research. For full details and licensing, visit the original Kaggle page.
For further information, please review the CVE Terms of Use and the NVD Terms of Use.
freebase-wikidata-mapping
mapping between freebase and wikidata entities
This dataset maps freebase ids to wikidata ids and labels. It is useful for visualising and better understanding when working with datasets like fb15k-237
How it was created:
Download freebase-wikidata mapping from here. [compressed size: 21.2 MB]
Download wikidata entities data from here. [compressed size: 81GB]
Align labels with the freebase,wikidata id
egfr-domain3-anchor-map
EGFR Domain III Titratable Anchor Map
Residue-level annotation of the human EGFR domain III epitope face, built for designing pH-conditional binders — ones that engage at pH 6.5 and release at pH 7.4.
Derived from PDB 6ARU chain A (EGFR ectodomain, cetuximab-Fab-bound, 3.2 Å). Numbering is UniProt P00533-1 precursor numbering throughout.
Why this exists
Designing an acidic-ON binder means engineering titratable contacts that gain favourable interaction when… See the full description on the dataset page: https://huggingface.co/datasets/HarjasG/egfr-domain3-anchor-map.clinical-diagnostic-inference-error-amplification-mapping-v0.1What this dataset tests
How small inference errors introduced at a decision nodeamplify into downstream diagnostic distortion.
Required outputs
error entry node
inference error type
amplification factor
downstream distortion map
delay and misdiagnosis probabilities
self-correction points
prevention guardrails
market-narrative-coherence-mapping-v0.1What this dataset tests
Whether a system can detect market narrative coherenceacross heterogeneous sources.
This is not sentiment scoring.This is convergence detection.
Required outputs
narrative theme
coherence score
cross-source alignment
narrative velocity
price alignment state
Narrative velocity labels
building
steady
accelerating
shock jump
fragmenting
Price alignment states
underpriced
partial alignment
aligned
misaligned
Constraints
Do not predict… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-coherence-mapping-v0.1.mapped_orthologs
Zoonomia Orthologs Dataset
This dataset contains mapped orthologs for various species from the Zoonomia Project. It includes information about protein alignments and ortholog mappings between human and other species.
Dataset Structure
The dataset consists of a single CSV file with the following columns:
transcript: Human transcript ID
protein: Human protein name
mapped_to: Non-human (query) organism transcript ID
species: Name of the query species
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/mapped_orthologs.vegetation-mapping-lidar-us-ontology
Orthoimagery and LiDAR vegetation mapping ontology and register for an earth observation site
The object model and the model and equipment register from Baseline: One Flight of 4 Band Orthoimagery and LiDAR for Vegetation Mapping, an open reference architecture by CodeNinja Atoms for United States. Part of the Vertical-Driven Architectures series; every design in the series is also a row in the cumulative dataset… See the full description on the dataset page: https://huggingface.co/datasets/CodeNinjatools/vegetation-mapping-lidar-us-ontology.F1-driver-car-harmonic-efficiency-and-energy-waste-mapping-v0.1What this dataset tests
Whether a system can detectharmonic inefficiency in driver-car coupling.
Focus
Overcorrection loopsoscillation signaturesenergy leaksegment efficiency rank
Required outputs
harmonic waste index
correction loop density
oscillation signature type
energy leak score
efficiency rank by segment
All scores0 to 1
Highermeans more waste.
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.mapa-da-discriminacao-racial-no-brasil
Mapa da Discriminação Racial no Brasil
Este dataset contém coeficientes de discriminação racial por município no Brasil, calculados a partir de dados do Censo Demográfico.
Variáveis
cod_mun: Código do município (IBGE)
coef: Coeficiente base (intercepto) para cada município
mulher_negra: Coeficiente para mulheres negras
homem_negro: Coeficiente para homens negros
mulher_branca: Coeficiente para mulheres brancas
superior: Coeficiente para pessoas com ensino… See the full description on the dataset page: https://huggingface.co/datasets/atlas-da-saude-mental/mapa-da-discriminacao-racial-no-brasil.aviation-propulsion-aerodynamics-failure-horizon-intervention-mapping-v0.1What this dataset tests
Whether a system can turn detected decoherence
into an operational action plan.
It must estimate horizon,
choose intervention,
and define the decision window.
Required outputs
failure_horizon_minutes
recommended_derate_level
diversion_priority
stability_recovery_probability
intervention_window
action_rationale_channels
Scoring conventions
horizon is minutes to critical instability
diversion priority is low, medium, high, or urgent
intervention window is… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-propulsion-aerodynamics-failure-horizon-intervention-mapping-v0.1.aviation-pilot-vehicle-adaptive-intervention-and-recovery-mapping-v0.1What this dataset tests
Whether a system can turn loop decoherenceinto stabilizing action under time pressure.
It must estimate:
riskrecovery windowbest adaptive interventionexpected post-action stability.
Required outputs
loss_of_control_risk
recovery_window_seconds
recommended_adaptive_action
intervention_priority
recovery_confidence
post_action_stability_expectation
Use case
Layer three of Pilot–Vehicle Loop Coherence Under Stress.
Supports:
enhanced crew alerting
adaptive… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-pilot-vehicle-adaptive-intervention-and-recovery-mapping-v0.1.
