spanish
Datasets
All datasets matching “spanish”spanish_spear_phishingDataset traducido del inglés al español mediante gpt4o mini.
Los mensajes del dataset contienen:
"email_subject": título del correo, no traducido
"sender_name": nombre del emisor, no traducido
"original_email_body": cuerpo del correo original, no traducido
"translated_email_body": cuerpo del correo traducido
El dataset corresponde al dataset de https://github.com/nahmiasd/Prompted-Contextual-Vectors-for-Spear-Phishing-Detection, el cual esta compuesto de:
"enron_ham": mensajes legítimos del… See the full description on the dataset page: https://huggingface.co/datasets/Darito/spanish_spear_phishing.messirveJuly 2025 UPDATE: We released version 1.1, adding almost 200k new queries 🎉🎉🎉.
v1.2 further adds the article titles as columns for convenience.
Use with:
country = "full" # "ar", "bo", ...
version = "1.2"
dataset = datasets.load_dataset("spanish-ir/messirve", country, revision=version)
print(dataset)
Dataset Card for MessIRve
MessIRve is a large-scale dataset for Spanish IR, designed to better capture the information needs of Spanish speakers across different countries.… See the full description on the dataset page: https://huggingface.co/datasets/spanish-ir/messirve.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.dpo-spanish-interpreted
Dataset Card for DPO Spanish Interpreted
Dataset Description
This dataset contains Direct Preference Optimization (DPO) pairs translated and culturally adapted into Spanish from an English source dataset (e.g. ultrafeedback or safe-rlhf).
Origin and Methodology
Model Used for Synthesis: kimi-k3 (via local endpoint)
Translation Strategy: The dataset was generated using a specialized prompt designed to prevent "translationese" and direct literal… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/dpo-spanish-interpreted.SpanishBCBL
DECOMEG — Brain Activity During Typing (MEG & EEG)
Non-invasive brain recordings (magnetoencephalography, MEG; and electroencephalography, EEG)
of healthy adults typing briefly-memorized sentences on a QWERTY keyboard. This is the dataset
underlying Brain2Qwerty (Lévy et al., 2025) and its companion neuroscience study
(Zhang et al., 2025).
Summary
Participants: 35 healthy adult volunteers recruited at the Basque Center on Cognition,
Brain and Language (BCBL), San… See the full description on the dataset page: https://huggingface.co/datasets/bcbl190626/SpanishBCBL.TransWeb-Edu-Spanish
