datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
curatorkit-testrun-CSVJSONParquet
curatorkit-testrun-CSVJSONParquet
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 05:51 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-CSVJSONParquet", "alpaca")
philosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.minimum_viable_articulation_v01.csv
Minimum Viable Articulation (MVA)
MVA measures a model’s ability to answer with the minimum viable output — no surplus explanation, no self-expansion, no tutorial behavior.
This dataset evaluates where models fail to stop:
Overcompletion
Hedging / padding
Teaching when not asked
Identity or stance leakage
Solving beyond scope
It exposes a behavior pattern where models confuse helpfulness with verbosity and treat extra tokens as value, rather than distortion.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/minimum_viable_articulation_v01.csv.dataset_aeroespacial_cultural_completo.csv
🚀 LATAM Aerospace Cultural QA
Dataset culturalmente alineado para modelos conversacionales en español y portugués, especializado en historia aeroespacial iberoamericana.
Desarrollado para el #HackathonSomosNLP 2026 — ¿Son los LLMs realmente multiculturales?
🛠 Metodología y Pipeline de Construcción
La versión actual del dataset ha sido refinada mediante un pipeline automatizado diseñado para maximizar la calidad y la diversidad cultural:
Generación Dinámica: Se generan… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset_aeroespacial_cultural_completo.csv.structeval-t-sft-hq-csv
StructEval-T SFT - High Quality CSV
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated CSV transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid CSV without errors are included.
Goal: To maximize single-format fine-tuning performance or to be used… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-csv.structeval-t-sft-v2-csv
StructEval-T SFT v2 - Full CSV
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated CSV transformations.
Key Features
Total Samples: 2,104
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid CSV without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-csv.Indian_language_community_chatbot.csvOpenHermes-CSVsentence2000_csv
