datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-viewerwikipedia dataset, now with viewer enabled! :D
Dataset Card for Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(https://dumps.wikimedia.org/) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
The articles are parsed using the mwparserfromhell tool, which can be… See the full description on the dataset page: https://huggingface.co/datasets/GGUFGuy/wikipedia-viewer.swiss_court_view_generationThis dataset contains court decision for court view generation task.longbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.auxiliary-views-knowledge-acquisition
Auxiliary Views Knowledge Acquisition
This repository contains the cleaned source documents and evaluation
probes used in Knowledge Acquisition During Pre-training? Large Language Models
Learn Better With Auxiliary Views (arXiv:2609.04180).
News
August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.
Configurations
Configuration
Split
Rows
documents
train
30
factual_cloze
test
6,435
factual_mcqa_5shot
test
4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.pi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.change-my-view-subreddit-cleaned
Opinionated LLM
awall-pi-strategy-views
Awall PI Strategy Views
4.551 linhas (1.342 âncoras) — um subconjunto do
quickium/awall-pi-quartets em que
cada prompt ganhou quatro descrições textuais geradas por LLM, olhando o mesmo ataque por
ângulos diferentes.
view
mediana
o que descreve
strategy_desc
279 chars
a técnica do prompt, direta
strategy_esp
349 chars
a técnica em contraste com a população de prompts parecidos
strategy_alt
476 chars
a leitura alternativa — como o prompt se defenderia de ser ataque… See the full description on the dataset page: https://huggingface.co/datasets/quickium/awall-pi-strategy-views.query-to-dataset-viewer-descriptions
Queries to Hugging Face Hub Datasets Views
Dataset Summary
This dataset consists of synthetically generated queries for datasets mapped to datasets on the Hugging Face Hub. The queries map to a datasets viewer API response summary of the dataset. The goal of the dataset is to train sentence transformer and ColBERT style models to map between a query from a user and a dataset without relying on a dataset card, i.e., using information in the dataset itself.
Quick… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/query-to-dataset-viewer-descriptions.Multi-view-CXR
EVOKE
Patient-Specific Multimodal Learning with Multi-View Contrastive Alignment for Chest X-ray Report Generation
Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology report generation presents a promising solution, existing methods often rely on single-view radiographs, which constrain… See the full description on the dataset page: https://huggingface.co/datasets/MK-runner/Multi-view-CXR.LGUplus_viewing_history
LGUplus Persona TV Viewing History
한국인 가상 페르소나 1만 명의 일주일치 TV 시청이력 합성 데이터셋입니다.
nvidia/Nemotron-Personas-Korea 페르소나와
실제 U+tv EPG(258채널 × 7일, 2026-06-03~09) 편성표를 기반으로, 교사 LLM(Qwen3-235B-A22B-Instruct-2507-FP8)이
2단계로 생성했습니다.
생성 방법
시청 스케줄 생성: 페르소나(직업·나이·가족·취미)를 보고 요일별 시청 시간대를 추정
(근거 문장을 schedule_reasoning으로 함께 생성 — 예: 주부/은퇴자는 평일 낮, 직장인은 저녁)
프로그램 선택: 각 (요일, 시간창)마다 해당 시간에 방영 중인 실제 편성표를 제시하고
페르소나가 시청할 프로그램을 순서대로 선택 (채널 현실성 지시: 주류 채널 위주,
전문채널은 취미·직업 일치 시에만 / 이미 본 (제목,회차)… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_viewing_history.VTSNLP-vietnamese-curated-1MThis dataset contain 1 000 000 examples from https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset
dataset-viewer-descriptions-processedNSFW_RP_Format_DPO_ViewerThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action*
This dataset contains NSFW content.
