Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GGUFGuy /wikipedia-viewerwikipedia dataset, now with viewer enabled! :D Dataset Card for Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). The articles are parsed using the mwparserfromhell tool, which can be… See the full description on the dataset page: https://huggingface.co/datasets/GGUFGuy/wikipedia-viewer.texttext-generation10M<n<100M0 likes442 downloads2mo agoHugging Face02rcds /swiss_court_view_generationThis dataset contains court decision for court view generation task.text-generation100K<n<1M3 likes323 downloads3y agoHugging Face03leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes131 downloads6mo agoHugging Face04jiosephlee /auxiliary-views-knowledge-acquisition Auxiliary Views Knowledge Acquisition This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180). News August 21, 2026: Our paper was accepted to Findings of EMNLP 2026. Configurations Configuration Split Rows documents train 30 factual_cloze test 6,435 factual_mcqa_5shot test 4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.texttext-generation10K<n<100K1 likes124 downloads1mo agoHugging Face05aaaaliou /pi-sessions-viewer Coding agent session traces for aaaaliou/pi-sessions-viewer This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.tabulartext-generationn<1K0 likes113 downloads6mo agoHugging Face06Siddish /change-my-view-subreddit-cleaned Opinionated LLM texttext-generation1K<n<10K1 likes96 downloads3y agoHugging Face07quickium /awall-pi-strategy-views Awall PI Strategy Views 4.551 linhas (1.342 âncoras) — um subconjunto do quickium/awall-pi-quartets em que cada prompt ganhou quatro descrições textuais geradas por LLM, olhando o mesmo ataque por ângulos diferentes. view mediana o que descreve strategy_desc 279 chars a técnica do prompt, direta strategy_esp 349 chars a técnica em contraste com a população de prompts parecidos strategy_alt 476 chars a leitura alternativa — como o prompt se defenderia de ser ataque… See the full description on the dataset page: https://huggingface.co/datasets/quickium/awall-pi-strategy-views.tabulartext-generation1K<n<10K0 likes81 downloads18d agoHugging Face08davanstrien /query-to-dataset-viewer-descriptions Queries to Hugging Face Hub Datasets Views Dataset Summary This dataset consists of synthetically generated queries for datasets mapped to datasets on the Hugging Face Hub. The queries map to a datasets viewer API response summary of the dataset. The goal of the dataset is to train sentence transformer and ColBERT style models to map between a query from a user and a dataset without relying on a dataset card, i.e., using information in the dataset itself. Quick… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/query-to-dataset-viewer-descriptions.textsentence-similarity10K<n<100K5 likes78 downloads2y agoHugging Face09MK-runner /Multi-view-CXRgated EVOKE Patient-Specific Multimodal Learning with Multi-View Contrastive Alignment for Chest X-ray Report Generation Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology report generation presents a promising solution, existing methods often rely on single-view radiographs, which constrain… See the full description on the dataset page: https://huggingface.co/datasets/MK-runner/Multi-view-CXR.imagetext-generationn<1K2 likes33 downloads23d agoHugging Face10ENERZAiKR /LGUplus_viewing_historygated LGUplus Persona TV Viewing History 한국인 가상 페르소나 1만 명의 일주일치 TV 시청이력 합성 데이터셋입니다. nvidia/Nemotron-Personas-Korea 페르소나와 실제 U+tv EPG(258채널 × 7일, 2026-06-03~09) 편성표를 기반으로, 교사 LLM(Qwen3-235B-A22B-Instruct-2507-FP8)이 2단계로 생성했습니다. 생성 방법 시청 스케줄 생성: 페르소나(직업·나이·가족·취미)를 보고 요일별 시청 시간대를 추정 (근거 문장을 schedule_reasoning으로 함께 생성 — 예: 주부/은퇴자는 평일 낮, 직장인은 저녁) 프로그램 선택: 각 (요일, 시간창)마다 해당 시간에 방영 중인 실제 편성표를 제시하고 페르소나가 시청할 프로그램을 순서대로 선택 (채널 현실성 지시: 주류 채널 위주, 전문채널은 취미·직업 일치 시에만 / 이미 본 (제목,회차)… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_viewing_history.texttext-generation10K<n<100K0 likes20 downloads3mo agoHugging Face11pre-view /VTSNLP-vietnamese-curated-1MThis dataset contain 1 000 000 examples from https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset texttext-generation1M<n<10M0 likes14 downloads2y agoHugging Face12davanstrien /dataset-viewer-descriptions-processedtexttext-generation1K<n<10K0 likes10 downloads2y agoHugging Face13ebowwa /NSFW_RP_Format_DPO_ViewerThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action* This dataset contains NSFW content. texttext-generationn<1K4 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.