Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thomasmustier /pi-for-excel-sessions Coding agent session traces for thomasmustier/pi-for-excel-sessions This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-for-excel-sessions.tabulartext-generationn<1K8 likes1.4k downloads4mo agoHugging Face02Azzindani /IDX_Financial_Statements_Excel IDX_Financial_Statements_Excel This repository contains a comprehensive collection of financial statement data from public companies listed on the Indonesia Stock Exchange (IDX). The data is structured to facilitate automated financial analysis, trend monitoring, and the training of machine learning models for the Indonesian capital market. Dataset Overview The dataset provides structured access to the primary components of corporate financial reporting: Statement of… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/IDX_Financial_Statements_Excel.table-question-answering0 likes1k downloads7mo agoHugging Face03openeurollm /OpenEuroLLM-Excellent-SFT OpenEuroLLM Excellent SFT Mixture OpenEuroLLM Excellent SFT is a multilingual supervised fine-tuning mixture containing 6,187,716 conversations. It combines instruction-following, reasoning, mathematics, code, STEM, multilingual, and selected tool-use data. The dataset was produced from decontaminated versions of five public datasets and curated using the mn5_propella pipeline. Dataset composition Source Examples Nemotron Post-Training Dataset v2 3,755… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/OpenEuroLLM-Excellent-SFT.texttext-generation1M<n<10M0 likes373 downloads12d agoHugging Face04liswei /Taiwan-Text-Excellence-2Bgated High quality corpus for Taiwanese culture and Traditional Chinese Taiwan Text Excellence (TTE) Contains high quality news and articles in Traditional Chinese. The data processing pipeline is optimized for LLM performance. Is de-duplicated and cleaned using both rule-based and learning-based filters. E.g., urls/emails/html tags/abnormal characters are cleaned, and numbers (full-width or half-width) are normalized. Contains ~2 billion tokens, measured using BPE tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/liswei/Taiwan-Text-Excellence-2B.texttext-generation1M<n<10M22 likes57 downloads2y agoHugging Face05birgermoell /oellm-excellent-dpo-v0.1 OELLM Curated Pair-Decontaminated DPO v0.1 73,964 preference pairs from three public source families, in Parquet train, validation, and test splits. This is an independent research release in the user's namespace, not an official OpenEuroLLM dataset. It has traceable original preference labels, a strict Propella chosen-response gate for Dolci, and pair-level evaluation decontamination. Read the short methods paper, machine-readable manifest, and independent validation report… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-excellent-dpo-v0.1.tabulartext-generation10K<n<100K0 likes52 downloads3d agoHugging Face06agentlans /Taiwan-Text-Excellence-sentencesgated 台灣文摘句資料集 概述 Taiwan-Text-Excellence 句子資料集是從較大的 liswei/Taiwan-Text-Excellence-2B 資料集中抽取的 200 萬個獨特中文句子的綜合集。這些句子是隨機選取的,並使用 chinese-sentence-processor 工具進行分割。此資料集非常適合各種自然語言處理任務,包括語言建模、文本生成和其他研究用途。 資料集統計 總句數: 2,000,000 訓練集: 1,600,000 個句子 測試集: 400,000 個句子 資料格式 資料集中的每一行都包含一個欄位: **text**:包含中文句子的字串。 範例 {"text": "而新郎和女方家人的脂燭在當晚亦會合二為一,再送到母屋帳前點燃一個燈籠,保持三日不滅。"} {"text": "這個時期簽署的現代劇至今仍是台灣戲劇的中堅力量,而這十年為之奮鬥也奠定了其後數十年的基礎。"} {"text": "曾柏瑜今天也車票,明天還請在規劃畫相關票活動中。"}… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Taiwan-Text-Excellence-sentences.texttext-generation1M<n<10M0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.