Team Ai
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-DataLab /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.tabular10K<n<100K76 likes965 downloads1y agoHugging Face02GG-samrt /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/GG-samrt/DataScience-Instruct-500K.tabular10K<n<100K0 likes363 downloads8mo agoHugging Face03DataScience-UIBK /PopMCQ 🎯 PopMCQ Does your model pick the famous answer or the correct one? 📌 Overview PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty. The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.tabularquestion-answering10K<n<100K1 likes187 downloads1mo agoHugging Face04Fan0718 /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/Fan0718/DataScience-Instruct-500K.tabular10K<n<100K0 likes177 downloads9mo agoHugging Face05RazinAleks /SO-Python_QA-Data_Science_and_Machine_Learning_classtabular1K<n<10K6 likes121 downloads3y agoHugging Face06DataScience-UIBK /OBLIQ-IR-Data OBLIQ-IR-Data The training data behind DataScience-UIBK/OBLIQ-IR-3B, plus the retrieval runs and evaluation outputs for every result in OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026). Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance, an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has little or no surface expression in the document. 🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.texttext-retrieval100K<n<1M4 likes118 downloads15d agoHugging Face07fantos /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/fantos/DataScience-Instruct-500K.tabular10K<n<100K0 likes104 downloads11mo agoHugging Face08nathansutton /data-science-job-descriptions Data Science Job DescriptionsThese data encompass the title, company, and description of the outer-join job board between October 2021 and today. license: wtfpl task_categories: - text-classification - feature-extraction language: - en tags: - jobs pretty_name: ds-jobs size_categories: - 1K<n<10K text1K<n<10K3 likes64 downloads3y agoHugging Face09binzhango /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/binzhango/DataScience-Instruct-500K.tabular10K<n<100K0 likes54 downloads11mo agoHugging Face10stindardlogic /data-science-workflows-sft-100k Data Science Workflows SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment. Motivation Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.texttext-generation100K<n<1M0 likes43 downloads3mo agoHugging Face11DevHunterAI /turkish-science-data Turkish Science Data Yüksek kaliteli Türkçe fen bilimleri sentetik veri seti. İçerik Alan Konu Başlıkları Fizik Kinematik, Newton yasaları, Enerji/İş, Elektrik devreleri Kimya Mol hesabı, Asit-baz/pH, Kimyasal denklemler, Periyodik tablo Biyoloji Mendel genetiği, DNA, Mitoz/Mayoz, Organeller, Hormonlar Format Her örnek soru-cevap formatında Türkçe düz metin: Soru: ... Çözüm: ... Cevap: ... Kullanım from datasets… See the full description on the dataset page: https://huggingface.co/datasets/DevHunterAI/turkish-science-data.text100K<n<1M0 likes28 downloads4mo agoHugging Face12Hamzasajjad38 /data-science-chatbot 📊 Data Science Chatbot Dataset (2000 Samples) 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts. This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a Data Science Tutor Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.texttext-generation1K<n<10K0 likes15 downloads6mo agoHugging Face13hadxs /Connor-Data-Sciences_humainestext10K<n<100K0 likes8 downloads4mo agoHugging Face14eshmoideas /DataScience-ML-DATASETSgated MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/eshmoideas/DataScience-ML-DATASETS.texttext-generation1M<n<10M1 likes6 downloads4mo agoHugging Face15hadxs /Connor-Data-Sciencestext10K<n<100K0 likes6 downloads4mo agoHugging Face16hadxs /Connor-Data-Datasciencetext10K<n<100K0 likes5 downloads4mo agoHugging Face17drxzero /Final-ALevel-Science-Datatext1K<n<10K1 likes4 downloads4mo agoHugging Face18hadxs /Connor-Data-AEX-sciencestext1K<n<10K0 likes4 downloads4mo agoHugging Face19DataScienceGroup /chat01textn<1K0 likes3 downloads8mo agoHugging Face20beldua /english-data-science-basics-30textn<1K0 likes3 downloads10mo agoHugging Face21hadxs /Connor-Data-AEX-sciences_humainestext1K<n<10K0 likes3 downloads4mo agoHugging Face22Pranav2640 /intrain_data_sciencetextn<1K0 likes2 downloads2y agoHugging Face23DataScienceGroup /imageimagen<1K0 likes2 downloads9mo agoHugging Face24DataScienceGroup /image-text-jsontextn<1K0 likes2 downloads9mo agoHugging Face25DataScienceGroup /image-text2textn<1K0 likes2 downloads9mo agoHugging Face26hadxs /Connor-Data-AEX-datasciencetext100K<n<1M0 likes2 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.