Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01juliensimon /space-track-tle-history Space-Track TLE History Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports. Quick Start from datasets import load_dataset # Load a specific year ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet") # Load everything (238M rows — use streaming for large-scale analysis) ds =… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-tle-history.tabulartime-series-forecasting100M<n<1B2 likes5.1k downloads3h agoHugging Face02juliensimon /starlink-tle-latest Latest Starlink & GPS TLEs Credit: NASA Part of the Orbital Mechanics Datasets collection on Hugging Face. Dataset description Latest Two-Line Element (TLE) orbital data for the Starlink and GPS constellations, sourced daily from CelesTrak. Two-Line Element sets (TLEs) are the standard format for representing satellite orbital elements, developed by NORAD in the 1960s and still used universally today. Each TLE encodes six Keplerian orbital elements plus… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/starlink-tle-latest.texttabular-regression10K<n<100K0 likes1.9k downloads10h agoHugging Face03juliensimon /constellation-tle-latest Constellation TLEs -- 18 Satellite Constellations Credit: NASA Part of the Orbital Mechanics Datasets collection on Hugging Face. Dataset description Daily Two-Line Element (TLE) snapshots for 18 satellite constellations sourced from CelesTrak. Covers GNSS navigation (GPS, Galileo, BeiDou, GLONASS, SBAS), LEO broadband (OneWeb, Kuiper, Qianfan, Hulianwang), LEO communications (Iridium, Globalstar, ORBCOMM), Earth observation (Planet Labs, Spire Global)… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/constellation-tle-latest.texttabular-regression1K<n<10K1 likes1.2k downloads10h agoHugging Face04oxzoid /space-track-tle-history Space-Track TLE History Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports. Quick Start from datasets import load_dataset # Load a specific year ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet") # Load everything (238M rows — use streaming for large-scale analysis)ds =… See the full description on the dataset page: https://huggingface.co/datasets/oxzoid/space-track-tle-history.tabulartime-series-forecasting100M<n<1B0 likes509 downloads6mo agoHugging Face05tlemenestrel /Smiles2Dockhttps://arxiv.org/pdf/2406.05738 text10M<n<100M1 likes221 downloads1y agoHugging Face06TLeonidas /this-person-does-not-existThis dataset consists of 8892 AI-generated profile pictures downloaded from (https://www.kaggle.com/datasets/pablobedolla/this-person-does-not-exist-data) image1K<n<10K0 likes107 downloads3y agoHugging Face07Arailym-tleubayeva /KazakhLawCorpus-clean KazakhLawCorpus-clean Dataset Summary KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset. The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.tabulartext-retrieval100K<n<1M1 likes68 downloads4d agoHugging Face08TLeonidas /twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below: tdavidson/hate_speech_offensive LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset ucberkeley-dlab/measuring-hate-speech It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data. text100K<n<1M1 likes59 downloads3y agoHugging Face09Arailym-tleubayeva /NK-Oil-Well-Sensor-Monitoring Oil Well Sensor Monitoring Dataset - NK Field Dataset Description This dataset contains hourly sensor readings from 10 oil wells at the NK field. The monitoring period covers approximately seven months, from January 1, 2026, to July 20, 2026. The data were provided by Galaz and Company LLP (ТОО «Галаз и Компания») within the research project: “Development and Implementation of Control Algorithms for Low-Production-Rate Wells in Mechanized Oil Production Systems… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/NK-Oil-Well-Sensor-Monitoring.texttime-series-forecasting100K<n<1M0 likes39 downloads3mo agoHugging Face10Arailym-tleubayeva /KazOilWellOps_Dataset Kazakhstan Oil Well Operational Dataset Description This dataset contains structured operational and production parameters of sucker rod pump (SRP) oil wells in Kazakhstan. It is intended for industrial AI research, oil production analysis, production forecasting, and predictive modeling of well performance under real field operating conditions. Location: North-West Konys oil field, Kyzylorda Region, Kazakhstan (≈150 km NW of Kyzylorda city). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazOilWellOps_Dataset.tabulartabular-regressionn<1K0 likes35 downloads8mo agoHugging Face11Arailym-tleubayeva /small_kazakh_corpus Dataset Card for Small Kazakh Language Corpus The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs. Dataset Details Dataset Description The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.textmask-generation10K<n<100K1 likes30 downloads2y agoHugging Face12Arailym-tleubayeva /KazakhTextDuplicates Dataset Card for KazakhTextDuplicates Dataset Details Dataset Description The KazakhTextDuplicates dataset is a collection of Kazakh-language texts containing duplicates with different levels of modification. The dataset includes exact duplicates, contextual duplicates, and partial duplicates, making it valuable for research in text similarity, duplicate detection, information retrieval, and plagiarism detection. Developed by: Arailym Tleubayeva Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicates.text10K<n<100K1 likes28 downloads2y agoHugging Face13Arailym-tleubayeva /legalup-laws Kazakhstan Legal Acts Dataset (LegalUp) Dataset Summary The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan. The current release includes 392,084 legislative document records extracted from a PostgreSQL database. The dataset is designed for: Legal Retrieval-Augmented Generation (Legal RAG) Information Retrieval Legal Search Question Answering Semantic Search Legal NLP Benchmark Construction Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.textquestion-answering100K<n<1M0 likes22 downloads1mo agoHugging Face14Arailym-tleubayeva /sist-kazakh-corpus SIST Kazakh Corpus Description SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles collected for research in text similarity detection, plagiarism analysis, and low-resource NLP tasks. The dataset was created to support: Text similarity detection in agglutinative languages Kazakh NLP benchmarking Scientific text analysis Retrieval-Augmented Generation (RAG) research Dataset Structure The dataset is provided in CSV format. Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.tabularsentence-similarityn<1K1 likes21 downloads8mo agoHugging Face15Arailym-tleubayeva /LegalRAG Kazakh Legal Text Chunks Dataset Summary Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language. The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in: legal retrieval, legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.tabularquestion-answering10K<n<100K0 likes17 downloads7mo agoHugging Face16Arailym-tleubayeva /sist-english-corpustabulartext-classificationn<1K0 likes15 downloads1y agoHugging Face17Arailym-tleubayeva /AITUAdmissionsGuideDataset AITU Admissions Guide Dataset Dataset Details Dataset Description This dataset contains questions, answers, and categories related to the admission process at Astana IT University (AITU). It is designed to assist in automating applicant consultations and can be used for chatbot training, recommendation systems, and NLP-based question-answering models. Curated by: Astana IT University Funded by [optional]: Arailym Tleubayeva, Alina Mitroshina, Alpar Arman… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/AITUAdmissionsGuideDataset.texttable-question-answeringn<1K0 likes12 downloads2y agoHugging Face18tleegwater /herb_sheetstextn<1K0 likes12 downloads1y agoHugging Face19SUSTech /tlem-leaderboardtabularn<1K0 likes8 downloads3y agoHugging Face20tleo /oi_docs_datasettextn<1K0 likes3 downloads2y agoHugging Face21tleo /oi_docs_synthetic_alpacatext1K<n<10K0 likes3 downloads2y agoHugging Face22tleo /hungary_history_alpacatext10K<n<100K0 likes3 downloads10mo agoHugging Face23tleo /jozsef_attila_osszestextn<1K0 likes3 downloads7mo agoHugging Face24isc-tleavitt /opus100-multilingualtabular1K<n<10K0 likes2 downloads1y agoHugging Face25isc-tleavitt /g23-projecttabular1K<n<10K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.