Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01csaybar /CloudSEN12-scribble🚨 New Dataset Version Released! We are excited to announce the release of Version [1.1] of our dataset! This update includes: [L2A & L1C support]. [Temporal support]. [Check the data without downloading (Cloud-optimized properties)]. 📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab CloudSEN12 NOLABEL A Benchmark Dataset for Cloud Semantic Understanding CloudSEN12 SCRIBBLE A Benchmark Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-scribble.tabular10K<n<100K0 likes2.7k downloads2y agoHugging Face02csaybar /CloudSEN12-high 🚨 New Dataset Version Released! We are excited to announce the release of Version [1.1] of our dataset! This update includes: [L2A & L1C support]. [Temporal support]. [Check the data without downloading (Cloud-optimized properties)]. 📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab CloudSEN12 HIGH-QUALITY A Benchmark Dataset for Cloud Semantic Understanding CloudSEN12… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-high.tabular10K<n<100K2 likes1.9k downloads2y agoHugging Face03csaybar /CloudSEN12-nolabel🚨 New Dataset Version Released! We are excited to announce the release of Version [1.1] of our dataset! This update includes: [L2A & L1C support]. [Temporal support]. [Check the data without downloading (Cloud-optimized properties)]. 📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab CloudSEN12 NOLABEL A Benchmark Dataset for Cloud Semantic Understanding CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-nolabel.tabular10K<n<100K0 likes1.6k downloads2y agoHugging Face04tracer-cloud /opensre OpenSRE / OpenRCA dataset Root cause analysis benchmark data: queries, incident records, telemetry (metrics, logs, traces), and query_alerts (per-row JSON derived from each query.csv). Telemetry CSVs use different schemas by file type; load them by path (they are not merged into the Hub subset configs above). Original archives are also described here: Google Drive. Regenerating query_alerts python3 scripts/query_csv_to_alert_json.py textn<1K0 likes273 downloads6mo agoHugging Face05gpueconomy /cloud-gpu-price-index Cloud GPU Price Index Canonical source: https://gpueconomy.com/price-index. That page is recomputed every hour; this record is a dated snapshot of it, version 2026-10-02, built from data updated 2026-10-02T13:39:58.768836+00:00. When you cite, cite GPU Economy and link the page; the snapshot is here so that a number you used keeps existing exactly as you used it. The index is the weekly median publicly listed on-demand price of one NVIDIA H100 SXM GPU-hour across the cloud GPU… See the full description on the dataset page: https://huggingface.co/datasets/gpueconomy/cloud-gpu-price-index.tabularn<1K0 likes190 downloads8d agoHugging Face06Float16-cloud /ThaiIDCardSynt Dataset Details Dataset Description Curated by: Matichon Maneegard Shared by [optional]: Matichon Maneegard Language(s) (NLP): image-to-text License: apache-2.0 Dataset Sources [optional] The dataset was entirely synthetic. It does not contain real information or pertain to any specific person. Uses Direct Use Using for tranning OCR or Multimodal. Dataset Structure This dataset contains 98 x 6 = 588 samples, and the… See the full description on the dataset page: https://huggingface.co/datasets/Float16-cloud/ThaiIDCardSynt.imageimage-to-textn<1K2 likes156 downloads3y agoHugging Face07unum-cloud /ann-arxiv-2m 2M Title-Abstract Arxiv Pairs title_abstract.tsv data from Cornell University Arxiv Dataset, preprocessed and coverted to TSV. title.e5-base-v2.fbin is a binary file with e5-base-v2 title embeddings. abstract.e5-base-v2.fbin is a binary file with e5-base-v2 abstract embeddings. text1M<n<10M6 likes142 downloads6mo agoHugging Face08ByteDance /CloudTimeSeriesData Intro The data organization follows TFB format: https://github.com/decisionintelligence/TFB. TFB data format TFB stores time series in a format of three column long tables, which we will introduce below: Format Introduction First column: date (the exact column name is required, the same applies below.) The columns stores the time information in the time series, which can be in either of the following formats: Timestamps in string, datetime, or other types… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/CloudTimeSeriesData.text10M<n<100M1 likes113 downloads2y agoHugging Face09fastgpu /cloud-gpu-prices Cloud GPU Rental Prices (FastGPU) What it costs to rent a GPU in the cloud, per GPU per hour, across marketplaces, neoclouds and hyperscalers. Mirrored here once a day from fastgpu.co/dataset, where the same files are served live and every price links to the provider it came from. Browse the prices themselves at fastgpu.co. A free, openly licensed dataset of GPU cloud rental prices across the whole market: marketplaces, neoclouds, and hyperscalers. It contains a normalized… See the full description on the dataset page: https://huggingface.co/datasets/fastgpu/cloud-gpu-prices.tabulartime-series-forecasting10K<n<100K0 likes71 downloads13h agoHugging Face10clouddress /recentprofit-official-savings-data RecentProfit Data Research v1.0 This repository packages RecentProfit's U.S. brand purchase-condition research as fact-level and comparison datasets. The canonical human-readable documentation, methodology, source links, and version history are on RecentProfit Data Research. Author: RecentProfit Schema version: v1.0 Data version: 2026-09-13 Region / currency: US / USD Brands: 33 Current fact rows: 212 License: MIT Data files data/recentprofit-brand-facts.csv —… See the full description on the dataset page: https://huggingface.co/datasets/clouddress/recentprofit-official-savings-data.tabularn<1K0 likes68 downloads23d agoHugging Face11sairamn /gcp-cloud-billing-costtabular100K<n<1M0 likes54 downloads2y agoHugging Face12cloudsurf-software /CloudSurf-4B-FC-bfcl-results CloudSurf-4B-FC — raw BFCL V4 result files Raw, unmodified BFCL V4 evaluation outputs backing the leaderboard submission PR ShishirPatil/gorilla#1357 for CloudSurf-4B-FC (a google/gemma-4-E4B-it fine-tune, Apache-2.0). Both sides are included: our tuned runs and the stock gemma-4-E4B-it baselines re-measured on the identical rig, so every number in the PR can be recomputed from primary files. Whiskers are the min–max across the three runs on each side. Stock wins Irrelevance… See the full description on the dataset page: https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results.tabularn<1K0 likes47 downloads2mo agoHugging Face13cola-cloud /whiskey-beer-and-california-wine-label-metadata Whiskey, Beer and California Wine Label Metadata This Hugging Face repository is an exact mirror of the Kaggle release published by COLA Cloud LLC on October 2, 2026. The CSV, original data card, data dictionary, license and manifest are byte-identical to that release (CSV SHA256 d54398f6c022eee8c49107ec7f05e2bc0fe9df5f778e0cd1d0819b43b4b4772b). The Kaggle entry remains the canonical publication; this mirror exists so Hub search and Hugging Face loaders can reach the same… See the full description on the dataset page: https://huggingface.co/datasets/cola-cloud/whiskey-beer-and-california-wine-label-metadata.texttext-retrievaln<1K0 likes43 downloads5d agoHugging Face14stackscan /cloud-hosting Cloud and CDN Adoption Among Large Organizations Overview This dataset records which cloud, hosting and CDN providers were detected on the websites of 21,325 large organizations, with firmographic context for each: industry, employee band, country, locality and founding year. Nineteen providers are covered, each as its own column, because organizations commonly use several at once. 3,259 of them show more than one. Collected in August 2026. Infrastructure changes… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/cloud-hosting.tabulartabular-classification10K<n<100K1 likes33 downloads2mo agoHugging Face15Viswa-k /cloud-gpu-pricing Cloud GPU Pricing Index Monthly snapshot of cloud GPU prices per GPU-hour in USD, maintained by Nodus, an AI compute platform. Use it to answer "how much does an H100 / H200 / B200 / B300 cost per hour" or to budget a training run. October 2026 snapshot (Oct 8, 2026) GPU VRAM Nodus supply price $/GPU-hr Modal listed $/GPU-hr B300 288 GB 6.943 7.099 B200 180 GB 5.500 6.250 H200 SXM 141 GB 3.593 4.540 H100 SXM 80 GB 2.600 3.949 A100 80GB 80 GB… See the full description on the dataset page: https://huggingface.co/datasets/Viswa-k/cloud-gpu-pricing.tabularn<1K0 likes28 downloads2d agoHugging Face16cloud19 /gelbooru-characters-enriched Gelbooru Characters Enriched This dataset is an enriched, fully-mapped version of Gelbooru character tags. It contains resolved franchise (copyright) associations and core appearance features (core tags) for 263,441 unique characters. Dataset Details The dataset maps the original character list to their corresponding copyrights (franchises) and general core attributes. It was constructed using a multi-stage hybrid extraction pipeline: Regex Extraction: Extracting… See the full description on the dataset page: https://huggingface.co/datasets/cloud19/gelbooru-characters-enriched.tabulartext-classification100K<n<1M0 likes27 downloads3mo agoHugging Face17ClarusC64 /cloud-load-latency-coherence-risk-v0.1What this repo is for Catch latency collapse early. It flags when: latency spikes without load headroom looks fine but queues rise overflow triggers too early saturation appears but monitoring hides i texttext-classificationn<1K0 likes25 downloads8mo agoHugging Face18knarayan /cloud_posture_checks Dataset Card for Dataset Name Prisma Cloud curated dataset for known misconfiguration checks across Compliance and Security issues tracked across its customer base. Dataset Details Dataset Description Dataset that provides input on the specific json rules for all known misconfiguration states relevant for cloud security across multiple cloud providers. Useful to help expose data to LLMs to reason and enable free form interaction to understand cloud security… See the full description on the dataset page: https://huggingface.co/datasets/knarayan/cloud_posture_checks.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face19Ishitaagarwal /cloud_incidentstextn<1K0 likes21 downloads11mo agoHugging Face20DebasishDhal99 /cloud-vs-temperature-data Why was it made? It was made by collocating data from two satellites, INSAT-3DR and CLOUDSAT, dedicated for meteorology. INSAT-3DR is a geostationary satellite over Indian region, while CLOUDSAT is a polar satellite dedicated for precise cloud observation. Since CLOUDSAT is a polar satellite (altitude 450 KM), it can only look at one minute region at a time. But because of this we get precise cloud data (Cloudy-Clear distinction, Cloud Height and Cloud Thickness). Since INSAT-3DR… See the full description on the dataset page: https://huggingface.co/datasets/DebasishDhal99/cloud-vs-temperature-data.tabular100K<n<1M0 likes20 downloads2y agoHugging Face21knarayan /cloud_posture_checks_name_and_desctext1K<n<10K0 likes14 downloads2y agoHugging Face22aroravce /CloudTrailThreatHuntingtext10K<n<100K0 likes13 downloads2y agoHugging Face23infinite-dataset-hub /CloudTrailPatterns CloudTrailPatterns tags: data science, cloud services, anomaly detection Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'CloudTrailPatterns' dataset is designed for the purpose of identifying and classifying various patterns of cloud service usage that may indicate anomalies or potential security threats. This dataset is tailored for data science and machine learning applications, particularly in the field of anomaly… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CloudTrailPatterns.tabularn<1K0 likes12 downloads2y agoHugging Face24Roy229 /google-cloud_github_fetch_huggingface_terminal_6737_r6981fjctabularn<1K0 likes12 downloads2mo agoHugging Face25cloud19 /gelbooru-characters-onlytabular100K<n<1M0 likes11 downloads4mo agoHugging Face26Roy229 /google-cloud_github_fetch_huggingface_terminal_6737_bv8nvl1etabularn<1K0 likes10 downloads2mo agoHugging Face27Roy229 /google-cloud_github_fetch_huggingface_terminal_6737_3qt2dfa4tabularn<1K0 likes10 downloads2mo agoHugging Face28Roy229 /huggingface_google_sheet_google-cloud_368_category_rulestextn<1K0 likes9 downloads2mo agoHugging Face29amanish /cloud-logstabular10K<n<100K0 likes8 downloads1y agoHugging Face30cokesea /CloudTimeSeriesData Intro The data organization follows TFB format: https://github.com/decisionintelligence/TFB. TFB data format TFB stores time series in a format of three column long tables, which we will introduce below: Format Introduction First column: date (the exact column name is required, the same applies below.) The columns stores the time information in the time series, which can be in either of the following formats: Timestamps in string, datetime, or… See the full description on the dataset page: https://huggingface.co/datasets/cokesea/CloudTimeSeriesData.text10M<n<100M0 likes7 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.