cloud
Datasets
All datasets matching “cloud”cloudcloudsen12
This dataset follows the TACO specification.
cloudsen12plus
Website: https://cloudsen12.github.io/
version: 1.1.2
The largest dataset of expert-labeled pixels for cloud and cloud shadow detection in Sentinel-2
CloudSEN12+ version 1.1.0 is a significant extension of the CloudSEN12 dataset, which doubles the number of
expert-reviewed labels, making it, by a large margin, the largest cloud detection dataset to
date for Sentinel-2. All labels from the previous version have… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/cloudsen12.clouda-ocr-canonical-v1
Clouda OCR canonical v1
The cleanup-approved training corpus contains 28,691 procedural synthetic
pages with stable IDs, unchanged images and GT, source provenance, and hashes.
Use manifests/train_approved.parquet and the three canonical_v1/shards/train_approved/train-*.tar
shards. Each page has <id>.meta.json, <id>.gt.txt, and <id>.img.png members.
The shards are deterministic and range from approximately 1.38 to 2.08 GB.
manifests/sources.parquet, licenses.parquet… See the full description on the dataset page: https://huggingface.co/datasets/cloudaocr/clouda-ocr-canonical-v1.clouds-decoded-rain-check
Sentinel-2 Cloud Property Retrievals over the UK and India
Rendered per-scene quicklook layers of cloud optical and microphysical properties
retrieved from Sentinel-2 (MSI) Level-1C imagery over eight tiles (four in the UK
and four in India). These images power the interactive report at
asterisk-labs-clouds-decoded-rain-check.static.hf.space. Cloud properties are estimated with the clouds-decoded retrieval code, developed as part of the clouds decoded project funded by ARIA.… See the full description on the dataset page: https://huggingface.co/datasets/asterisk-labs/clouds-decoded-rain-check.alania-synthetic-speech-tr
Alania Turkish Synthetic Speech
English · Türkçe
3,411 hours of 48 kHz Turkish speech in 2,037,461 clips, spoken by 2,752 designed voices plus
one-off voices, each clip paired with the written sentence, its spoken form and a plain-English description of the
voice. Every clip is AI-generated: no real person's voice is in this dataset.
We made it while building Alania-2, the Turkish text-to-speech model behind
speech.patientdesk.ai. Openly licensed Turkish speech for TTS is… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-synthetic-speech-tr.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.
