sin
Datasets
All datasets matching “sin”audiofolder_single_config_in_metadataPhysicalAI-Robotics-Manipulation-SingleArm
Dataset Description:
PhysicalAI-Robotics-Manipulation-SingeArm is a collection of datasets of automatic generated motions of a Franka Panda robot performing operations such as block stacking, opening cabinets and drawers. The dataset was generated in IsaacSim leveraging task and motion planning algorithms to find solutions to the tasks automatically [1, 3]. The environments are table-top scenes where the object layouts and asset textures are procedurally generated [2].This dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-SingleArm.wikipedia-paragraphs
wikipedia-paragraphs
wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research.
Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID.
Dataset structure
Configurations
The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.4KLSDB
4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation
DataCV @ CVPR 2026 · Accepted 🎉
4KLSDB is a native-4K image dataset with 129,484 train / 2,000 val / 1,984 test images, spanning nature, urban scenes, people, food, artwork, CGI, animals, and architecture. It supports both image restoration (super-resolution) and 4K text-to-image generation.
Quick links · 🌐 Project page · 💻 Code (GitHub) · 📄 Paper (arXiv) · 🤗 Dataset · 🧱 Checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/SingleBicycle/4KLSDB.SpeechRu
Russian Podcasts (unlabeled)
~186k unlabeled Russian-language podcast episodes scraped from the web,
packaged as Parquet shards with the audio bytes embedded. The audio has no
transcripts — this is an unsupervised / self-supervised audio corpus,
suitable for ASR pre-training, speech-representation learning, TTS data
mining, audio classification, and similar tasks.
Each row contains:
audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo),
decoded on-the-fly via the… See the full description on the dataset page: https://huggingface.co/datasets/Sinoosoida/SpeechRu.Sindhi-texts-big-dataset
Sindhi Texts (big dataset)
A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It
combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia,
newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora.
3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a
Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5
characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.
