Team Ai
20 results

sin

hf-internal-testing /audiofolder_single_config_in_metadataaudion<1K0 likes101k downloads3y agoHugging Facenvidia /PhysicalAI-Robotics-Manipulation-SingleArm Dataset Description: PhysicalAI-Robotics-Manipulation-SingeArm is a collection of datasets of automatic generated motions of a Franka Panda robot performing operations such as block stacking, opening cabinets and drawers. The dataset was generated in IsaacSim leveraging task and motion planning algorithms to find solutions to the tasks automatically [1, 3]. The environments are table-top scenes where the object layouts and asset textures are procedurally generated [2].This dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-SingleArm.robotics21 likes18k downloads1y agoHugging Facesingletongue /wikipedia-paragraphs wikipedia-paragraphs wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research. Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID. Dataset structure Configurations The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.tabular100M<n<1B2 likes17k downloads3mo agoHugging FaceSingleBicycle /4KLSDB 4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation DataCV @ CVPR 2026 · Accepted 🎉 4KLSDB is a native-4K image dataset with 129,484 train / 2,000 val / 1,984 test images, spanning nature, urban scenes, people, food, artwork, CGI, animals, and architecture. It supports both image restoration (super-resolution) and 4K text-to-image generation. Quick links · 🌐 Project page · 💻 Code (GitHub) · 📄 Paper (arXiv) · 🤗 Dataset · 🧱 Checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/SingleBicycle/4KLSDB.imageimage-to-image100K<n<1M15 likes12k downloads5mo agoHugging FaceSinoosoida /SpeechRu Russian Podcasts (unlabeled) ~186k unlabeled Russian-language podcast episodes scraped from the web, packaged as Parquet shards with the audio bytes embedded. The audio has no transcripts — this is an unsupervised / self-supervised audio corpus, suitable for ASR pre-training, speech-representation learning, TTS data mining, audio classification, and similar tasks. Each row contains: audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo), decoded on-the-fly via the… See the full description on the dataset page: https://huggingface.co/datasets/Sinoosoida/SpeechRu.audioautomatic-speech-recognition100K<n<1M4 likes10k downloads3mo agoHugging Facearnizamani /Sindhi-texts-big-dataset Sindhi Texts (big dataset) A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia, newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora. 3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5 characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.texttext-generation100K<n<1M4 likes8k downloads2mo agoHugging Face