Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucazhou2000 /sciencemysterybench-transcriptsimagen<1K0 likes2.3k downloads24d agoHugging Face02QasimHussain /spatial-transcriptomics-atlas-demo Spatial Transcriptomics Atlas: Human Lymph Node Architecture and Immune Microenvironment Integrated multi-modal analysis of spatial gene expression in human lymph node tissue. Panels depict high-resolution histology (A), annotated tissue domains (B), gene detection density (C), expression patterns of top spatially variable genes (D--F), and neighborhood enrichment statistics (G). Abstract This repository presents a reproducible computational… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/spatial-transcriptomics-atlas-demo.imagen<1K0 likes85 downloads2mo agoHugging Face03BDRC /monlamai-transcriptions Tibetan OCR — MonlamAI transcriptions 3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode transcriptions, contributed by MonlamAI over BDRC manuscript scans and aligned page by page. This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes: filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.imageimage-to-text1K<n<10K0 likes56 downloads2mo agoHugging Face04BDRC /berkeley-transcriptions Tibetan OCR — Berkeley 8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small woodblock (uchen) portion. These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley. This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.imageimage-to-text1K<n<10K0 likes55 downloads2mo agoHugging Face05BDRC /palri-parkhang-transcriptions Tibetan OCR — Palri Parkhang 11,133 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small uchen portion. The transcriptions were produced by Palri Parkhang, an input project led by Chris Tomlinson (former BDRC's CTO) in Nepal in 2006-2013 and aligned to BDRC scans; the material is largely Nyingma collected works and gter ma cycles. The transcriptions prioritized legibility over fidelity to the enscribed text and… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/palri-parkhang-transcriptions.imageimage-to-text10K<n<100K0 likes54 downloads2mo agoHugging Face06ashraq /youtube-transcriptionThis is YouTube video transcription dataset built from YTTTS Speech Collection for semantic search.image1K<n<10K5 likes45 downloads4y agoHugging Face07BSC-CSSH /AMSMB-line-transcription Dataset Card Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.imageimage-to-text1K<n<10K0 likes31 downloads1y agoHugging Face08pinecone /yt-transcriptionsimage10K<n<100K1 likes29 downloads4y agoHugging Face09justinsunqiu /multilingual_transcriptions_finalimage1K<n<10K0 likes19 downloads1y agoHugging Face10justinsunqiu /multilingual_transcriptions_fullimage1K<n<10K0 likes13 downloads1y agoHugging Face11justinsunqiu /multilingual_transcriptions_summarized_by_english_backtranslated_finalimage10K<n<100K0 likes13 downloads1y agoHugging Face12justinsunqiu /multilingual_transcriptions_translated_rawimage1K<n<10K0 likes12 downloads1y agoHugging Face13QasimHussain /transcriptome-health-dashboard-demo Transcriptome Health Dashboard v2.0 A professional-grade RNA-Seq quality control and analysis pipeline implementing biologically-rigorous normalization, interactive visualizations, and comprehensive sample QC metrics. Principal Component Analysis of 424 TCGA-LIHC samples visualizing transcriptomic structure. Overview This pipeline performs comprehensive quality control analysis for bulk RNA-Seq datasets, implementing industry-standard bioinformatics… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/transcriptome-health-dashboard-demo.imagen<1K0 likes11 downloads2mo agoHugging Face14justinsunqiu /multilingual_transcriptions_translated_english_finalimage1K<n<10K0 likes10 downloads1y agoHugging Face15justinsunqiu /transcription_changesimagen<1K0 likes10 downloads1y agoHugging Face16jdabello /yt_transcriptionsimage10K<n<100K0 likes9 downloads3y agoHugging Face17justinsunqiu /multilingual_transcriptions_summarized_by_native_nonnativeimage1K<n<10K0 likes9 downloads1y agoHugging Face18justinsunqiu /multilingual_transcriptions_summarizedimage1K<n<10K0 likes8 downloads1y agoHugging Face19justinsunqiu /multilingual_transcriptions_cleanedimage1K<n<10K0 likes8 downloads1y agoHugging Face20justinsunqiu /multilingual_transcriptions_summarized_by_type_finalimage1K<n<10K0 likes6 downloads1y agoHugging Face21justinsunqiu /multilingual_transcriptions_rawimage1K<n<10K0 likes5 downloads2y agoHugging Face22justinsunqiu /transcription_changes_classifiedimagen<1K0 likes5 downloads1y agoHugging Face23justinsunqiu /multilingual_transcriptionsimage1K<n<10K0 likes4 downloads2y agoHugging Face24curiousmrk /transcription-coding-wiki-500kgated Transcription Dataset: Code & Wiki (390K) Text-to-image rendered dataset for training vision-language models to read code and text from images. Schema Column Type Description image Image Rendered grayscale JPEG prompt string Transcription instruction (varied) response string Ground truth text language string python/javascript/java/c++/rust/go/english domain string code or english length_bucket string short/medium/long/gundam resolution string… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.image100K<n<1M0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.