Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k19862217 /simpsons_script_linestext10K<n<100K0 likes1.1k downloads3y agoHugging Face02Lines /Open-Domain-Oral-Disease-QA-Dataset Open-Domain-Oral-Disease-QA-Dataset Dataset Details Dataset Description This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in the domain of oral disease. We currently offer a suite of evaluation datasets encompassing models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B. More data is under reviewed. This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/Lines/Open-Domain-Oral-Disease-QA-Dataset.textn<1K5 likes47 downloads2y agoHugging Face03lxe /receipt-household-lines-odbl Receipt Household & Grocery Lines (ODbL) Synthetic US receipt line items built from real products in the Open Food Facts family of databases, each with a POS-style printed label, a readable product name and one of 39 spending categories (taxonomy v2, taxonomy.json). This is the share-alike part of the Receipt Kitten item-namer training data. The permissive part (real receipts, USDA-based synthetic lines, teacher-generated household inventory) is lxe/receipt-line-items-taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/lxe/receipt-household-lines-odbl.texttext-classification10K<n<100K0 likes29 downloads12d agoHugging Face04TeraSpace /glados_ru_linesRussian glados lines with links to audio from https://i1.theportalwiki.net/ parsed from https://theportalwiki.com/wiki/GLaDOS_voice_lines/ru textn<1K4 likes27 downloads3y agoHugging Face05AlhitawiMohammed22 /lines_hu_v5image10K<n<100K0 likes17 downloads3y agoHugging Face06prashant0919 /nepali-synthetic-ocr-lines Nepali Synthetic OCR/HTR Document Line Dataset A dataset of synthetic Devanagari text line images imitating historical and official scanned document conditions, designed for OCR (Optical Character Recognition) and HTR (Handwritten Text Recognition) models such as TrOCR, CRNN, and PaddleOCR. This dataset was generated using the Mountmind PeakOCR Studio synthetic corpus generator pipeline, introducing realistic document aging artifacts like: Skew Angle Rotations (Hough line… See the full description on the dataset page: https://huggingface.co/datasets/prashant0919/nepali-synthetic-ocr-lines.imageimage-to-text1K<n<10K1 likes17 downloads4mo agoHugging Face07ray0rf1re /ada-linestext1K<n<10K0 likes16 downloads3mo agoHugging Face08jingxize /Classic-lines-from-the-movie-Nezhatextn<1K0 likes6 downloads2y agoHugging Face09perkros /netlist-snippets-40-linestext10K<n<100K0 likes5 downloads2y agoHugging Face10perkros /netlist-snippets-80-linestext10K<n<100K0 likes5 downloads2y agoHugging Face11semran1 /labeled_lines2_1mtabular100K<n<1M0 likes4 downloads11mo agoHugging Face12perkros /netlist-snippets-20-linestext10K<n<100K0 likes3 downloads2y agoHugging Face13semran1 /labeled_lines_datatabular100K<n<1M0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.