Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PS4Research /marimo-workshop-catalogue marimo workshop product catalogue A small, ready-to-search product catalogue for a hands-on marimo workshop, where students build a multimodal (text and photo) product search engine that runs on a CPU. File What it is catalogue.csv 3,000 products, 150 in each of 20 categories images.zip images/<product_id>.jpg, 240×320 JPEG image_vectors.npy (3000, 512) float32 CLIP image vectors, one per catalogue row, each of length 1 products.csv a random 500-row subset used… See the full description on the dataset page: https://huggingface.co/datasets/PS4Research/marimo-workshop-catalogue.image1K<n<10K0 likes86 downloads24d agoHugging Face02asiom /cyp3a4_marimotabular1K<n<10K0 likes61 downloads8d agoHugging Face03uv-scripts /marimo Marimo UV Scripts Marimo notebooks that work as both interactive tutorials and batch scripts. What is this? Marimo notebooks are pure Python files that can be: Edited interactively with a reactive notebook interface Run as scripts with uv run - same as any UV script This makes them perfect for tutorials and educational content where you want users to explore step-by-step, but also run the whole thing as a batch job. Available Scripts Script Description… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/marimo.0 likes30 downloads8mo agoHugging Face04mari-lab /mari-monolingual-corpusA monolingual corpus of the Mari language in various genres, containing over 20 million word occurrences. The presented genres: Genre Russian English мутер словарь dictionary газетысе увер газетные новости periodical news прозо проза prose фольклор фольклор folklore публицистике публицистика publicistic literature поэзий поэзия poetry трагикомедийтрагикомедия tragicomedy пьесе пьеса play драме драма drama комедий-водевиль водевиль vaudeville комедий комедия… See the full description on the dataset page: https://huggingface.co/datasets/mari-lab/mari-monolingual-corpus.tabular1M<n<10M4 likes28 downloads3y agoHugging Face05Marimonald /ojisan-translation-corpus-ja Ojisan Translation Corpus (Japanese) 日本語のドキュメントはこちら The Ojisan Translation Corpus (ojisan-translation-corpus-ja) is a curated, high-quality Japanese dataset designed for style-transfer and preference alignment into the culturally unique "Ojisan-dialect" (おじさん構文 / Ojisan-koubun). This dataset powers the fine-tuning of Ojisan-translator-v1-Qwen-3.5-9B, providing structured subsets for Supervised Fine-Tuning (SFT) and Kahneman-Tversky Optimization (KTO). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Marimonald/ojisan-translation-corpus-ja.text-generation1K<n<10K0 likes22 downloads1mo agoHugging Face06goldenfox /Marimo-thought0 likes11 downloads1y agoHugging Face07kgdrathan /marimo-manim-sft-datatextn<1K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.