Team Ai
21 results

bit

BitRobot /HIW-500 HIW-500: Humanoids In-the-Wild Dataset https://bitrobot-foundation.github.io/humanoids-in-the-wild-500-hours/ HIW-500: Humanoids In-the-Wild Dataset is a large-scale dataset for whole-body humanoid robot learning in natural home environments. It captures human teleoperation demonstrations on Unitree G1 across real homes in Southeast Asia, where layouts, object states, lighting, clutter, and operator styles vary from episode to episode. The dataset is designed for research on… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/HIW-500.53 likes149k downloads12d agoHugging FaceBitRobot /HIW-500-LeRobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.images.head": { "dtype": "video", "shape": [ 480, 1280, 3 ], "names": [ "height", "width", "channels" ], "info": { "video.height":… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/HIW-500-LeRobot.tabularrobotics10M<n<100M24 likes22k downloads3mo agoHugging Faceihavespoons /bite-baseline bite-baseline — artifacts for extreme (ternary) quantization of Qwen3.6-35B-A3B Companion dataset for ihavespoons/bite — an open pipeline for compressing a Mixture-of-Experts LLM (Qwen/Qwen3.6-35B-A3B, 35B total / ~3B active, 256 experts) toward ternary {-1,0,+1} weights (1.71 bpw) via PTQ init + quantization-aware distillation. See the repo's docs/report-extreme-quant-moe.md for the full technical report. Contents Path What it is baseline.json… See the full description on the dataset page: https://huggingface.co/datasets/ihavespoons/bite-baseline.0 likes21k downloads2mo agoHugging Faceloose-bits /uq-hiddenstates uq-hiddenstates — residual-stream states of reasoning traces at a fixed depth Every token position of a reasoning trace, recorded at relative model depth 0.75, with correctness labels. Built for studying whether uncertainty is legible in the residual stream while the model reasons, rather than only at the answer. Layout Per model and dataset: {model}_{ds}_L{idx}.part{k}.npy — fp16 [rows, hidden], traces concatenated, raw states, not normalized. Sharded at ~20 GB… See the full description on the dataset page: https://huggingface.co/datasets/loose-bits/uq-hiddenstates.question-answering0 likes11k downloads20d agoHugging Facebitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K199 likes11k downloads2y agoHugging FaceBitAgent /tool_callingtext100K<n<1M12 likes4.4k downloads2y agoHugging Face