Team Ai
20 results

language

IPEC-COMMUNITY /language_table_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "xarm", "total_episodes": 442226, "total_frames": 7045476, "total_tasks": 127605, "total_videos": 442226, "total_chunks": 443, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:442226" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot.robotics29 likes118k downloads2y agoHugging FaceLanguageBind /Open-Sora-Plan-v1.1.0 Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.text100K<n<1M46 likes64k downloads2y agoHugging FacesaaduddinM /rbo_oxe_base_language_table_lerobot Language Table (LeRobot) — Task-Pruned, Reindexed Subset This release is a task-pruned subset of the original IPEC-COMMUNITY/language_table_lerobot. We subsampled by task text and rebuilt the package so it remains internally consistent (indices, splits, stats, paths). Robot: xArm Modality: RGB video + states + actions FPS / Resolution: 10 FPS, 360×640, AV1 License: apache-2.0 (inherits from source) What’s different in this subset Kept ~0.85% of unique… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/rbo_oxe_base_language_table_lerobot.robotics0 likes22k downloads1y agoHugging FaceCohereLabs /aya_collection_language_split This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.tabular100M<n<1B122 likes14k downloads1y agoHugging Facelivebench /language Dataset Card for "livebench/language" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/language.textn<1K1 likes9.7k downloads2y agoHugging Facehasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes8.2k downloads10mo agoHugging Face