Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ev-tlt /MACE_finetuning_supplementary MACE Fine-Tuning Supplementary Supplementary data and scripts for: Tompa, T. L.; Varga-Umbrich, E.; Batatia, I.; Elena, A. M.; Bernstein, N.; Csányi, G. Fine-tuning MLIP foundation models: strategies for accuracy and transferability (2026). arXiv:2606.12704. The repository contains training datasets, mace_run_train launch scripts, Slurm logs, fine-tuned model checkpoints (.model), evaluation scripts, and processed results for the paper. Benchmark systems: lithium argyrodite… See the full description on the dataset page: https://huggingface.co/datasets/ev-tlt/MACE_finetuning_supplementary.0 likes753 downloads27d agoHugging Face02shiqiao123 /Muon-MACE-data Muon-MACE: selection inputs, benchmark results and figure data Data companion to Muon-MACE, a MACE research implementation with hybrid Muon–Adam optimization. This repository contains 52 individually accessible data files (220,102,043 bytes). Browse a table, download an array or clone the directory tree; no ZIP extraction is needed. Code, configurations and plotting programs live on GitHub. 中文指南 · File catalog · Checksums and download mapping · Data dictionary Find… See the full description on the dataset page: https://huggingface.co/datasets/shiqiao123/Muon-MACE-data.tabular1K<n<10K0 likes165 downloads1mo agoHugging Face03LVSTCK /macedonian-llm-eval Macedonian LLM Eval This repository is adapted from the original work by Aleksa Gordić. If you find this work useful, please consider citing or acknowledging the original source. You can find the Macedonian LLM eval on GitHub. To run evaluation just follow the guide. Info: You can run the evaluation for Serbian and Slovenian as well, just swap Macedonian with either one of them. What is currently covered: Common sense reasoning: Hellaswag, Winogrande, PIQA… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-llm-eval.1 likes149 downloads1y agoHugging Face04mteb /MacedonianTweetSentimentClassification MacedonianTweetSentimentClassification An MTEB dataset Massive Text Embedding Benchmark An Macedonian dataset for tweet sentiment classification. Task category t2c Domains Social, Written Reference https://aclanthology.org/R15-1034/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MacedonianTweetSentimentClassification"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MacedonianTweetSentimentClassification.texttext-classification1K<n<10K0 likes140 downloads1y agoHugging Face05GD-ML /MACE-Dance 🎵 MACE-Dance Dataset MACE-Dance is a large-scale dataset for music-driven dance video generation, released with our SIGGRAPH 2026 paper: MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation It is designed to support research on generating dance videos that are both: 🕺 kinematically plausible 🎨 visually coherent 🎼 well aligned with music ✨ Overview The dataset contains approximately: 70K dance video clips 5–10 seconds per… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MACE-Dance.text10K<n<100K2 likes115 downloads6mo agoHugging Face06fzzh /MACE MACE Dataset Dataset for paper "Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers". textquestion-answering10K<n<100K0 likes81 downloads8mo agoHugging Face07LVSTCK /macedonian-corpus-raw Macedonian Corpus - Raw 🌟 Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.text10M<n<100M0 likes77 downloads1y agoHugging Face08jharrymoore /maceoff_r6_l1_128chtext100K<n<1M0 likes73 downloads2y agoHugging Face09LVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes56 downloads1y agoHugging Face10shunyalabs /macedonian-speech-datasetaudio1K<n<10K0 likes55 downloads1y agoHugging Face11ryannnchang /macexo0 likes51 downloads11mo agoHugging Face12lucaporte /record-maceo24This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 4, "total_frames": 2982, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo24.tabularrobotics1K<n<10K0 likes49 downloads1y agoHugging Face13LVSTCK /macedonian-corpus-cleaned Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.texttext-generation1M<n<10M0 likes44 downloads1y agoHugging Face14macerj /Forensic_Toolkit_DatasetForensic Toolkit Dataset Overview The Forensic Toolkit Dataset is a comprehensive collection of 300 digital forensics and incident response (DFIR) tools, designed for training AI models, supporting forensic investigations, and enhancing cybersecurity workflows. The dataset includes both mainstream and unconventional tools, covering disk imaging, memory analysis, network forensics, mobile forensics, cloud forensics, blockchain analysis, and AI-driven forensic techniques. Each entry provides… See the full description on the dataset page: https://huggingface.co/datasets/macerj/Forensic_Toolkit_Dataset.text-generationn<1K0 likes41 downloads29d agoHugging Face15lucaporte /record-maceo23This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 1487, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo23.tabularrobotics1K<n<10K0 likes39 downloads1y agoHugging Face16saillab /alpaca_macedonian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_macedonian_taco.text10K<n<100K0 likes33 downloads2y agoHugging Face17lucaporte /record-maceo22This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 1280, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo22.tabularrobotics1K<n<10K0 likes33 downloads1y agoHugging Face18lucaporte /record-maceo-goodThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 15, "total_frames": 8836, "total_tasks": 1, "total_videos": 15, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:15" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo-good.tabularrobotics1K<n<10K0 likes31 downloads1y agoHugging Face19lucaporte /record-maceo11This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 2330, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo11.tabularrobotics1K<n<10K0 likes28 downloads1y agoHugging Face20lucaporte /record-maceo6This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 5, "total_frames": 2674, "total_tasks": 1, "total_videos": 5, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo6.tabularrobotics1K<n<10K0 likes27 downloads1y agoHugging Face21lucaporte /record-maceo10This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 6, "total_frames": 4839, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:6" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo10.tabularrobotics1K<n<10K0 likes27 downloads1y agoHugging Face22lucaporte /record-maceo20This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 932, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo20.tabularroboticsn<1K0 likes26 downloads1y agoHugging Face23lucaporte /record-maceo7This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 3, "total_frames": 2280, "total_tasks": 1, "total_videos": 3, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucaporte/record-maceo7.tabularrobotics1K<n<10K0 likes25 downloads1y agoHugging Face24isaacchung /macedonian-tweet-sentiment-classification Dataset Card for Macedonian Tweet Sentiment Classification Dataset Description Dataset Summary This is a Macedonian dataset is a collection of tweets for sentiment classification. Supported Tasks and Leaderboards text-classification, sentiment-classification: The dataset can be used to train a model for sentiment classification. The model performance is evaluated based on the accuracy of the predicted labels as compared to the given labels in the… See the full description on the dataset page: https://huggingface.co/datasets/isaacchung/macedonian-tweet-sentiment-classification.texttext-classification1K<n<10K1 likes21 downloads2y agoHugging Face25AngstromAI /maceoff_24_preprocessed0 likes19 downloads2y agoHugging Face26paulktothicloudcom /mac-egpu-llm-stack Mac Mini M4 Pro + RX 7900 XTX Local LLM Stack A fully automated setup for running a four-component local LLM inference stack on a Mac Mini M4 Pro with an AMD RX 7900 XTX eGPU via TinyGPU. All inference runs locally — no cloud API keys, no telemetry, no code leaves the machine. Hardware Target Component Specification Host Apple Mac Mini M4 Pro Unified Memory 64 GB (hard minimum — see Memory Requirements) eGPU SAPPHIRE NITRO+ RX 7900 XTX VAPOR-X 24 GB… See the full description on the dataset page: https://huggingface.co/datasets/paulktothicloudcom/mac-egpu-llm-stack.n<1K1 likes19 downloads4mo agoHugging Face27DGurgurov /macedonian_sa Sentiment Analysis Data for the Macedonian Language Dataset Description: This dataset contains a sentiment analysis dataset from Jovanski et al. (2015). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @inproceedings{jovanoski-etal-2015-sentiment, title = "Sentiment Analysis in {T}witter for {M}acedonian", author = "Jovanoski, Dame and Pachovski, Veno and Nakov, Preslav"… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/macedonian_sa.texttext-classification1K<n<10K0 likes17 downloads2y agoHugging Face28saillab /alpaca-macedonian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-macedonian-cleaned.text10K<n<100K0 likes17 downloads2y agoHugging Face29harrytq /MACE 🎵 MACE-Dance Dataset MACE-Dance is a large-scale dataset for music-driven dance video generation, released with our SIGGRAPH 2026 paper: MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation It is designed to support research on generating dance videos that are both: 🕺 kinematically plausible 🎨 visually coherent 🎼 well aligned with music ✨ Overview The dataset contains approximately: 70K dance video clips 5–10 seconds per… See the full description on the dataset page: https://huggingface.co/datasets/harrytq/MACE.10K<n<100K0 likes16 downloads5mo agoHugging Face30Speech-data /Macedonian-Speech-Dataset 🎧 Macedonian Speech Dataset The Macedonian Speech Dataset is a high-quality speech audio dataset designed to provide structured and reliable audio data for AI and machine learning systems. It includes 128 hours of audio data distributed across 654 files, delivered in MP3 and WAV formats, with a total size of 279 MB. This carefully curated audio dataset ensures diverse and representative voice data, with 55% female and 45% male speakers, and an age range spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Macedonian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.