Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hosseinbv /dim58-cpuData-31cases Dim58 CPU Data — 31 Cases Dataset uploaded from: /mnt/data/ubuntu/research/outputs/data_cpu_geodesic58 Dataset summary Property Value Repository hosseinbv/dim58-cpuData-31cases Number of files 64 Total size 17.87 GB Source folder data_cpu_geodesic58 File types Extension File count .npz 62 .json 1 .csv 1 Top-level contents 0000_internal_case1_data.npz 0001_internal_B_10.npz… See the full description on the dataset page: https://huggingface.co/datasets/hosseinbv/dim58-cpuData-31cases.tabularn<1K0 likes715 downloads3mo agoHugging Face02AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes562 downloads25d agoHugging Face03malaiwah /glm-moe-dsa-tiny-cpu-repro-v1 Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay Reproducibility evidence for malaiwah/glm-moe-dsa-tiny-random-bf16, checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37. This is a synthetic pipeline test, not a quality benchmark, quantization measurement, qualified production reference, or registry submission. The model is random-init. No GPU or paid cloud job was used. Observed result Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.tabularn<1K0 likes431 downloads1mo agoHugging Face04malaiwah /kimi-k3-tiny-cpu-repro-v1 Kimi K3 complete tiny random BF16 CPU fixture Untrained independently seeded random weights; no upstream weights or training data. This is a reproducibility fixture, not useful language modeling or production quality evidence. Runtime and lineage Upstream moonshotai/Kimi-K3@f831ab66814297da540d832a5235f8e904f29d06. Actual loaded class: KimiK3ForConditionalGeneration. Complete untied head and real small vision tower/projector. Parameters: 269688; vision parameters:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k3-tiny-cpu-repro-v1.tabularn<1K0 likes382 downloads1mo agoHugging Face05malaiwah /kimi-k25-tiny-cpu-repro-v1 Kimi K2.5 / K2.6 / K2.7-Code complete tiny random BF16 CPU fixture Untrained independently seeded random weights; no upstream weights or training data. This is a reproducibility fixture, not useful language modeling or production quality evidence. Runtime and lineage Upstream moonshotai/Kimi-K2.7-Code@74797c9c62378b951a1f6fcf5c4631024e9b8bef. Actual loaded class: Kimi_K25ForConditionalGeneration. Complete untied head and real small vision tower/projector.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-cpu-repro-v1.tabularn<1K0 likes372 downloads1mo agoHugging Face06AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes342 downloads2mo agoHugging Face07malaiwah /k2-horizon-tiny-cpu-repro-v1 K2-Horizon MoVA tiny random CPU fixture Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name. Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa. No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used. Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes314 downloads1mo agoHugging Face08malaiwah /deepseek-v4-tiny-cpu-repro-v1 DeepSeek-V4 tiny corrected-native-primitives CPU text fixture Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction. No upstream weights, paid GPU/cloud compute or useful-model claim. This is not unmodified native Transformers or the complete production release. Architecture and scope Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes278 downloads1mo agoHugging Face09malaiwah /glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root. first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset. GLM5-Next tiny native CPU fixture This is a complete untrained random-initialized native Glm5NextForConditionalGeneration wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes253 downloads1mo agoHugging Face10malaiwah /qwen3-5-tiny-cpu-repro-v1 Qwen3.5 tiny native random CPU fixture Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint. This is a pipeline/reproducibility fixture, not a useful language model, distillation, quantization, quality benchmark, or claim about the performance of Qwen3.8-27B. No upstream model weights or training data were used. No paid GPU/cloud compute. Architecture and lineage Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes215 downloads1mo agoHugging Face11malaiwah /minimax-m2-tiny-cpu-repro-v1 minimax-m2 complete native tiny random CPU fixture Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b. No upstream weights, training data, paid GPU or cloud compute were used. Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes210 downloads1mo agoHugging Face12malaiwah /spark2-5-tiny-cpu-repro-v1 Spark2.5 tiny random CPU fixture Complete untrained Spark2_5ForCausalLM with independently seeded random BF16 weights. This is a reproducibility fixture, not a useful language model, distilled model, quality benchmark, or production registry measurement. No upstream weights, training data, paid GPU or cloud rentals were used. Architecture, code and license Source: XHToken/Spark-X2.5-4B at 5e10fcc0286756aebf7c41dc52c1e42d95c70281. The complete text causal model… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-cpu-repro-v1.tabularn<1K0 likes202 downloads1mo agoHugging Face13malaiwah /minimax-m3-tiny-cpu-repro-v1 minimax-m3 complete native tiny random CPU fixture Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0. No upstream weights, training data, paid GPU or cloud compute were used. Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes201 downloads1mo agoHugging Face14malaiwah /qwen4-exp-tiny-cpu-repro-v1 Qwen4-Exp tiny CPU reproduction receipts This is an artifact/receipt bundle, not training data and not a single root-format QFS dataset. All four readable synthetic documents are embedded in panel/panel.receipt.json. Provenance and limitations This is an independently generated, untrained random checkpoint inspired by Qwen/Qwen3.8-Flash-Next@de4b8e4d43b917e7706784d8bb445c9af86a3540, not a quantization, distillation, behavioral replica, or fine-tune. No source… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen4-exp-tiny-cpu-repro-v1.tabularn<1K0 likes119 downloads1mo agoHugging Face15malaiwah /qfs-existing-tiny-cpu-format-v1 existing tiny CPU format fixture reproducibility Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim. Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-existing-tiny-cpu-format-v1.tabularn<1K0 likes114 downloads1mo agoHugging Face16oscarcole03 /clear_table_pi_cpu_test_20261001_152121This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/oscarcole03/clear_table_pi_cpu_test_20261001_152121.tabularroboticsn<1K0 likes114 downloads10d agoHugging Face17oscarcole03 /clear_table_pi_cpu_test_20261001_153851This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/oscarcole03/clear_table_pi_cpu_test_20261001_153851.tabularroboticsn<1K0 likes110 downloads10d agoHugging Face18malaiwah /qfs-affine-tiny-cpu-format-v1 affine tiny CPU format fixture reproducibility Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim. Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training quality.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-affine-tiny-cpu-format-v1.tabularn<1K0 likes105 downloads1mo agoHugging Face19sauravsingla08 /MemVanta-CPU-LLM-Memory-Benchmark MemVanta CPU LLM Memory Benchmark Reproducible CPU LLM inference benchmark evidence comparing memory usage and throughput for MemVanta and a pinned comparison runtime on the same GGUF artifact. The initial configuration is the committed OpenLLaMA 7B v2 Q4_0 same-model CPU A/B benchmark from the public MemVanta repository. This dataset publishes benchmark evidence and provenance only; it does not redistribute model weights. Canonical result Metric MemVanta… See the full description on the dataset page: https://huggingface.co/datasets/sauravsingla08/MemVanta-CPU-LLM-Memory-Benchmark.tabularn<1K0 likes53 downloads13d agoHugging Face20radna0 /harmony-nemotron-cpu-artifacts Harmony CPU artifacts: Nemotron datasets (normalized + candidate pools) This dataset repo is an artifact store produced on an EPYC CPU box. It contains: normalized/ — CPU-normalized Parquet shards with a text-first Harmony format (text) plus meta_* and quality_* fields. pools/ — candidate pool Parquet shards (subsets) for later GPU scoring (Modal NLL/PPL). No GPU scoring has been run yet. reports/ — summary tables of counts per dataset/split/pool. Directory layout… See the full description on the dataset page: https://huggingface.co/datasets/radna0/harmony-nemotron-cpu-artifacts.tabular10M<n<100M0 likes48 downloads9mo agoHugging Face21malaiwah /qfs-microfloat-tiny-cpu-format-v1 microfloat tiny CPU format fixture reproducibility Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim. Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-microfloat-tiny-cpu-format-v1.tabularn<1K0 likes48 downloads1mo agoHugging Face22malaiwah /qfs-qwen-gguf-tiny-cpu-format-v1 qwen-gguf tiny CPU format fixture reproducibility Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim. Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-qwen-gguf-tiny-cpu-format-v1.tabularn<1K0 likes40 downloads1mo agoHugging Face23imstevenpmwork /final_test_stream_encoding_linux_midres_cpuThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 4, "total_frames": 3567, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_linux_midres_cpu.tabularrobotics1K<n<10K0 likes39 downloads8mo agoHugging Face24AdityaMayukhSom /MixSub-LLaMA-3.2-Text-Only-Overlap-CPU-Scoretabular1K<n<10K0 likes36 downloads2y agoHugging Face25imstevenpmwork /final_test_stream_encoding_linux_highres_cpuThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 4, "total_frames": 3529, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_linux_highres_cpu.tabularrobotics1K<n<10K0 likes34 downloads8mo agoHugging Face26imstevenpmwork /final_test_stream_encoding_mac_midhres_cpuThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 4, "total_frames": 3596, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_mac_midhres_cpu.tabularrobotics1K<n<10K0 likes30 downloads8mo agoHugging Face27Ahmef810211 /intel-cpu-datasettabular1K<n<10K0 likes25 downloads5d agoHugging Face28thewisp /skewer_luncheon_jan_28_experiment_cpuThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so107_follower", "total_episodes": 1, "total_frames": 820, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/thewisp/skewer_luncheon_jan_28_experiment_cpu.tabularroboticsn<1K0 likes23 downloads8mo agoHugging Face29ICOS-AI /scaphandre_cpu_usage Scaphandre CPU Usage Dataset Dataset Description This dataset contains CPU usage monitoring data collected using Scaphandre for the ICOS Federated Learning infrastructure. Overview Source: Scaphandre energy monitoring tool Collection Method: Live system monitoring (continuous fetching) Purpose: Training data for ICOS FL Update Frequency: Real-time collection with 3s intervals Data Schema Column Type Description timestamp float Unix… See the full description on the dataset page: https://huggingface.co/datasets/ICOS-AI/scaphandre_cpu_usage.tabulartime-series-forecastingn<1K0 likes22 downloads1y agoHugging Face30imstevenpmwork /final_test_stream_encoding_mac_highres_cpuThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 4, "total_frames": 3524, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_mac_highres_cpu.tabularrobotics1K<n<10K0 likes22 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.