datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wan2.2-I2V-Activations-INT4LTX-Video-Activations-INT4all-Meta-Llama-3.1-70B-Instruct-AWQ-INT4eval_act_int4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 15,
"total_frames": 9687,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kivod/eval_act_int4.details_qualis2006__llama-2-7b-int4-python-code-18k
Dataset Card for Evaluation run of qualis2006/llama-2-7b-int4-python-code-18k
Dataset Summary
Dataset automatically created during the evaluation run of model qualis2006/llama-2-7b-int4-python-code-18k on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_qualis2006__llama-2-7b-int4-python-code-18k.Wan2.2-T2V-Activations-INT4Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
glm53-fixture-0.1B-fidelity-quant-int4-v1
GLM-5.3-Flash-0.1B fixture — candidate fidelity dataset, toy RTN-int4 routed experts (hidden form)
The numbers in this dataset are meaningless as quantization quality.
The weights are random (inference-optimization/GLM-5.3-Flash-0.1B-A0.1B is an
architectural fixture), and the quantizer is deliberately crude. This exists so
that step 3 of the three-step fidelity architecture has two real datasets to
compare, and so that anyone can see what a candidate capture looks like
next to… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fixture-0.1B-fidelity-quant-int4-v1.INT4
Also just a few numbers ¯_(ツ)_/¯
eval-qwen2_5-coder-32b-instruct-gptq-int4-post_books_20254090-svd-int4eval-qwen2_5-coder-32b-instruct-gptq-int4-post_books_2025_v2qwen72b_gptq_int4-rstp_bilingual_rstpdahiliye_int4_chunkseval-qwen2_5-coder-32b-instruct-gptq-int4-post_llama31qwen2_5_gptq_int4_step168plover-qa-extractions-qwen35-9b-int4-autoround-intelgemma3-4b-it-int4.taskgemma-2b-int4
