datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.sakhan10__quantized_open_llama_3b_v2-details
Dataset Card for Evaluation run of sakhan10/quantized_open_llama_3b_v2
Dataset automatically created during the evaluation run of model sakhan10/quantized_open_llama_3b_v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sakhan10__quantized_open_llama_3b_v2-details.bitandbytes_quantized
/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350 · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-24 06:13:30Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/bitandbytes_quantized.fp8_quantized
/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350-fp8 · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350-fp8
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-24 05:49:28Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/fp8_quantized.
