datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.quantized-retrieval-dataMNLP_M3_quantized_datasetMNLP_M3_quantized_datasetToneWebinars_quantized-bigcodecquantized-llama-3.1-humaneval-evals
Coding Benchmark Results
The coding benchmark results were obtained with the EvalPlus library.
HumanEvalpass@1
HumanEval+pass@1
meta-llama_Meta-Llama-3.1-405B-Instruct
67.3
67.5
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-FP8
66.7
66.6
neuralmagic_Meta-Llama-3.1-405B-Instruct-W4A16
66.5
66.4
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-INT8
64.3
64.8
neuralmagic_Meta-Llama-3.1-70B-Instruct-W8A8-FP8
58.1
57.7
neuralmagic_Meta-Llama-3.1-70B-Instruct-W4A16
57.1… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-humaneval-evals.MNLP_M2_quantized_dataset1quantized_burgers_vqMNLP_M3_quantized_datasetemilia_en_quantized-wav-unifyMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetMNLP_M2_quantized_dataset
Dataset Card for MNLP_M2_sft_dataset
Dataset Description
A unified STEM instruction-following dataset comprising 240,500 examples drawn from six existing benchmarks: SciQ, Deepmind Code Contests, TIGER-Lab MathInstruct, TULU Algebra, TULU Code, and Facebook Natural Reasoning. Each example is formatted as a chat-style message pair for supervised fine-tuning of instruction-following models.
Curated by: Sarra Chabane
Shared by: GingerBled (https://huggingface.co/GingerBled)… See the full description on the dataset page: https://huggingface.co/datasets/arthurrpp/MNLP_M2_quantized_dataset.quantized-llama-3.1-arena-hard-evals
Arena-Hard Benchmark Results
This artifact contains all the data neccessary to reproduce the results of the Arena-Hard benchmark for Neural Magic's quantized Llama-3.1 models.
The model_answers directory includes the generated answers from all models, and the model_judgements directory contains the evaluations by gpt-4-1106-preview.
The Arena-Hard version used for benchmarking is v0.1.0, corresponding to commit efc012e192b88024a5203f5a28ec8fc0342946df.
All model answers were… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-arena-hard-evals.MNLP_M2_quantized_datasetMNLP_M2_quantized_datasetgiant-midi-quantized
Dataset Card for "giant-midi-quantized"
More Information needed
MNLP_M3_quantized_datasetMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetMNLP_M3_quantized_dataset3MNLP_M3_quantized_datasetToneSlavic_quantized-bigcodecMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetmaestro-sustain-quantized
Dataset Card for "maestro-sustain-quantized"
More Information needed
MNLP_M3_quantized_datasetfiltered_common_voice_tamil_english-preprocessed-quantizedemilia_multilang_quantized-wav-uni
