datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda
ATX Swift 1.5 Qwen3.8-27B Uncensored 2.3 bpw — LiveCodeBench and Pagoda tests
Test package for jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF, a native 2.3 bpw re-encode of Swift 1.5 (with its MTP head) built for 12 GB cards at about 200K context. It holds three tests:
LiveCodeBench v6 on rented 12 GB cards (RTX 3060 12 GB, RTX 3080 12 GB, RTX 3080 Ti): 7 complete runs.
Pagoda: a voxel-pagoda coding-agent prompt run against this model and three other ~2-bit 27B… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda.livecodebench-plus
LiveCodeBench-v6-Plus
A curated coding benchmark of 91 problems selected by hardness/discrimination
(lcb-v6-plus). It combines two sources, all in one clean schema:
64 evolved problems — mutated/evolved variants from LiveCodeBench-v6
(each carries its seed_problem).
27 original problems — un-evolved AtCoder problems taken directly from
livecodebench/code_generation_lite
release v6 (seed_problem is null).
About BenchEvolver
The evolved problems were produced by… See the full description on the dataset page: https://huggingface.co/datasets/BenchEvolver/livecodebench-plus.LiveCodeBench-CodeGenerationlivecodebench-merging-leaderboard
LiveCodeBench v6 Evaluation Leaderboard
Evaluation results for cross-capability merging of OLMo-3 and OLMo-3.1 RL-Zero models on 454 coding problems.
Evaluation
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
Leaderboard
Model
pass@4
pass@1
Loop Rate
Qwen/Qwen3-4B-Thinking-2507
54.6%
45.4%
0.4%
pmahdavi/Olmo-3-7B-Think-Math-Code… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/livecodebench-merging-leaderboard.livecodebench
LiveCodeBench for Code-LLaVA
This dataset contains the LiveCodeBench code generation benchmark prepared for
Code-LLaVA evaluation.
Source
Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Repository: https://github.com/LiveCodeBench/LiveCodeBench
Version: release_v6 (May 2023 - Apr 2025, 1055 problems)
Dataset Structure
Two configurations are available:
memwrap: Problems with <|memory_start|> /… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/livecodebench.LiveCodeBench-v6-R182
LiveCodeBench-v6-R182
The 182 problems obtained by taking the release_v6 slice of livecodebench/code_generation_lite and keeping only those with contest_date >= 2025-01-01 (contest dates span 2025-01-04 to 2025-04-06).
Usage
from datasets import load_dataset
ds = load_dataset("jwu323/LiveCodeBench-v6-R182", split="test")
print(ds[0]["question_title"], ds[0]["contest_date"])
