datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LiveCodeBench-Proexecution-v2livecodebench-merging-leaderboard
LiveCodeBench v6 Evaluation Leaderboard
Evaluation results for cross-capability merging of OLMo-3 and OLMo-3.1 RL-Zero models on 454 coding problems.
Evaluation
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
Leaderboard
Model
pass@4
pass@1
Loop Rate
Qwen/Qwen3-4B-Thinking-2507
54.6%
45.4%
0.4%
pmahdavi/Olmo-3-7B-Think-Math-Code… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/livecodebench-merging-leaderboard.Unroll-Qwen2.5-7B-Instruct_1754646934_eval_6419_livecodebench_skip_ffn_idx_8_v2
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754646934_eval_6419_livecodebench_skip_ffn_idx_8_v2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 20.38% ± 0.49%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
21.33%
109
511
2
20.35%
104
511
3
20.55%
105
511
4
18.59%
95
511
5
21.92%
112
511
6
19.57%
100
511
Custom-OpenThinker-32B_1754028628_eval_07bc_livecodebench_prn_attn_1
chengfu0118/Custom-OpenThinker-32B_1754028628_eval_07bc_livecodebench_prn_attn_1
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 69.11% ± 0.45%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
70.84%
362
511
2
69.08%
353
511
3
69.28%
354
511
4
67.51%
345
511
5
68.49%
350
511
6
69.47%
355
511
LiveCodeBench_evalchemyCustom-OpenThinker-32B_1754028672_eval_07bc_livecodebench_prn_attn_2
chengfu0118/Custom-OpenThinker-32B_1754028672_eval_07bc_livecodebench_prn_attn_2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 25.15% ± 0.69%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
27.20%
139
511
2
24.27%
124
511
3
26.61%
136
511
4
26.03%
133
511
5
23.48%
120
511
6
23.29%
119
511
Unroll-Qwen2.5-7B-Instruct_1754614568_eval_6419_livecodebench_skip_attn_idx_23
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754614568_eval_6419_livecodebench_skip_attn_idx_23
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 16.47% ± 0.19%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
16.44%
84
511
2
17.22%
88
511
3
16.63%
85
511
4
16.24%
83
511
5
16.44%
84
511
6
15.85%
81
511
Unroll-Qwen2.5-7B-Instruct_1754647690_eval_6419_livecodebench_skip_ffn_idx_26_v2
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754647690_eval_6419_livecodebench_skip_ffn_idx_26_v2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 28.70% ± 0.40%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
28.77%
147
511
2
27.98%
143
511
3
29.16%
149
511
4
29.16%
149
511
5
27.20%
139
511
6
29.94%
153
511
Custom-OpenThinker-32B_1754028845_eval_07bc_livecodebench_prn_attn_6
chengfu0118/Custom-OpenThinker-32B_1754028845_eval_07bc_livecodebench_prn_attn_6
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 0.00% ± 0.00%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
511
2
0.00%
0
511
3
0.00%
0
511
4
0.00%
0
511
5
0.00%
0
511
6
0.00%
0
511
Unroll-Qwen2.5-7B-Instruct_1754647312_eval_6419_livecodebench_skip_ffn_idx_17_v2
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754647312_eval_6419_livecodebench_skip_ffn_idx_17_v2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 20.48% ± 0.55%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
18.79%
96
511
2
21.92%
112
511
3
21.14%
108
511
4
19.57%
100
511
5
21.92%
112
511
6
19.57%
100
511
Custom-OpenThinker-32B_1754028885_eval_07bc_livecodebench_prn_attn_7
chengfu0118/Custom-OpenThinker-32B_1754028885_eval_07bc_livecodebench_prn_attn_7
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 0.00% ± 0.00%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
511
2
0.00%
0
511
3
0.00%
0
511
4
0.00%
0
511
5
0.00%
0
511
6
0.00%
0
511
Custom-OpenThinker-32B_1754028759_eval_07bc_livecodebench_prn_attn_4
chengfu0118/Custom-OpenThinker-32B_1754028759_eval_07bc_livecodebench_prn_attn_4
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 0.78% ± 0.14%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
0.59%
3
511
2
0.59%
3
511
3
0.39%
2
511
4
1.37%
7
511
5
0.78%
4
511
6
0.98%
5
511
Custom-OpenThinker-32B_1754028718_eval_07bc_livecodebench_prn_attn_3
chengfu0118/Custom-OpenThinker-32B_1754028718_eval_07bc_livecodebench_prn_attn_3
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 3.65% ± 0.17%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
4.11%
21
511
2
3.13%
16
511
3
4.11%
21
511
4
3.72%
19
511
5
3.33%
17
511
6
3.52%
18
511
Unroll-Qwen2.5-7B-Instruct_1754924337_eval_6419_livecodebench_geometric_num_prune_ffn_5_run-002
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754924337_eval_6419_livecodebench_geometric_num_prune_ffn_5_run-002
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 7.66% ± 0.28%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
8.41%
43
511
2
8.61%
44
511
3
7.44%
38
511
4
6.85%
35
511
5
7.24%
37
511
6
7.44%
38
511
Unroll-Qwen2.5-7B-Instru_1754928730_eval_6419_livecodebench_top-5-voting_num_prune_ffn_3_run-002
chengfu0118/Unroll-Qwen2.5-7B-Instru_1754928730_eval_6419_livecodebench_top-5-voting_num_prune_ffn_3_run-002
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 19.99% ± 0.70%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
17.81%
91
511
2
20.35%
104
511
3
18.00%
92
511
4
20.94%
107
511
5
22.11%
113
511
6
20.74%
106
511
Unroll-Qwen2.5-7B-Instruct_1754901542_eval_6419_livecodebench_top-5-voting_num_prune_attn_5
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754901542_eval_6419_livecodebench_top-5-voting_num_prune_attn_5
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 21.69% ± 0.36%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
20.35%
104
511
2
21.33%
109
511
3
22.50%
115
511
4
22.31%
114
511
5
21.14%
108
511
6
22.50%
115
511
Unroll-Qwen2.5-7B-Instruct_1754912452_eval_6419_livecodebench_geometric_num_prune_ffn_6_run-002
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754912452_eval_6419_livecodebench_geometric_num_prune_ffn_6_run-002
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 2.90% ± 0.20%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
3.33%
17
511
2
3.33%
17
511
3
3.33%
17
511
4
2.54%
13
511
5
2.74%
14
511
6
2.15%
11
511
Unroll-Qwen2.5-7B-Instruct_1754928557_eval_6419_livecodebench_geometric_num_prune_ffn_3_run-002
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754928557_eval_6419_livecodebench_geometric_num_prune_ffn_3_run-002
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 16.54% ± 0.46%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
15.26%
78
511
2
15.66%
80
511
3
17.03%
87
511
4
18.20%
93
511
5
17.22%
88
511
6
15.85%
81
511
qwq_livecodebench_distilledCustom-OpenThinker-32B_1754028585_eval_07bc_livecodebench_prn_attn_0
chengfu0118/Custom-OpenThinker-32B_1754028585_eval_07bc_livecodebench_prn_attn_0
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 68.59% ± 0.58%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
70.06%
358
511
2
66.93%
342
511
3
70.45%
360
511
4
67.91%
347
511
5
68.69%
351
511
6
67.51%
345
511
Unroll-Qwen2.5-7B-Instruct_1754613783_eval_6419_livecodebench_skip_attn_idx_6
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754613783_eval_6419_livecodebench_skip_attn_idx_6
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 27.82% ± 0.24%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
28.57%
146
511
2
28.18%
144
511
3
27.01%
138
511
4
27.59%
141
511
5
28.18%
144
511
6
27.40%
140
511
Unroll-Qwen2.5-7B-Instruct_1754614119_eval_6419_livecodebench_skip_attn_idx_14
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754614119_eval_6419_livecodebench_skip_attn_idx_14
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 25.54% ± 0.30%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
24.07%
123
511
2
25.83%
132
511
3
26.03%
133
511
4
25.64%
131
511
5
25.64%
131
511
6
26.03%
133
511
Unroll-Qwen2.5-7B-Instruct_1754646835_eval_6419_livecodebench_skip_ffn_idx_5_v2
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754646835_eval_6419_livecodebench_skip_ffn_idx_5_v2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 27.82% ± 0.49%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
25.44%
130
511
2
27.98%
143
511
3
28.18%
144
511
4
28.38%
145
511
5
28.77%
147
511
6
28.18%
144
511
Unroll-Qwen2.5-7B-Instruct_1754687379_eval_6419_livecodebench_geometric_num_prune_attn_6
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754687379_eval_6419_livecodebench_geometric_num_prune_attn_6
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 7.40% ± 0.31%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
7.44%
38
511
2
8.22%
42
511
3
7.05%
36
511
4
7.83%
40
511
5
6.07%
31
511
6
7.83%
40
511
Unroll-Qwen2.5-7B-Instruct_1754896720_eval_6419_livecodebench_weighted-avg_num_prune_attn_5
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754896720_eval_6419_livecodebench_weighted-avg_num_prune_attn_5
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 15.53% ± 0.58%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
16.83%
86
511
2
15.26%
78
511
3
16.24%
83
511
4
14.29%
73
511
5
17.03%
87
511
6
13.50%
69
511
Unroll-Qwen2.5-7B-Instruct_1754897307_eval_6419_livecodebench_simple-avg_num_prune_attn_5
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754897307_eval_6419_livecodebench_simple-avg_num_prune_attn_5
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 7.60% ± 0.22%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
7.83%
40
511
2
7.44%
38
511
3
8.41%
43
511
4
7.83%
40
511
5
7.05%
36
511
6
7.05%
36
511
Unroll-Qwen2.5-7B-Instruct_1754898337_eval_6419_livecodebench_rank-based_num_prune_attn_6
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754898337_eval_6419_livecodebench_rank-based_num_prune_attn_6
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 4.01% ± 0.19%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
3.33%
17
511
2
3.91%
20
511
3
3.91%
20
511
4
3.91%
20
511
5
4.31%
22
511
6
4.70%
24
511
Unroll-Qwen2.5-7B-Instruct_1754900584_eval_6419_livecodebench_geometric_num_prune_attn_2
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754900584_eval_6419_livecodebench_geometric_num_prune_attn_2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 23.52% ± 0.52%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
24.85%
127
511
2
21.14%
108
511
3
24.27%
124
511
4
23.48%
120
511
5
23.68%
121
511
6
23.68%
121
511
Unroll-Qwen2.5-7B-Instruct_1754900883_eval_6419_livecodebench_geometric_num_prune_attn_6
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754900883_eval_6419_livecodebench_geometric_num_prune_attn_6
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 7.40% ± 0.33%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
7.44%
38
511
2
8.41%
43
511
3
7.05%
36
511
4
7.63%
39
511
5
6.07%
31
511
6
7.83%
40
511
