kcxain/KernelZero-CUDA-7B
0339
KernelZero-CUDA-7B
KernelZero-CUDA-7B is the CUDA kernel generation checkpoint reported as the final KernelZero CUDA model. It is based on Qwen2.5-Coder-7B-Instruct and trained with the KernelZero pipeline.
Paper
KernelZero: Continual Co-Evolution of LLMs for Kernel Generation
KernelBench results
The paper reports the following results with 10 sampled candidates per problem.
Intended use
This checkpoint generates CUDA extensions from PyTorch modules. Use the same prompt template and validation environment as KernelZero for reproducible evaluation.
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "kcxain/KernelZero-CUDA-7B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)Checkpoint provenance
This release contains the merged Hugging Face inference weights from the final 80-step CUDA Coder followed by the 40-step continuation used for the paper's main CUDA result. Optimizer states and distributed training shards are excluded.
