Team Ai
Modelpublic

kcxain/KernelZero-CUDA-7B

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes339downloads
Model Card

KernelZero-CUDA-7B

KernelZero-CUDA-7B is the CUDA kernel generation checkpoint reported as the final KernelZero CUDA model. It is based on Qwen2.5-Coder-7B-Instruct and trained with the KernelZero pipeline.

Paper

KernelZero: Continual Co-Evolution of LLMs for Kernel Generation

KernelBench results

The paper reports the following results with 10 sampled candidates per problem.

Benchmarkpass@1pass@5pass@10fast\_1@1fast\_1@10fast\_2@1fast\_2@10
Level 175.898.87100.017.629.07.212.0
Level 269.693.7097.02.412.01.26.0

Intended use

This checkpoint generates CUDA extensions from PyTorch modules. Use the same prompt template and validation environment as KernelZero for reproducible evaluation.

Loading

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "kcxain/KernelZero-CUDA-7B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)

Checkpoint provenance

This release contains the merged Hugging Face inference weights from the final 80-step CUDA Coder followed by the 40-step continuation used for the paper's main CUDA result. Optimizer states and distributed training shards are excluded.