Team Ai
Apppublic

nareshns2004/GPU-Inference-Optimization-Lab

sourceHugging Faceapache-2.0updated 15h agoView on Hugging Face
0likes
README.md25 linesDownload Raw Back to root
1---2title: GPU Inference Optimization Lab3emoji: ๐Ÿ“ˆ4colorFrom: blue5colorTo: blue6sdk: static7pinned: false8license: apache-2.09short_description: LLM Inference performance and GPU efficiency10---11 12# GPU Inference Optimization Lab13 14Measure how optimizations change LLM inference speed and GPU efficiency: precision (fp32/fp16/bf16), batch size, `torch.compile`, and later quantization and attention kernels.15 16- `lab/measure.py` โ€“ times generation, records tokens/s and peak GPU memory17- `lab/run_experiments.py` โ€“ runs a grid of settings and prints a comparison table18 19## Run locally20 21```bash22pip install -r requirements.txt23python lab/run_experiments.py24```25