jangwon-kim-cocel/Bayesian-Policy-Distillation
<div align="center"> <h1>Bayesian Policy Distillation</h1> <h3>Towards Lightweight and Fast Neural Policy Networks</h3>
<a href="https://www.python.org/"> <img src="https://img.shields.io/badge/Python-3.7+-blue?logo=python&style=flat-square" alt="Python Badge"/> </a> <a href="https://pytorch.org/"> <img src="https://img.shields.io/badge/PyTorch-1.8+-EE4C2C?logo=pytorch&style=flat-square" alt="PyTorch Badge"/> </a> <a href="https://doi.org/10.1016/j.engappai.2025.113539"> <img src="https://img.shields.io/badge/EAAI%202026-Published-success?style=flat-square" alt="EAAI Badge"/> </a> <a href="https://www.elsevier.com/"> <img src="https://img.shields.io/badge/Elsevier-Journal-orange?style=flat-square" alt="Elsevier Badge"/> </a> <br/><br/> <img src="./gifforreadme.gif" width="550px"/>
</div>
Engineering Applications of Artificial Intelligence (EAAI 2026)
PyTorch Implementation
This repository contains a PyTorch implementation of Bayesian Policy Distillation (BPD) of the paper:
Bayesian policy distillation: Towards lightweight and fast neural policy networks Jangwon Kim, Yoonsu Jang, Jonghyeok Park, Yoonhee Gil, Soohee Han Engineering Applications of Artificial Intelligence, Volume 166, 2026
๐ Paper Link
DOI: https://doi.org/10.1016/j.engappai.2025.113539 Journal: Engineering Applications of Artificial Intelligence
Bayesian Policy Distillation
BPD achieves extreme policy compression through offline reinforcement learning by:
- Bayesian Neural Networks: Uncertainty-driven dynamic weight pruning
- Sparse Variational Dropout: Automatic sparsity induction via KL regularization
- Offline RL Framework: Value optimization + behavior cloning
$$ \mathcal{L}{BPD}(\theta, \alpha) = -\lambda Q{\psi1}(s, \pi\omega(s)) + \frac{|\mathcal{D}|}{M}\sum{m=1}^{M}(\pi{\omegam}(sm) - am)^2 + \eta \cdot D{KL}(q(\omega|\theta,\alpha) \| p(\omega)) $$
Key Results:
- ~98% compression (1.5-2.5% sparsity) while maintaining performance
- 4.5ร faster inference on embedded systems
- Successfully deployed on real inverted pendulum with 78% inference time reduction
Quick Start
Basic Training
python main.py --env-name Hopper-v3 --level expert --random-seed 1Custom Configuration
python main.py \
--env-name Walker2d-v3 \
--level medium \
--student-hidden-dims "(128, 128)" \
--alpha-threshold 2 \
--nu 4 \
--h 0.5Available Environments
Hopper-v3,Walker2d-v3,HalfCheetah-v3,Ant-v3
Teacher Policy Levels
expert: High-performance teacher policymedium: Moderate-performance teacher policy
Key Hyperparameters
Adjusting Compression:
--alpha-threshold 3-4: Conservative pruning--alpha-threshold 2: Balanced [default]--alpha-threshold 1: Aggressive pruning
Results
MuJoCo Benchmark (Expert Teacher)
Real Hardware (Inverted Pendulum)
- Inference: 1.36ms โ 0.30ms (4.5ร faster)
- Memory: 290.82KB โ 4.43KB (98.5% reduction)
- Parameters: 72,705 โ 1,108 (65.6ร compression) ---
Citation
@article{kim2026bayesian,
title={Bayesian policy distillation: Towards lightweight and fast neural policy networks},
author={Kim, Jangwon and Jang, Yoonsu and Park, Jonghyeok and Gil, Yoonhee and Han, Soohee},
journal={Engineering Applications of Artificial Intelligence},
volume={166},
pages={113539},
year={2026},
publisher={Elsevier}
}