Team Ai
Modelpublic

jangwon-kim-cocel/Bayesian-Policy-Distillation

sourceHugging Facemitupdated 8mo agoView on Hugging Face
1likes
Model Card

<div align="center"> <h1>Bayesian Policy Distillation</h1> <h3>Towards Lightweight and Fast Neural Policy Networks</h3>

<a href="https://www.python.org/"> <img src="https://img.shields.io/badge/Python-3.7+-blue?logo=python&style=flat-square" alt="Python Badge"/> </a> &nbsp;&nbsp; <a href="https://pytorch.org/"> <img src="https://img.shields.io/badge/PyTorch-1.8+-EE4C2C?logo=pytorch&style=flat-square" alt="PyTorch Badge"/> </a> &nbsp;&nbsp; <a href="https://doi.org/10.1016/j.engappai.2025.113539"> <img src="https://img.shields.io/badge/EAAI%202026-Published-success?style=flat-square" alt="EAAI Badge"/> </a> &nbsp;&nbsp; <a href="https://www.elsevier.com/"> <img src="https://img.shields.io/badge/Elsevier-Journal-orange?style=flat-square" alt="Elsevier Badge"/> </a> <br/><br/> <img src="./gifforreadme.gif" width="550px"/>

</div>


Engineering Applications of Artificial Intelligence (EAAI 2026)

PyTorch Implementation

This repository contains a PyTorch implementation of Bayesian Policy Distillation (BPD) of the paper:

Bayesian policy distillation: Towards lightweight and fast neural policy networks Jangwon Kim, Yoonsu Jang, Jonghyeok Park, Yoonhee Gil, Soohee Han Engineering Applications of Artificial Intelligence, Volume 166, 2026

๐Ÿ“„ Paper Link

DOI: https://doi.org/10.1016/j.engappai.2025.113539 Journal: Engineering Applications of Artificial Intelligence

Bayesian Policy Distillation

BPD achieves extreme policy compression through offline reinforcement learning by:

  1. 1.Bayesian Neural Networks: Uncertainty-driven dynamic weight pruning
  2. 2.Sparse Variational Dropout: Automatic sparsity induction via KL regularization
  3. 3.Offline RL Framework: Value optimization + behavior cloning

$$ \mathcal{L}{BPD}(\theta, \alpha) = -\lambda Q{\psi1}(s, \pi\omega(s)) + \frac{|\mathcal{D}|}{M}\sum{m=1}^{M}(\pi{\omegam}(sm) - am)^2 + \eta \cdot D{KL}(q(\omega|\theta,\alpha) \| p(\omega)) $$

Key Results:

  • โ€”~98% compression (1.5-2.5% sparsity) while maintaining performance
  • โ€”4.5ร— faster inference on embedded systems
  • โ€”Successfully deployed on real inverted pendulum with 78% inference time reduction

Quick Start

Basic Training

bash
python main.py --env-name Hopper-v3 --level expert --random-seed 1

Custom Configuration

bash
python main.py \
    --env-name Walker2d-v3 \
    --level medium \
    --student-hidden-dims "(128, 128)" \
    --alpha-threshold 2 \
    --nu 4 \
    --h 0.5

Available Environments

  • โ€”Hopper-v3, Walker2d-v3, HalfCheetah-v3, Ant-v3

Teacher Policy Levels

  • โ€”expert: High-performance teacher policy
  • โ€”medium: Moderate-performance teacher policy

Key Hyperparameters

ParameterDefaultDescription
--student-hidden-dims(128, 128)Student network hidden layer sizes
--alpha-threshold2Pruning threshold for log(ฮฑ) (higher = less compression)
--nu4KL weight annealing speed
--h0.5Q-value loss coefficient
--batch-size256Mini-batch size
--max-teaching-count1000000Total training iterations
--eval-freq5000Evaluation frequency

Adjusting Compression:

  • โ€”--alpha-threshold 3-4: Conservative pruning
  • โ€”--alpha-threshold 2: Balanced [default]
  • โ€”--alpha-threshold 1: Aggressive pruning

Results

MuJoCo Benchmark (Expert Teacher)

EnvironmentTeacherBPD (Ours)SparsityCompression
Ant-v3536454552.40%41.7ร—
Walker2d-v3535748171.68%59.5ร—
Hopper-v3358331341.35%74.1ร—
HalfCheetah-v311432103552.21%45.2ร—

Real Hardware (Inverted Pendulum)

  • โ€”Inference: 1.36ms โ†’ 0.30ms (4.5ร— faster)
  • โ€”Memory: 290.82KB โ†’ 4.43KB (98.5% reduction)
  • โ€”Parameters: 72,705 โ†’ 1,108 (65.6ร— compression) ---

Citation

bibtex
@article{kim2026bayesian,
  title={Bayesian policy distillation: Towards lightweight and fast neural policy networks},
  author={Kim, Jangwon and Jang, Yoonsu and Park, Jonghyeok and Gil, Yoonhee and Han, Soohee},
  journal={Engineering Applications of Artificial Intelligence},
  volume={166},
  pages={113539},
  year={2026},
  publisher={Elsevier}
}