supplyscience/forecasting-feature-engineering
Machine Learning for Retail Sales Forecasting
A Streamlit application to study the impact of LightGBM hyperparameters on sales forecasting accuracy, based on the M5 Walmart dataset.
The app provides:
- Context tab — article content explaining the dataset, approach, and 6 feature engineering buckets
- Experiment tab — interactive hyperparameter tuning with per-store LightGBM training, learning curves, and feature importance analysis
Project Structure
forecasting/
├── app.py # Streamlit application
├── preprocess.py # Offline data preprocessing pipeline
├── Dockerfile # Docker image for deployment
├── requirements.txt # pip dependencies (Docker)
├── pyproject.toml # uv dependencies (local dev)
├── data/
│ ├── processed_data.parquet # Pre-processed dataset (9.7 MB)
│ ├── features.json # Feature list
│ └── label_encodings.json # Label encoding mappings
├── Data/ # Raw M5 dataset (not shipped in Docker)
├── images/ # Visuals used in the Context tab
├── archives/
│ ├── initial_notebooks/ # Original Jupyter notebooks
│ └── article_screens/ # Screenshots of the published article
└── docs/
├── notebooks_analysis.md # Analysis of the original notebooks
└── article_summary.md # Summary of the published articleLocal Setup
Option 1: uv (recommended)
# Install uv if you don't have it
curl -LsSf https://astral.sh/uv/install.sh | sh
# Clone and enter the project
git clone <repo-url> && cd forecasting
# Install dependencies
uv sync
# Run the app
uv run streamlit run app.pyThe app will be available at http://localhost:8501.
Option 2: Docker Compose — Dev Mode (recommended for development)
docker compose up --buildThis mounts your local app.py, data/ and images/ into the container. Streamlit watches for file changes and auto-reloads — just edit app.py and the app refreshes in your browser.
The app will be available at http://localhost:7860.
To stop: Ctrl+C or docker compose down.
Option 3: Docker (standalone)
# Build the image
docker build -t forecast-app .
# Run the container
docker run -p 7860:7860 forecast-appThe app will be available at http://localhost:7860.
Option 4: pip
# Create a virtual environment
python -m venv .venv
source .venv/bin/activate # Linux/macOS
# .venv\Scripts\activate # Windows
# Install dependencies
pip install -r requirements.txt
# Run the app
streamlit run app.pyRegenerating the Processed Data
The data/ folder already contains the pre-processed dataset. If you need to regenerate it from the raw M5 files (in Data/):
uv run python preprocess.pyThis filters the raw dataset to Wisconsin (3 stores, 200 sampled items), runs the full feature engineering pipeline, and outputs data/processed_data.parquet.
Required raw files in Data/:
sales_train_evaluation.csvcalendar.csvsell_prices.csv
