Chima207/distilbert_amazon_goodreads_book_classification
<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->
distilbertamazongoodreadsbookclassification
This model was trained from scratch on the Goodreads-Books dataset. It achieves the following results on the evaluation set:
- Loss: 1.7274
- Accuracy: 0.5135
- F1 Score: 0.4989
- Precision: 0.5060
- Recall: 0.5135
Model description
This model evaluates a two-stage sequential transfer learning approach (Amazon to Goodreads). Starting from distilbert-base-uncased, the model was first fine-tuned on structured Amazon book metadata and subsequently fine-tuned on the heuristically mapped, oversampled Goodreads dataset) across 31 categories. While this sequential adaptation yields a peak Accuracy of 51.38% (a marginal +0.80% gain over single-stage Goodreads fine-tuning), the empirical results demonstrate diminishing returns. The additional computational overhead of two-stage fine-tuning is rarely justified for real-world deployment compared to direct single-stage adaptation.
Intended uses & limitations
- Research and analysis of domain adaptation mechanisms between structured e-commerce data and noisy user-generated web content.
- High-accuracy deployment scenarios where every marginal percentage increase in classification performance is critical regardless of compute budget.
- Requires substantial training time and computational resources for a minimal performance gain over single-stage fine-tuning.
- Inherits the classification domain boundaries of the 31 target Amazon categories.
Datasets
- Goodreads-Dataset: Hugging Face Repository (Original: [BrightData/Goodreads-Books](https://huggingface.co/datasets/BrightData/Goodreads-Books))
- Amazon-Dataset: [Kaggle Amazon Kindle Books Dataset](https://www.kaggle.com/datasets/asaniczka/amazon-kindle-books-dataset-2023-130k-books))
Training and evaluation data
- Stage 1 (Pre-training / Source Domain):
- Source: Structured Amazon Kindle metadata.
- Purpose: Learn clean domain representations and sentence semantics on structured book descriptions.
- Stage 2 (Target Adaptation / Target Domain):
- Source: Preprocessed Goodreads book metadata.
- Preprocessing: 1,003 crowdsourced Goodreads shelves heuristically mapped into the 31 Amazon Kindle categories and balanced using Random Oversampling.
- Purpose: Adapt the pre-trained weights to the noisy, community-driven text characteristics of Goodreads entries.
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- trainbatchsize: 4
- evalbatchsize: 2
- seed: 42
- gradientaccumulationsteps: 4
- totaltrainbatch_size: 16
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lrschedulertype: linear
- num_epochs: 2
- mixedprecisiontraining: Native AMP
Training results
Framework versions
- Transformers 4.45.2
- Pytorch 2.5.1
- Datasets 4.1.1
- Tokenizers 0.20.1
Academic Context & Citation / Akademischer Kontext
This repository and model were developed as part of a Bachelor's thesis in 2026.
- Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
- License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)
Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.
- Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
- Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)
