Team Ai
Modelpublic

Chima207/distilbert_amazon_goodreads_book_classification

sourceHugging Faceccupdated 7d agoView on Hugging Face
1likes36downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

distilbertamazongoodreadsbookclassification

This model was trained from scratch on the Goodreads-Books dataset. It achieves the following results on the evaluation set:

  • —Loss: 1.7274
  • —Accuracy: 0.5135
  • —F1 Score: 0.4989
  • —Precision: 0.5060
  • —Recall: 0.5135

Model description

This model evaluates a two-stage sequential transfer learning approach (Amazon to Goodreads). Starting from distilbert-base-uncased, the model was first fine-tuned on structured Amazon book metadata and subsequently fine-tuned on the heuristically mapped, oversampled Goodreads dataset) across 31 categories. While this sequential adaptation yields a peak Accuracy of 51.38% (a marginal +0.80% gain over single-stage Goodreads fine-tuning), the empirical results demonstrate diminishing returns. The additional computational overhead of two-stage fine-tuning is rarely justified for real-world deployment compared to direct single-stage adaptation.

Intended uses & limitations

  • —Research and analysis of domain adaptation mechanisms between structured e-commerce data and noisy user-generated web content.
  • —High-accuracy deployment scenarios where every marginal percentage increase in classification performance is critical regardless of compute budget.
  • —Requires substantial training time and computational resources for a minimal performance gain over single-stage fine-tuning.
  • —Inherits the classification domain boundaries of the 31 target Amazon categories.

Datasets

  • —Goodreads-Dataset: Hugging Face Repository (Original: [BrightData/Goodreads-Books](https://huggingface.co/datasets/BrightData/Goodreads-Books))
  • —Amazon-Dataset: [Kaggle Amazon Kindle Books Dataset](https://www.kaggle.com/datasets/asaniczka/amazon-kindle-books-dataset-2023-130k-books))

Training and evaluation data

  1. 1.Stage 1 (Pre-training / Source Domain):
  2. 2.Source: Structured Amazon Kindle metadata.
  3. 3.Purpose: Learn clean domain representations and sentence semantics on structured book descriptions.
  4. 4.Stage 2 (Target Adaptation / Target Domain):
  5. 5.Source: Preprocessed Goodreads book metadata.
  6. 6.Preprocessing: 1,003 crowdsourced Goodreads shelves heuristically mapped into the 31 Amazon Kindle categories and balanced using Random Oversampling.
  7. 7.Purpose: Adapt the pre-trained weights to the noisy, community-driven text characteristics of Goodreads entries.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 2e-05
  • —trainbatchsize: 4
  • —evalbatchsize: 2
  • —seed: 42
  • —gradientaccumulationsteps: 4
  • —totaltrainbatch_size: 16
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —num_epochs: 2
  • —mixedprecisiontraining: Native AMP

Training results

Training LossEpochStepValidation LossAccuracyF1 ScorePrecisionRecall
0.4141.000085191.72740.51350.49890.50600.5135
0.2371.9999170382.04050.52100.51180.51460.5210

Framework versions

  • —Transformers 4.45.2
  • —Pytorch 2.5.1
  • —Datasets 4.1.1
  • —Tokenizers 0.20.1

Academic Context & Citation / Akademischer Kontext

This repository and model were developed as part of a Bachelor's thesis in 2026.

  • —Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
  • —License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)

Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.

  • —Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
  • —Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)