dennisjooo/emotion_classification
Emotion Classification
This model is a fine-tuned version of google/vit-base-patch16-224-in21k on the FastJobs/Visual_Emotional_Analysis dataset.
In theory, the accuracy for a random guess on this dataset is 0.125 (8 labels and you need to choose one).
It achieves the following results on the evaluation set:
- Loss: 1.0511
- Accuracy: 0.6687
- Precision: 0.7104
- F1: 0.6713
Model description
The Vision Transformer base version trained on ImageNet-21K released by Google. Further details can be found on their repo.
Training and evaluation data
Data Split
Trained on FastJobs/Visual_Emotional_Analysis dataset. Used a 4:1 ratio for training and development sets and a random seed of 42. Also used a seed of 42 for batching the data, completely unrelated lol.
Pre-processing Augmentation
The main pre-processing phase for both training and evaluation includes:
- Bilinear interpolation to resize the image to (224, 224, 3) because it uses ImageNet images to train the original model
- Normalizing images using a mean and standard deviation of [0.5, 0.5, 0.5] just like the original model
Other than the aforementioned pre-processing, the training set was augmented using:
- Random horizontal & vertical flip
- Color jitter
- Random resized crop
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-05
- trainbatchsize: 64
- evalbatchsize: 64
- seed: 42
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lrschedulertype: cosinewithrestarts
- lrschedulerwarmup_steps: 150
- num_epochs: 300
Training results
Framework versions
- Transformers 4.33.0
- Pytorch 2.0.0
- Datasets 2.1.0
- Tokenizers 0.13.3
