TheRealVigilante/Processed_Top_15k_Anime
π¦ Anime Recommender Dataset (Sentence-BERT Ready) This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT. π Original Dataset Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025)Author: Quan ThanLicense: Apache 2.0 π§β¦ See the full description on the dataset page: https://huggingface.co/datasets/TheRealVigilante/Processed_Top_15k_Anime.
π¦ Anime Recommender Dataset (Sentence-BERT Ready)
This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT.
π Original Dataset
Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025) Author: Quan Than License: Apache 2.0
π§ Modifications & Enhancements
The original dataset has been modified for better use in AI-driven recommendation tasks:
β Column Cleanup
- Removed excessive trailing spaces and newline artifacts in all fields.
- Standardized formatting across fields like
genres,studios,producers.
β Feature Engineering
- Added a new column called `combined_features`, which concatenates the most relevant text fields:
genrestypestudiosproducerssourceratingsynopsis
This column is intended for use in embedding generation via NLP models (e.g. Sentence-BERT, TfidfVectorizer).
β Consistency Fixes
- Ensured that all text fields are valid UTF-8.
- Filled missing or null values with empty strings (
'') for compatibility with model pipelines.
π§ Use Case
This dataset is designed to be used in content-based anime recommender systems, especially those using semantic similarity techniques. Typical workflow:
- Load the
combined_featurescolumn. - Generate embeddings using a model like
all-MiniLM-L6-v2(fromsentence-transformers). - Use cosine similarity to find top matches.
- Filter or rank based on score, popularity, or genre.
π Columns
π Format
- File:
anime.csv - Encoding: UTF-8
- Format: Standard CSV
- Rows: \~15,000 (depending on filtering)
- Columns: 23 including
combined_features
π Author Notes
This dataset was prepared by:
*π¨βπ» Youssef ElNahas β aka TheVigilante***
- π GitHub Profile
- π¬ For collaboration or feedback, feel free to reach out!
π‘οΈ Disclaimer
This dataset is a derivative of the publicly shared dataset on Kaggle. Please ensure you review and comply with the original datasetβs license and usage terms.
