Team Ai
Datasetpublic

TheRealVigilante/Processed_Top_15k_Anime

πŸ“¦ Anime Recommender Dataset (Sentence-BERT Ready) This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT. πŸ“Œ Original Dataset Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025)Author: Quan ThanLicense: Apache 2.0 πŸ”§β€¦ See the full description on the dataset page: https://huggingface.co/datasets/TheRealVigilante/Processed_Top_15k_Anime.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes41downloads
Dataset Card

πŸ“¦ Anime Recommender Dataset (Sentence-BERT Ready)

This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT.


πŸ“Œ Original Dataset

Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025) Author: Quan Than License: Apache 2.0


πŸ”§ Modifications & Enhancements

The original dataset has been modified for better use in AI-driven recommendation tasks:

βœ… Column Cleanup

  • β€”Removed excessive trailing spaces and newline artifacts in all fields.
  • β€”Standardized formatting across fields like genres, studios, producers.

βœ… Feature Engineering

  • β€”Added a new column called `combined_features`, which concatenates the most relevant text fields:
  • β€”genres
  • β€”type
  • β€”studios
  • β€”producers
  • β€”source
  • β€”rating
  • β€”synopsis

This column is intended for use in embedding generation via NLP models (e.g. Sentence-BERT, TfidfVectorizer).

βœ… Consistency Fixes

  • β€”Ensured that all text fields are valid UTF-8.
  • β€”Filled missing or null values with empty strings ('') for compatibility with model pipelines.

🧠 Use Case

This dataset is designed to be used in content-based anime recommender systems, especially those using semantic similarity techniques. Typical workflow:

  1. 1.Load the combined_features column.
  2. 2.Generate embeddings using a model like all-MiniLM-L6-v2 (from sentence-transformers).
  3. 3.Use cosine similarity to find top matches.
  4. 4.Filter or rank based on score, popularity, or genre.

πŸ—‚ Columns

ColumnDescription
anime_idMyAnimeList unique ID
anime_urlURL to the anime’s MAL page
image_urlLink to the anime’s poster image
namePrimary title
english_nameEnglish-translated title (if any)
japanese_namesJapanese name(s)
scoreAverage user score
genresComma-separated genres
synopsisFull text description
typeFormat: TV, Movie, OVA, etc.
episodesNumber of episodes
premieredSeason/year it first aired
producersCompanies involved in production
studiosAnimation studio(s)
sourceOriginal work type (Manga, Light Novel, etc.)
durationTime per episode
ratingContent rating (e.g., PG-13, R)
rankRank on MAL
popularityPopularity rank
favoritesNumber of users marking it as favorite
scored_byNumber of users who rated it
membersNumber of users in total interested
combined_featuresCustom field combining genres, studios, synopsis, etc. for embeddings use

πŸ“‚ Format

  • β€”File: anime.csv
  • β€”Encoding: UTF-8
  • β€”Format: Standard CSV
  • β€”Rows: \~15,000 (depending on filtering)
  • β€”Columns: 23 including combined_features

πŸ™Œ Author Notes

This dataset was prepared by:

*πŸ‘¨β€πŸ’» Youssef ElNahas β€” aka TheVigilante***

  • β€”πŸ”— GitHub Profile
  • β€”πŸ’¬ For collaboration or feedback, feel free to reach out!

πŸ›‘οΈ Disclaimer

This dataset is a derivative of the publicly shared dataset on Kaggle. Please ensure you review and comply with the original dataset’s license and usage terms.