datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Personal-Cambodian-Content-Creators
Personal Cambodian Content Creators
Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) focused on generating realistic personas and profiles for digital content creators and influencers in Cambodia.
π Project Vision
I hope this dataset helps your project succeed. Whether you are building a creative writing assistant, a marketing simulation tool, or conducting research on⦠See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Cambodian-Content-Creators.world-cup-creators
World Cup 2026 TikTok Creators β Open Dataset
Per-nation profile of TikTok creators across the 48 nations qualified for the
2026 FIFA World Cup, from Crawlora's creators_search dataset (β3.33M
discoverable creators), June 2026 snapshot. Aggregate rollups only β counts
and rates by nation, never individual creators.
π Study: https://crawlora.net/blog/world-cup-creators-2026
Coverage
A creator's country comes primarily from TikTok's own account region (with a⦠See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/world-cup-creators.smolified-creatorsei
π€ smolified-creatorsei
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-creatorsei.
π¦ Asset Details
Origin: Smolify Foundry (Job ID: b95889f4)
Records: 7722
Type: Synthetic Instruction Tuning Data
βοΈ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
