FreshCrawl/rednote-xiaohongshu-notes
RedNote (Xiaohongshu) Notes with Engagement and Save Rates 10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography. Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot. The finding this dataset exists for A post can be useful or it can be… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/rednote-xiaohongshu-notes.
RedNote (Xiaohongshu) Notes with Engagement and Save Rates
10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography.
Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot.
The finding this dataset exists for
A post can be useful or it can be entertaining. Likes measure the second. Saves measure the first. RedNote exposes both, so you can separate them.
Median save-to-like ratio, by vertical:
A 2.4x spread, and it orders exactly as the utility reading predicts. Recipes and parenting advice get saved for later. Fitness and fashion get liked and scrolled past. Overall median is 0.51, which is itself remarkable: on most platforms saves are an order of magnitude rarer than likes.
This is a median comparison across verticals, not a controlled model. Creator size, post age and format are uncontrolled. Take it as a well-supported starting point and a reason the columns are here, not as a finished result.
At a glance
Top regions by volume: Guangdong (1,540), Zhejiang (941), Jiangsu (519), Sichuan (448), Shanghai (434), Shandong (429), Beijing (425), Fujian (406).
Why this is useful for ML
- Multi-target engagement regression. Four independent targets (likes, saves, comments, shares) over shared Chinese text. Predicting saves is a genuinely different task from predicting likes, and this is the rare corpus where you can test that claim.
- Chinese social text at scale. 10,427 posts of real user-written Simplified Chinese, median 175 characters, with hashtags inline. Not translated, not synthetic.
- Geography without GPS. Province-level origin on 88% of rows supports regional discourse analysis, a task usually blocked by the absence of any location signal in social corpora.
- Content-type classification. Eight labelled verticals with a strong engagement signature each, which makes for a non-trivial classification target where the text and the metrics disagree in interesting ways.
from datasets import load_dataset
ds = load_dataset("FreshCrawl/rednote-xiaohongshu-notes", split="train")
print(ds)Fields
Collection method
- Collected on 2026-09-03 via keyword search across 10 verticals, then per-post detail retrieval for full text and province.
- Keywords were chosen for behavioural variance rather than volume, so the engagement analysis has contrast to work with: 护肤, 美食, 穿搭, 旅行, 健身, 育儿, 家居, 考研, 数码, 宠物 and related terms.
- Search returns a truncated post body, so every row here went through a second detail fetch. That is why
post_textis real prose rather than a hashtag stub. - Deduplicated on
note_id. 2.3% of search results were cross-keyword duplicates.
Limitations and bias
- Eight verticals, not ten. Collection stopped early, so
petsandtechare absent despite being in the keyword plan, andcareeris under-sampled at 495 posts against roughly 1,300 for the others. Do not read the vertical counts as relative platform popularity. - Search-ranked, not random. Posts come from keyword search results, which favour engagement. This is a sample of discoverable content, and it will over-represent successful posts. Absolute engagement figures are inflated relative to a random draw from the platform.
- Survivorship. Deleted and private posts are invisible by construction.
- Engagement is a snapshot. Counts were read once, at collection time. An older post has had longer to accumulate them, and
published_atspans 2020 to 2026, so age is a live confounder in any engagement model. Control for it. - Province is self-reported by the platform, derived from network origin, missing on 12%, and it reflects where a poster was, not where they are from.
- Simplified Chinese only.
Privacy
Creator identity is not in this dataset. Display name, avatar URL, user id and access tokens were dropped before publication. author_hash is a truncated SHA-256 of the platform id with an unpublished salt, which allows per-creator grouping without carrying an identifier.
Post text is reproduced as written and may mention people or places. No attempt was made to rewrite it. @-mentions of other users were dropped.
Licence
Post text remains the intellectual property of its individual authors and of Xiaohongshu. This compilation is published for research and educational use. Cite the dataset and link back if you publish work based on it. Do not redistribute the raw file as a commercial product.
Fresher data
This is a static snapshot from 2026-09-03 and it will not be updated. It ages from the day it was published, which is fine for research and useless for anything operational.
The scrapers that produced it are public, and the same schema comes back live:
- RedNote Search Scraper: keyword search, 500 notes per 30s
- RedNote Note Detail Scraper: full text, province, engagement, media
- RedNote Comments Scraper: threaded comments, per-commenter province
Code written against this file works unchanged against fresh data.
