Team Ai
Datasetpublic

FreshCrawl/rednote-xiaohongshu-notes

RedNote (Xiaohongshu) Notes with Engagement and Save Rates 10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography. Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot. The finding this dataset exists for A post can be useful or it can be… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/rednote-xiaohongshu-notes.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes666downloads
Dataset Card

RedNote (Xiaohongshu) Notes with Engagement and Save Rates

10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography.

Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot.

The finding this dataset exists for

A post can be useful or it can be entertaining. Likes measure the second. Saves measure the first. RedNote exposes both, so you can separate them.

Median save-to-like ratio, by vertical:

VerticalSave/likeVerticalSave/like
美食 food0.74家居 home0.62
育儿 parenting0.72美妆 beauty0.46
旅行 travel0.63穿搭 fashion0.37
职场 career0.62健身 fitness0.31

A 2.4x spread, and it orders exactly as the utility reading predicts. Recipes and parenting advice get saved for later. Fitness and fashion get liked and scrolled past. Overall median is 0.51, which is itself remarkable: on most platforms saves are an order of magnitude rarer than likes.

This is a median comparison across verticals, not a controlled model. Creator size, post age and format are uncontrolled. Take it as a well-supported starting point and a reason the columns are here, not as a finished result.

At a glance

Posts10,427
Verticals8
Distinct creators (hashed)8,921
Province / region values110
Province populated87.9%
Post date range2020-01-01 to 2026-09-03
Snapshot collected2026-09-03
Median likes / saves712 / 331
Median post length175 chars
Video posts14.2%

Top regions by volume: Guangdong (1,540), Zhejiang (941), Jiangsu (519), Sichuan (448), Shanghai (434), Shandong (429), Beijing (425), Fujian (406).

Why this is useful for ML

  • —Multi-target engagement regression. Four independent targets (likes, saves, comments, shares) over shared Chinese text. Predicting saves is a genuinely different task from predicting likes, and this is the rare corpus where you can test that claim.
  • —Chinese social text at scale. 10,427 posts of real user-written Simplified Chinese, median 175 characters, with hashtags inline. Not translated, not synthetic.
  • —Geography without GPS. Province-level origin on 88% of rows supports regional discourse analysis, a task usually blocked by the absence of any location signal in social corpora.
  • —Content-type classification. Eight labelled verticals with a strong engagement signature each, which makes for a non-trivial classification target where the text and the metrics disagree in interesting ways.
python
from datasets import load_dataset

ds = load_dataset("FreshCrawl/rednote-xiaohongshu-notes", split="train")
print(ds)

Fields

ColumnTypeDescription
note_idstringRedNote's unique post id. Primary key, deduplicated
verticalstringOne of 8 content categories, from the search cluster that surfaced the post
search_keywordstringThe Chinese keyword that surfaced this post
note_typestringnormal (image post) or video
titlestringPost headline
post_textstringFull post body in Simplified Chinese, hashtags inline
post_text_lengthintCharacter count of post_text
published_atstringPublication date, ISO 8601
ip_provincestringProvince or country RedNote displays for the poster. Empty on 12%
liked_countfloatLikes
collected_countfloatSaves. The distinctive metric
comments_countfloatComments
shared_countfloatShares
save_to_like_ratiofloatcollected_count / liked_count, precomputed
comment_to_like_ratiofloatcomments_count / liked_count, precomputed
tagsstringPipe-separated hashtags
tag_countintNumber of hashtags
image_countintImages attached
has_videoint1 if the post carries video
author_hashstringSalted, truncated hash of the creator id. Groups posts by creator without identifying anyone. The salt is not published
note_urlstringSource post URL

Collection method

  • —Collected on 2026-09-03 via keyword search across 10 verticals, then per-post detail retrieval for full text and province.
  • —Keywords were chosen for behavioural variance rather than volume, so the engagement analysis has contrast to work with: 护肤, 美食, 穿搭, 旅行, 健身, 育儿, 家居, 考研, 数码, 宠物 and related terms.
  • —Search returns a truncated post body, so every row here went through a second detail fetch. That is why post_text is real prose rather than a hashtag stub.
  • —Deduplicated on note_id. 2.3% of search results were cross-keyword duplicates.

Limitations and bias

  • —Eight verticals, not ten. Collection stopped early, so pets and tech are absent despite being in the keyword plan, and career is under-sampled at 495 posts against roughly 1,300 for the others. Do not read the vertical counts as relative platform popularity.
  • —Search-ranked, not random. Posts come from keyword search results, which favour engagement. This is a sample of discoverable content, and it will over-represent successful posts. Absolute engagement figures are inflated relative to a random draw from the platform.
  • —Survivorship. Deleted and private posts are invisible by construction.
  • —Engagement is a snapshot. Counts were read once, at collection time. An older post has had longer to accumulate them, and published_at spans 2020 to 2026, so age is a live confounder in any engagement model. Control for it.
  • —Province is self-reported by the platform, derived from network origin, missing on 12%, and it reflects where a poster was, not where they are from.
  • —Simplified Chinese only.

Privacy

Creator identity is not in this dataset. Display name, avatar URL, user id and access tokens were dropped before publication. author_hash is a truncated SHA-256 of the platform id with an unpublished salt, which allows per-creator grouping without carrying an identifier.

Post text is reproduced as written and may mention people or places. No attempt was made to rewrite it. @-mentions of other users were dropped.

Licence

Post text remains the intellectual property of its individual authors and of Xiaohongshu. This compilation is published for research and educational use. Cite the dataset and link back if you publish work based on it. Do not redistribute the raw file as a commercial product.

Fresher data

This is a static snapshot from 2026-09-03 and it will not be updated. It ages from the day it was published, which is fine for research and useless for anything operational.

The scrapers that produced it are public, and the same schema comes back live:

Code written against this file works unchanged against fresh data.