FreshCrawl/g2-software-reviews
G2 Software Reviews 111,441 B2B software reviews from G2, covering the 79 most-reviewed products, spanning 2012 to 2026. The largest public G2 review corpus by a wide margin. Before this, the biggest available was a sample of under 1,000 rows. What is in here that is not in other review datasets A structured pros-and-cons split on 35,137 reviews. G2 asks "what do you like best" and "what do you dislike" as separate prompts, so those are separate columns rather… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/g2-software-reviews.
G2 Software Reviews
111,441 B2B software reviews from G2, covering the 79 most-reviewed products, spanning 2012 to 2026.
The largest public G2 review corpus by a wide margin. Before this, the biggest available was a sample of under 1,000 rows.
What is in here that is not in other review datasets
- A structured pros-and-cons split on 35,137 reviews. G2 asks "what do you like best" and "what do you dislike" as separate prompts, so those are separate columns rather than something a model has to extract. Same for
problems_solved. - Eight labelled solicitation channels. Not a boolean.
Organic,G2 invite,Seller invite,G2 invite on behalf of seller,In-app,G2 Gives Campaign,Thank You page,Organic Review from User Profile. You can distinguish platform-solicited from vendor-solicited, which almost no review dataset lets you do. - Reviewer verification method on 16,122 reviews:
Business EmailorLinkedIn. G2 tells you how it verified the reviewer. - Markdown-rendered review on every row, in addition to the plain text.
- 14 years of history, 2012-08-30 to 2026-09-03.
At a glance
A result worth knowing before you model this
G2 invite reviews average 3.61 stars against 4.35 for organic ones, a gap of nearly three quarters of a star. Seller invite sits at 4.10.
Products do not solicit reviews at random, though, so part of that gap could simply be which products run invite campaigns. Comparing each channel against organic reviews of the same product tests that:
Controlling for product shrinks the gap by about a third but does not remove it. G2-invited reviews still run roughly half a star below organic reviews of the same product, and are higher in only 6 of 18 products.
Two honest caveats. Only 18 to 27 products carry 30 or more reviews on both sides, so the magnitude is unsettled. And the same comparison run across the full 6,120-product cache this was drawn from gives a much weaker effect, which suggests the pattern is specific to heavily-reviewed products rather than universal.
Either way, product_slug matters as much as the text columns: the uncontrolled comparison overstates the effect by a third.
Why this is useful for ML
- Aspect-based sentiment with real labels. 35,137 reviews where the positive and negative aspects were written into separate boxes by the author. That is supervision you normally have to annotate.
- Rating regression over 14 years of the same products, so temporal drift is directly observable.
- Solicitation-bias research, with the product-level grouping needed to do it correctly.
- Long-form B2B English, median 317 characters of plain text plus a 650-character markdown rendering.
from datasets import load_dataset
ds = load_dataset("FreshCrawl/g2-software-reviews", split="train")
print(ds)Fields
Privacy: what is redacted and why it is still a column
G2 shows a reviewer's first name and last initial, plus a profile photo. Both are omitted here.
reviewer_name is kept as a column set to [redacted] rather than deleted, so the schema honestly reflects what the source carries. reviewer_has_avatar keeps the analytical signal (does having a photo correlate with anything) without shipping the image URL.
What remains is non-identifying segment data: job title, company-size band, industry. The live source additionally returns the display name, avatar URL and reviewer monogram; those are deliberately not reproduced here.
Review text is published as written and was not rewritten. It occasionally names a product, a company or a colleague, as review text does.
Collection method
- Cached continuously from G2 between 2026-03-30 and 2026-09-03, exported 2026-09-03.
- Curation rule: every product with 1,000 or more cached reviews, which yields 79 products and 111,441 reviews. This is a deliberate, statable subset, not a census of G2 and not a random sample.
- Reviews per product are complete within the cache, so a product's rating distribution here is not truncated the way a "most recent N" cut would be.
Limitations and bias
- Popular products only. The 1,000-review threshold selects for widely adopted software. Ratings here will not represent the long tail of small G2 listings.
- Mean rating is 4.31 of 5 and heavily left-skewed, as review corpora are. Accuracy on a naive classifier is misleading; use balanced metrics.
- The structured split covers 31.5%, not all rows.
has_structured_splitis provided so you can filter cleanly rather than discovering it mid-analysis. - Coverage varies by column and is stated per field above.
reviewer_industryat 6.7% andreview_urlat 24.9% are thin, andreviewer_titlecovers half the rows because G2 supplies a company-size band instead of a job title on the rest. - Engagement and rating are a single snapshot. Reviews span 14 years, so review age is a live confounder.
- English only.
Licence
Review text remains the intellectual property of its individual authors and of G2. This compilation is published for research and educational use. Cite the dataset and link back if you publish work based on it. Do not redistribute the raw file as a commercial product.
Fresher data
This is a static snapshot from 2026-09-03 and it will not be updated. It ages from the day it was published, which is fine for research and useless for anything operational.
The scrapers that produced it are public and return the same fields live, including the ones redacted here:
- G2 Reviews Scraper: any G2 product, star filters, full reviewer fields
- Software Review Scraper: the same shape across G2, Capterra, TrustRadius and Gartner in one run
Code written against this file works unchanged against fresh data.
