datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Web_Scraper_Dataproxy-scraper-snapshots
Verified free proxies – daily snapshots
One Parquet file per UTC day with every public HTTP, SOCKS4 and SOCKS5 proxy that passed all checks in the
first hourly run of that day. Collected by proxy-scraper,
which pulls from 700+ public lists and keeps only proxies that actually relay traffic.
from datasets import load_dataset
ds = load_dataset("Taventix/proxy-scraper-snapshots", split="train")
df = ds.to_pandas()
# residential HTTPS exits that got through to Reddit, by country… See the full description on the dataset page: https://huggingface.co/datasets/Taventix/proxy-scraper-snapshots.korean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.gleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.finra-brokercheck-scraper
FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures
Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows.
Rows in this dataset
1,430
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.neptun.scraper
Data in this dataset
Docker & NPM
Scraped using crawl4ai.
The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl.
The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps.
GitHub
Scraped using firecrawl.
The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.defillama-protocols-scraper-sample-data
DefiLlama Protocols Scraper
Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape.
What the actor scrapes
🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.tvmaze-scraper
TVmaze Scraper · TV Shows, Episodes, Casts & Networks
Scrape TV shows, episode details, cast members, ratings, genres, and network broadcast data from TVmaze's public database. Pay-per-event pricing per show record.
Rows in this dataset
1,205
Fields
32
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/tvmaze-scraper.web-scraper-dataset
Web Scraper Dataset
Whole-site crawl of 650+ domains across ~70 categories (news, wikis, fanfiction,
education, NSFW, social, gaming, code forges, papers, shadow libraries and more).
Common Crawl first (with retry/backoff), live crawlers (wget2/Playwright/Selenium/
requests) as fallback, uncapped per-domain page counts.
Files
MASTER_text.parquet — crawled page text
(id, modality, language, source_domain, category, title, url, text, text_length, crawl_date)… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/web-scraper-dataset.shopify-store-products-scraper
Shopify Store Scraper
Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser.
Rows in this dataset
11,720
Fields
33
Collector runs behind it
72
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages
Run the collector yourself
https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.greenhouse-jobs-scraper
Greenhouse Jobs Scraper
Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL.
Rows in this dataset
14,091
Fields
36
Collector runs behind it
61
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.kalshi-scraper
Kalshi Scraper · Event Contracts, Markets, Prices & Volume
Scrape Kalshi prediction markets, event contracts, option pricing, order book quotes, trading volume, open interest, and resolution rules. Export structured JSON, CSV, or Excel data.
Rows in this dataset
7,649
Fields
34
Collector runs behind it
49
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/kalshi-scraper.instagram-scraperIndia-Lok-Sabha-Debates-Dataset-Scrapergeckoterminal-dex-pools-scraper-sample-data
GeckoTerminal DEX Pools Scraper
Scrape live DEX liquidity pools from GeckoTerminal across 100+ blockchains — price, FDV, market cap, liquidity, 1h/6h/24h volume, price change and transaction counts. Thousands of pools per run. Schedule it for a continuously fresh on-chain feed.
What the actor scrapes
🦎 GeckoTerminal DEX Pools Scraper — Scrape On-Chain DEX Liquidity & Price Data Scrape live DEX liquidity pools from GeckoTerminalacross 100+ blockchains — Ethereum… See the full description on the dataset page: https://huggingface.co/datasets/logiover/geckoterminal-dex-pools-scraper-sample-data.discord-messages
Discord Messages Dataset
Description
This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains.
The data is formatted as plain text with one message per line, making it ideal for:
Language model pre-training
Fine-tuning chatbots
Sentiment analysis
Toxicity detection
Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.lagou-tech-jobs-scraper-sample-data
Lagou Tech Jobs Scraper (拉勾网)
Extract thousands of tech job listings from Lagou.com (拉勾网), China's largest IT recruitment platform. Scrape salary ranges, tech stacks, company details, funding stages, and more from ByteDance, Alibaba, Tencent, Baidu, and 100,000+ Chinese tech companies. No browser needed — fast, cheap, scalable.
What the actor scrapes
Lagou Tech Jobs Scraper (拉勾网) — Scrape China Tech Jobs, Salaries & Company Data Scrape Lagou.com (拉勾网), China's #1… See the full description on the dataset page: https://huggingface.co/datasets/logiover/lagou-tech-jobs-scraper-sample-data.lever-jobs-scraper
Lever Jobs Scraper · Job Postings, Teams, Locations & Salary
Scrape open job postings, departments, locations, remote status, compensation and application URLs from any Lever company job board.
Rows in this dataset
8,716
Fields
20
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a keyword list: a row exists… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/lever-jobs-scraper.linkedin-top-content-scraper-sample-data
LinkedIn Top Content & Top Voices Scraper
Scrapes LinkedIn's public Top Content directory to extract curated high-engagement posts and Top Voice influencers across 40+ categories. Get post text, author profiles, follower counts, reaction metrics, and Top Voice badges. No login, no cookies, no account ban risk. $2 per 1,000 posts.
What the actor scrapes
LinkedIn Top Content & Top Voices Scraper Scrape LinkedIn's public Top Content directory — a curated archive of… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-top-content-scraper-sample-data.usaspending-gov-scraper-sample-data
USASpending.gov Federal Awards Scraper
Scrape US federal contracts, grants and awards from the official USASpending.gov API — no login, no API key, no blocking. Award ID, recipient, amount, agency, dates and place of performance. Filter by type, date and keyword. Hundreds of thousands of awards per run.
What the actor scrapes
🏛️ USASpending.gov Federal Awards Scraper — US Contracts, Grants & Awards to JSON & CSV Scrape US federal contracts, grants, loans and… See the full description on the dataset page: https://huggingface.co/datasets/logiover/usaspending-gov-scraper-sample-data.boardgamegeek-scraper
BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics
Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing.
Rows in this dataset
437
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.grants-gov-scraper
Grants.gov Scraper · Grant Opportunities, Agencies & Awards
Scrape US federal grant opportunities, funding announcements, and agency award notices from Grants.gov by keyword, agency, category, eligibility, and status.
Rows in this dataset
1,488
Fields
12
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/grants-gov-scraper.uk-companies-house-bulk-scraper-sample-data
UK Companies House Bulk Scraper
Extract UK companies from the official Companies House register. 5M+ businesses by industry (SIC code), location, age, status. Monthly bulk snapshot. No API key. B2B lead-gen, KYC, recruitment, market research. $1/1K companies.
What the actor scrapes
UK Companies House Bulk Scraper — 5M+ UK Companies, No API Key Extract UK companies from the official Companies House government register — 5+ million active and dissolved businesses… See the full description on the dataset page: https://huggingface.co/datasets/logiover/uk-companies-house-bulk-scraper-sample-data.open-library-scraper
Open Library Scraper · Books, Authors, Editions & Subjects
Scrape Open Library books, authors, subjects, editions, and metadata via Open Library API. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
3,880
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/open-library-scraper/ — 3,880 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/open-library-scraper.doaj-scraper
DOAJ Scraper · Open Access Journals, Articles & Publishers
Scrape open access research articles, DOIs, authors, subjects, publishers, and abstracts from the Directory of Open Access Journals (DOAJ) API. Features pay-per-event pricing and automatic backoff.
Rows in this dataset
1,720
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/doaj-scraper.shopify-app-store-scraper
Shopify App Store Scraper
List Shopify App Store apps and get one structured row per app, with developer, star rating, review count, every advertised pricing plan and the app's rank in its category.
Rows in this dataset
3,880
Fields
26
Collector runs behind it
87
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-app-store-scraper/ — 1,625 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-app-store-scraper.github-repo-scraper
GitHub Repo Scraper · Repositories, Stars, Topics & Languages
Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication.
Rows in this dataset
2,481
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.openstreetmap-business-poi-scraper-sample-data
OpenStreetMap Business & POI Scraper
Scrape businesses and points of interest from OpenStreetMap via Overpass API. Extract name, address, phone, website, opening hours and GPS coordinates for any city worldwide. Free alternative to Google Maps API. No API key needed.
What the actor scrapes
🗺️ OpenStreetMap Business & POI Scraper — Scrape Businesses & Points of Interest, No API Key Scrape businesses and points of interest from OpenStreetMap using the free… See the full description on the dataset page: https://huggingface.co/datasets/logiover/openstreetmap-business-poi-scraper-sample-data.defillama-yields-scraper-sample-data
DefiLlama Yields Scraper
Scrape DeFi yield & APY pools from DefiLlama — APY, TVL, base/reward yield, 1d/7d/30d APY trend, impermanent-loss risk and volume for 20,000+ pools across every chain. Filter by chain, protocol, TVL and APY. Schedule it daily to track the best yields.
What the actor scrapes
💰 DefiLlama Yields Scraper — DeFi APY & TVL Pool Data Across All Chains Scrape DeFi yield and APY pools from DefiLlama, the most trusted DeFi data source. This Apify… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-yields-scraper-sample-data.app-store-reviews-scraper
App Store Reviews Scraper
Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key.
Rows in this dataset
23,048
Fields
43
Collector runs behind it
88
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages
Run the collector yourself
https://apify.com/reapx/app-store-reviews-scraper
What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.
