Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Taventix /proxy-scraper-snapshots Verified free proxies – daily snapshots One Parquet file per UTC day with every public HTTP, SOCKS4 and SOCKS5 proxy that passed all checks in the first hourly run of that day. Collected by proxy-scraper, which pulls from 700+ public lists and keeps only proxies that actually relay traffic. from datasets import load_dataset ds = load_dataset("Taventix/proxy-scraper-snapshots", split="train") df = ds.to_pandas() # residential HTTPS exits that got through to Reddit, by country… See the full description on the dataset page: https://huggingface.co/datasets/Taventix/proxy-scraper-snapshots.tabular10K<n<100K0 likes383 downloads22h agoHugging Face02rbtrprjkt /scrape_residueThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue.tabularrobotics10K<n<100K0 likes348 downloads16d agoHugging Face03rbtrprjkt /scrape_residue_20260916_110935This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue_20260916_110935.tabularrobotics1K<n<10K0 likes195 downloads24d agoHugging Face04reapxdev /finra-brokercheck-scraper FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows. Rows in this dataset 1,430 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.tabular1K<n<10K0 likes161 downloads2mo agoHugging Face05logiover /defillama-protocols-scraper-sample-data DefiLlama Protocols Scraper Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape. What the actor scrapes 🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.tabularn<1K0 likes137 downloads5mo agoHugging Face06reapxdev /tvmaze-scraper TVmaze Scraper · TV Shows, Episodes, Casts & Networks Scrape TV shows, episode details, cast members, ratings, genres, and network broadcast data from TVmaze's public database. Pay-per-event pricing per show record. Rows in this dataset 1,205 Fields 32 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/tvmaze-scraper.image1K<n<10K1 likes128 downloads2mo agoHugging Face07reapxdev /shopify-store-products-scraper Shopify Store Scraper Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser. Rows in this dataset 11,720 Fields 33 Collector runs behind it 72 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages Run the collector yourself https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.image10K<n<100K0 likes97 downloads2mo agoHugging Face08reapxdev /greenhouse-jobs-scraper Greenhouse Jobs Scraper Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL. Rows in this dataset 14,091 Fields 36 Collector runs behind it 61 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.tabular10K<n<100K0 likes87 downloads2mo agoHugging Face09reapxdev /kalshi-scraper Kalshi Scraper · Event Contracts, Markets, Prices & Volume Scrape Kalshi prediction markets, event contracts, option pricing, order book quotes, trading volume, open interest, and resolution rules. Export structured JSON, CSV, or Excel data. Rows in this dataset 7,649 Fields 34 Collector runs behind it 49 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/kalshi-scraper.tabular1K<n<10K1 likes86 downloads2mo agoHugging Face10logiover /geckoterminal-dex-pools-scraper-sample-data GeckoTerminal DEX Pools Scraper Scrape live DEX liquidity pools from GeckoTerminal across 100+ blockchains — price, FDV, market cap, liquidity, 1h/6h/24h volume, price change and transaction counts. Thousands of pools per run. Schedule it for a continuously fresh on-chain feed. What the actor scrapes 🦎 GeckoTerminal DEX Pools Scraper — Scrape On-Chain DEX Liquidity & Price Data Scrape live DEX liquidity pools from GeckoTerminalacross 100+ blockchains — Ethereum… See the full description on the dataset page: https://huggingface.co/datasets/logiover/geckoterminal-dex-pools-scraper-sample-data.tabularn<1K0 likes73 downloads5mo agoHugging Face11logiover /lagou-tech-jobs-scraper-sample-data Lagou Tech Jobs Scraper (拉勾网) Extract thousands of tech job listings from Lagou.com (拉勾网), China's largest IT recruitment platform. Scrape salary ranges, tech stacks, company details, funding stages, and more from ByteDance, Alibaba, Tencent, Baidu, and 100,000+ Chinese tech companies. No browser needed — fast, cheap, scalable. What the actor scrapes Lagou Tech Jobs Scraper (拉勾网) — Scrape China Tech Jobs, Salaries & Company Data Scrape Lagou.com (拉勾网), China's #1… See the full description on the dataset page: https://huggingface.co/datasets/logiover/lagou-tech-jobs-scraper-sample-data.tabularn<1K0 likes62 downloads5mo agoHugging Face12logiover /linkedin-top-content-scraper-sample-data LinkedIn Top Content & Top Voices Scraper Scrapes LinkedIn's public Top Content directory to extract curated high-engagement posts and Top Voice influencers across 40+ categories. Get post text, author profiles, follower counts, reaction metrics, and Top Voice badges. No login, no cookies, no account ban risk. $2 per 1,000 posts. What the actor scrapes LinkedIn Top Content & Top Voices Scraper Scrape LinkedIn's public Top Content directory — a curated archive of… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-top-content-scraper-sample-data.tabularn<1K0 likes56 downloads5mo agoHugging Face13logiover /usaspending-gov-scraper-sample-data USASpending.gov Federal Awards Scraper Scrape US federal contracts, grants and awards from the official USASpending.gov API — no login, no API key, no blocking. Award ID, recipient, amount, agency, dates and place of performance. Filter by type, date and keyword. Hundreds of thousands of awards per run. What the actor scrapes 🏛️ USASpending.gov Federal Awards Scraper — US Contracts, Grants & Awards to JSON & CSV Scrape US federal contracts, grants, loans and… See the full description on the dataset page: https://huggingface.co/datasets/logiover/usaspending-gov-scraper-sample-data.tabularn<1K0 likes53 downloads5mo agoHugging Face14reapxdev /boardgamegeek-scraper BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing. Rows in this dataset 437 Fields 21 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.imagen<1K0 likes51 downloads2mo agoHugging Face15reapxdev /open-library-scraper Open Library Scraper · Books, Authors, Editions & Subjects Scrape Open Library books, authors, subjects, editions, and metadata via Open Library API. Fast HTTP scraper charging per returned record with tiered pricing. Rows in this dataset 3,880 Fields 21 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/open-library-scraper/ — 3,880 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/open-library-scraper.tabular1K<n<10K0 likes47 downloads2mo agoHugging Face16rbtrprjkt /tool-scraper_daniel_20260827_124546This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-scraper_daniel_20260827_124546.tabularrobotics10K<n<100K0 likes47 downloads1mo agoHugging Face17reapxdev /shopify-app-store-scraper Shopify App Store Scraper List Shopify App Store apps and get one structured row per app, with developer, star rating, review count, every advertised pricing plan and the app's rank in its category. Rows in this dataset 3,880 Fields 26 Collector runs behind it 87 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/shopify-app-store-scraper/ — 1,625 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-app-store-scraper.image1K<n<10K1 likes39 downloads2mo agoHugging Face18reapxdev /github-repo-scraper GitHub Repo Scraper · Repositories, Stars, Topics & Languages Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication. Rows in this dataset 2,481 Fields 27 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.tabular1K<n<10K0 likes39 downloads2mo agoHugging Face19logiover /openstreetmap-business-poi-scraper-sample-data OpenStreetMap Business & POI Scraper Scrape businesses and points of interest from OpenStreetMap via Overpass API. Extract name, address, phone, website, opening hours and GPS coordinates for any city worldwide. Free alternative to Google Maps API. No API key needed. What the actor scrapes 🗺️ OpenStreetMap Business & POI Scraper — Scrape Businesses & Points of Interest, No API Key Scrape businesses and points of interest from OpenStreetMap using the free… See the full description on the dataset page: https://huggingface.co/datasets/logiover/openstreetmap-business-poi-scraper-sample-data.tabularn<1K0 likes38 downloads5mo agoHugging Face20logiover /defillama-yields-scraper-sample-data DefiLlama Yields Scraper Scrape DeFi yield & APY pools from DefiLlama — APY, TVL, base/reward yield, 1d/7d/30d APY trend, impermanent-loss risk and volume for 20,000+ pools across every chain. Filter by chain, protocol, TVL and APY. Schedule it daily to track the best yields. What the actor scrapes 💰 DefiLlama Yields Scraper — DeFi APY & TVL Pool Data Across All Chains Scrape DeFi yield and APY pools from DefiLlama, the most trusted DeFi data source. This Apify… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-yields-scraper-sample-data.tabularn<1K0 likes37 downloads5mo agoHugging Face21reapxdev /app-store-reviews-scraper App Store Reviews Scraper Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key. Rows in this dataset 23,048 Fields 43 Collector runs behind it 88 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages Run the collector yourself https://apify.com/reapx/app-store-reviews-scraper What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.image10K<n<100K0 likes37 downloads2mo agoHugging Face22logiover /otodom-pl-scraper-polish-real-estate-data-sample-data Otodom.pl Scraper — Polish Real Estate Data Extractor Scrape real estate listings from Otodom.pl — Poland's #1 property portal. Extract apartments, houses, plots and commercial properties by location, price (PLN), area, rooms and market. Direct NEXT_DATA access (no browser): price, m², czynsz, GPS, address, agency and image gallery per ad. What the actor scrapes Otodom.pl Scraper — Scrape Polish Real Estate Listings & Prices Scrape property listings from… See the full description on the dataset page: https://huggingface.co/datasets/logiover/otodom-pl-scraper-polish-real-estate-data-sample-data.tabularn<1K0 likes34 downloads5mo agoHugging Face23logiover /sec-edgar-form-d-scraper-sample-data SEC EDGAR Form D Scraper - Startup Funding Leads Scrape SEC EDGAR Form D filings to find recently funded startups. Extract company name, funding amount, CEO/director names, contact info, industry & more. Perfect for B2B sales prospecting & investor research. What the actor scrapes SEC EDGAR Form D Scraper — Find Funded Startups & B2B Leads Scrape SEC EDGAR Form D (Regulation D) filings and turn public startup-funding disclosures into structured, actionable lead… See the full description on the dataset page: https://huggingface.co/datasets/logiover/sec-edgar-form-d-scraper-sample-data.tabularn<1K0 likes33 downloads5mo agoHugging Face24logiover /greenhouse-job-board-scraper-sample-data Greenhouse Job Board API — Jobs, Departments & Offices Unofficial Greenhouse Job Board API in one Apify actor. Scrape jobs, full descriptions, departments and offices from any company on Greenhouse — Airbnb, Stripe, Anthropic, Mistral AI, Doctolib, Datadog, Notion. Pure HTTP, no auth, parallel batch. For HR tech, ATS, lead gen and AI agents. What the actor scrapes 🌱 Greenhouse Job Board API — Scrape Tech Jobs, Departments & Offices The unofficial Greenhouse Job… See the full description on the dataset page: https://huggingface.co/datasets/logiover/greenhouse-job-board-scraper-sample-data.tabularn<1K0 likes31 downloads5mo agoHugging Face25reapxdev /arxiv-papers-scraper arXiv Papers Scraper Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window. Rows in this dataset 21,722 Fields 26 Collector runs behind it 92 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.tabular10K<n<100K0 likes31 downloads2mo agoHugging Face26reapxdev /crossref-scraper Crossref Scraper · DOI Metadata, Authors, Journals & Citations Scrape scholarly DOI metadata, works, journal articles, authors, citations, funding, and licenses from the Crossref REST API. Fast HTTP scraper with pay-per-event pricing. Rows in this dataset 2,492 Fields 29 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/crossref-scraper/ — 2,492 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/crossref-scraper.tabular1K<n<10K0 likes30 downloads2mo agoHugging Face27reapxdev /pypi-scraper PyPI Scraper · Python Packages, Releases, Authors & Licences Scrape Python packages, release histories, dependencies, authors, maintainers, licenses, and download statistics from PyPI. HTTP only, pay-per-event pricing. Rows in this dataset 451 Fields 27 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/pypi-scraper.tabularn<1K1 likes30 downloads2mo agoHugging Face28bensblueprints /atomic-scraper-leads Atomic Scraper Leads Public business listings scraped from Google Maps for leads.benjaminboyce.com. This dataset is an export of the leads table from the Atomic / gmaps-scraper-suite pipeline. Each row is a local business listing with contact and location fields plus scrape metadata. Freshness Use scraped_at (UTC timestamps) as the freshness signal. This snapshot was exported on 2026-08-25. Oldest scraped_at: 2026-08-09 Newest scraped_at: 2026-08-25 Listings… See the full description on the dataset page: https://huggingface.co/datasets/bensblueprints/atomic-scraper-leads.tabulartabular-to-text10K<n<100K0 likes30 downloads2mo agoHugging Face29logiover /coingecko-derivatives-scraper-sample-data CoinGecko Derivatives Scraper Scrape 22,000+ crypto derivative tickers from CoinGecko in one run — price, 24h change, funding rate, open interest, basis, spread and 24h volume across every derivatives exchange. Schedule it for a continuously fresh feed. What the actor scrapes 📑 CoinGecko Derivatives Scraper — Scrape Crypto Futures & Perpetual Tickers Scrape 22,000+ crypto derivative tickers from CoinGeckoin a single run — every perpetual and futures contract… See the full description on the dataset page: https://huggingface.co/datasets/logiover/coingecko-derivatives-scraper-sample-data.tabularn<1K0 likes29 downloads5mo agoHugging Face30reapxdev /apple-podcasts-scraper Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets. Rows in this dataset 2,189 Fields 20 Collector runs behind it 51 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.image1K<n<10K1 likes29 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.