Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes542 downloads4h agoHugging Face02mkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes187 downloads1mo agoHugging Face03reapxdev /finra-brokercheck-scraper FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows. Rows in this dataset 1,430 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.tabular1K<n<10K0 likes161 downloads2mo agoHugging Face04neptun-org /neptun.scraper Data in this dataset Docker & NPM Scraped using crawl4ai. The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl. The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps. GitHub Scraped using firecrawl. The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.textquestion-answering100K<n<1M1 likes144 downloads2y agoHugging Face05reapxdev /tvmaze-scraper TVmaze Scraper · TV Shows, Episodes, Casts & Networks Scrape TV shows, episode details, cast members, ratings, genres, and network broadcast data from TVmaze's public database. Pay-per-event pricing per show record. Rows in this dataset 1,205 Fields 32 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/tvmaze-scraper.image1K<n<10K1 likes128 downloads2mo agoHugging Face06reapxdev /shopify-store-products-scraper Shopify Store Scraper Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser. Rows in this dataset 11,720 Fields 33 Collector runs behind it 72 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages Run the collector yourself https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.image10K<n<100K0 likes97 downloads2mo agoHugging Face07reapxdev /greenhouse-jobs-scraper Greenhouse Jobs Scraper Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL. Rows in this dataset 14,091 Fields 36 Collector runs behind it 61 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.tabular10K<n<100K0 likes87 downloads2mo agoHugging Face08reapxdev /kalshi-scraper Kalshi Scraper · Event Contracts, Markets, Prices & Volume Scrape Kalshi prediction markets, event contracts, option pricing, order book quotes, trading volume, open interest, and resolution rules. Export structured JSON, CSV, or Excel data. Rows in this dataset 7,649 Fields 34 Collector runs behind it 49 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/kalshi-scraper.tabular1K<n<10K1 likes86 downloads2mo agoHugging Face09reapxdev /lever-jobs-scraper Lever Jobs Scraper · Job Postings, Teams, Locations & Salary Scrape open job postings, departments, locations, remote status, compensation and application URLs from any Lever company job board. Rows in this dataset 8,716 Fields 20 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a keyword list: a row exists… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/lever-jobs-scraper.text1K<n<10K0 likes60 downloads2mo agoHugging Face10reapxdev /boardgamegeek-scraper BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing. Rows in this dataset 437 Fields 21 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.imagen<1K0 likes51 downloads2mo agoHugging Face11reapxdev /grants-gov-scraper Grants.gov Scraper · Grant Opportunities, Agencies & Awards Scrape US federal grant opportunities, funding announcements, and agency award notices from Grants.gov by keyword, agency, category, eligibility, and status. Rows in this dataset 1,488 Fields 12 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/grants-gov-scraper.text1K<n<10K0 likes49 downloads2mo agoHugging Face12reapxdev /open-library-scraper Open Library Scraper · Books, Authors, Editions & Subjects Scrape Open Library books, authors, subjects, editions, and metadata via Open Library API. Fast HTTP scraper charging per returned record with tiered pricing. Rows in this dataset 3,880 Fields 21 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/open-library-scraper/ — 3,880 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/open-library-scraper.tabular1K<n<10K0 likes47 downloads2mo agoHugging Face13reapxdev /doaj-scraper DOAJ Scraper · Open Access Journals, Articles & Publishers Scrape open access research articles, DOIs, authors, subjects, publishers, and abstracts from the Directory of Open Access Journals (DOAJ) API. Features pay-per-event pricing and automatic backoff. Rows in this dataset 1,720 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/doaj-scraper.text1K<n<10K0 likes40 downloads2mo agoHugging Face14reapxdev /shopify-app-store-scraper Shopify App Store Scraper List Shopify App Store apps and get one structured row per app, with developer, star rating, review count, every advertised pricing plan and the app's rank in its category. Rows in this dataset 3,880 Fields 26 Collector runs behind it 87 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/shopify-app-store-scraper/ — 1,625 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-app-store-scraper.image1K<n<10K1 likes39 downloads2mo agoHugging Face15reapxdev /github-repo-scraper GitHub Repo Scraper · Repositories, Stars, Topics & Languages Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication. Rows in this dataset 2,481 Fields 27 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.tabular1K<n<10K0 likes39 downloads2mo agoHugging Face16reapxdev /app-store-reviews-scraper App Store Reviews Scraper Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key. Rows in this dataset 23,048 Fields 43 Collector runs behind it 88 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages Run the collector yourself https://apify.com/reapx/app-store-reviews-scraper What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.image10K<n<100K0 likes37 downloads2mo agoHugging Face17reapxdev /sec-edgar-scraper SEC EDGAR Scraper Search SEC EDGAR by ticker, CIK, SIC industry code, or full-text phrase and get one structured row per filing, with company classification, period dates, direct document links and optional XBRL financials. Rows in this dataset 2,691 Fields 27 Collector runs behind it 64 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/sec-edgar-scraper/ — 771 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/sec-edgar-scraper.text1K<n<10K0 likes35 downloads2mo agoHugging Face18reapxdev /arxiv-papers-scraper arXiv Papers Scraper Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window. Rows in this dataset 21,722 Fields 26 Collector runs behind it 92 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.tabular10K<n<100K0 likes31 downloads2mo agoHugging Face19reapxdev /crossref-scraper Crossref Scraper · DOI Metadata, Authors, Journals & Citations Scrape scholarly DOI metadata, works, journal articles, authors, citations, funding, and licenses from the Crossref REST API. Fast HTTP scraper with pay-per-event pricing. Rows in this dataset 2,492 Fields 29 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/crossref-scraper/ — 2,492 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/crossref-scraper.tabular1K<n<10K0 likes30 downloads2mo agoHugging Face20reapxdev /pypi-scraper PyPI Scraper · Python Packages, Releases, Authors & Licences Scrape Python packages, release histories, dependencies, authors, maintainers, licenses, and download statistics from PyPI. HTTP only, pay-per-event pricing. Rows in this dataset 451 Fields 27 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/pypi-scraper.tabularn<1K1 likes30 downloads2mo agoHugging Face21reapxdev /apple-podcasts-scraper Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets. Rows in this dataset 2,189 Fields 20 Collector runs behind it 51 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.image1K<n<10K1 likes29 downloads2mo agoHugging Face22reapxdev /usaspending-scraper USAspending Scraper · Federal Contracts, Grants & Recipients Scrape US federal contracts, grants, and direct payments from USAspending.gov by agency, NAICS code, PSC code, award type, state, fiscal year, and dollar amount. Rows in this dataset 1,924 Fields 19 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/usaspending-scraper.text1K<n<10K0 likes29 downloads2mo agoHugging Face23reapxdev /wikipedia-scraper Wikipedia Scraper · Articles, Extracts, Categories & Links Extract Wikipedia articles, full text, lead summaries, categories, internal links, page views, and metadata across languages. HTTP-only Wikipedia API scraper for research, LLM datasets, and knowledge graphs. Rows in this dataset 2,500 Fields 11 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wikipedia-scraper.text1K<n<10K0 likes29 downloads2mo agoHugging Face24reapxdev /wordpress-plugins-scraper WordPress Plugins Scraper · Plugins, Installs & Ratings Scrape WordPress plugins directory by tags, search queries, author accounts, install bands, and rating filters. Extract ratings, active installs, tags, release details, and author links. Rows in this dataset 1,660 Fields 23 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wordpress-plugins-scraper.tabular1K<n<10K0 likes29 downloads2mo agoHugging Face25reapxdev /hackernews-scraper Hacker News Scraper · Stories, Comments, Points & Domains Scrape Hacker News stories, comments, Ask HN, Show HN, point thresholds, date ranges, and linked web domains via the official HN Search API by Algolia with rich search filters and domain extraction. Rows in this dataset 2,375 Fields 15 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/hackernews-scraper.tabular1K<n<10K1 likes27 downloads2mo agoHugging Face26reapxdev /pubmed-scraper PubMed Scraper · Papers, Authors, Journals & MeSH Terms Scrape academic research papers, authors, journals, MeSH terms, abstracts, and open-access metadata from NCBI PubMed API. Features HTTP backoff resilience and pay-per-event pricing. Rows in this dataset 1,847 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/pubmed-scraper.text1K<n<10K0 likes27 downloads2mo agoHugging Face27reapxdev /stackoverflow-scraper StackOverflow Scraper Scrape Stack Overflow questions, answers, tags and user profiles through the public Stack Exchange API. Filter by tag, score, date, accepted status and full-text search. No login, no browser. Rows in this dataset 16,719 Fields 47 Collector runs behind it 57 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/stackoverflow-scraper/ — 9,841 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/stackoverflow-scraper.tabular10K<n<100K0 likes26 downloads2mo agoHugging Face28reapxdev /docker-hub-scraper Docker Hub Scraper · Images, Tags, Pulls & Publishers Scrape Docker Hub container images, tags, pull counts, star counts, publishers, official vs community status, and categories. Rows in this dataset 2,843 Fields 15 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/docker-hub-scraper/ — 2,805 entity pages Run the collector yourself https://apify.com/reapx/docker-hub-scraper What… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/docker-hub-scraper.tabular1K<n<10K0 likes26 downloads2mo agoHugging Face29reapxdev /openalex-scraper OpenAlex Scraper · Works, Authors, Institutions & Citations Scrape scholarly works, papers, citations, authors, institutions, and open-access metadata from the OpenAlex API. Fast HTTP scraper charging per returned record with tiered pricing. Rows in this dataset 2,387 Fields 18 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/openalex-scraper/ — 2,387 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/openalex-scraper.tabular1K<n<10K1 likes26 downloads2mo agoHugging Face30reapxdev /flippa-scraper Flippa Scraper · Website, App & Online Business Listings Scrape website, mobile app, domain name and digital asset listings from Flippa. Filter by property type, price range, revenue, profit, or search terms. Rows in this dataset 2,301 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a keyword… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/flippa-scraper.tabular1K<n<10K0 likes24 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.