Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hudsongouge /AoPS-Scrape AoPS-Scrape Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints. Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session. Splits Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split: Split Rows Notes deduplicated 29,964 One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tabularquestion-answering10K<n<100K2 likes8.9k downloads3mo agoHugging Face02SKT-NRS /GIT-SCRAPED 🚀 SKT-NRS / GIT-SCRAPED This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets. Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures. 📂 Repository Structure All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.1K<n<10K1 likes1.4k downloads4mo agoHugging Face03thuerey-group /apebench-scraped APEBench Scraped A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0. ⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.1 likes852 downloads2y agoHugging Face04eagle0504 /ysa-web-scrape-dataset-qa-formatted-small-versiontextn<1K1 likes787 downloads3y agoHugging Face05AbstractPhil /IMDB-PUBLIC-SCRAPED Hello World with Hugging Face Current Date: 2025-03-19 04:36:42.698271 So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later. It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks. I'll be working out the problems and getting the scraper working correctly at some point soon. 1 likes562 downloads5mo agoHugging Face06Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes542 downloads23h agoHugging Face07thuerey-group /apebench-scraped-old APEBench-scraped (old) All datasets scraped from the APEBench benchmark suite with the version used for the Neurips submission. Download Download without large files GIT_LFS_SKIP_SMUDGE=1 git clone git@hf.co:datasets/thuerey-group/apebench-scraped-old Afterwards, you can inspect the repository and download the files you need. For example, for 1d_diff_adv: git lfs install git lfs pull -I "data/1d_diff_adv*" Alternatively, you can download the entire repository with large… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped-old.0 likes480 downloads2y agoHugging Face08thonyyy /tatoeba-nusax-scrape-mt-concattext100M<n<1B0 likes460 downloads2y agoHugging Face09Taventix /proxy-scraper-snapshots Verified free proxies – daily snapshots One Parquet file per UTC day with every public HTTP, SOCKS4 and SOCKS5 proxy that passed all checks in the first hourly run of that day. Collected by proxy-scraper, which pulls from 700+ public lists and keeps only proxies that actually relay traffic. from datasets import load_dataset ds = load_dataset("Taventix/proxy-scraper-snapshots", split="train") df = ds.to_pandas() # residential HTTPS exits that got through to Reddit, by country… See the full description on the dataset page: https://huggingface.co/datasets/Taventix/proxy-scraper-snapshots.tabular10K<n<100K0 likes383 downloads18h agoHugging Face10trentmkelly /DrugHub-scrape DrugHub Market Snapshot, September 2026 A complete, text-only capture of the public listing, vendor, and review pages of DrugHub, a Monero-only darknet market operating since 2023. Everything here was visible to any visitor without an account. Doesn't include any images. Collected 16-17 September 2026. Enriched with model-derived labels (typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment table and the Enrichment section below. What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.tabular100K<n<1M5 likes366 downloads21d agoHugging Face11rbtrprjkt /scrape_residueThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue.tabularrobotics10K<n<100K0 likes348 downloads16d agoHugging Face12mondk /vi.wikipedia-scraped-dataSorry for not having English. bru 4 likes309 downloads2mo agoHugging Face13Asib27 /github_repo_scrapedScraped python packages from github. 0 likes248 downloads2y agoHugging Face14endomorphosis /legal_scrapers JusticeDAO Legal Scrapers Research collectors that download official legislative sources only (national gazettes and official open-data portals). Intended for building reproducible legal-text corpora, not for production legal research products. Not legal advice. The official gazette of each jurisdiction prevails. License AGPL-3.0 for the collector scripts in this repository. Contents scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.texttext-retrievaln<1K0 likes244 downloads1mo agoHugging Face15Isamu136 /penetration_testing_scraped_dataset Dataset Card for "penetration_testing_scraped_dataset" More Information needed text100K<n<1M14 likes240 downloads3y agoHugging Face16scrapegraphai /scrapegraphai-100k ScrapeGraphAI-100k Dataset Summary ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.tabulartext-generation10K<n<100K27 likes240 downloads3mo agoHugging Face17sayurio /english-manhwa-scrape English Manhwa Scrape Dataset Request More ScrapesOrder Private Scrapes Note: The total files are about 150+GB and my internet connection is slow af. So I'll be uploading them in batches File Structure: files/ - Manhwa 1 Name - Chapter 1.cbz - Chapter 2.cbz - Chapter 3.cbz - ...... - Chapter n.cbz - Manhwa 2 Name - Chapter 1.cbz - Chapter 2.cbz - Chapter 3.cbz - ...... - Chapter n.cbz Due to maximum 10000 files in a repo limit of huggingface, I had to further… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/english-manhwa-scrape.image-to-text100K<n<1M2 likes227 downloads6mo agoHugging Face18rbtrprjkt /scrape_residue_20260916_110935This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue_20260916_110935.tabularrobotics1K<n<10K0 likes195 downloads24d agoHugging Face19cassanof /arxiv_cs_scrape0 likes194 downloads2y agoHugging Face20mkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes187 downloads1mo agoHugging Face21logiover /gleif-lei-scraper-sample-data GLEIF LEI Scraper Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run. What the actor scrapes 🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.textn<1K0 likes181 downloads5mo agoHugging Face22reapxdev /finra-brokercheck-scraper FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows. Rows in this dataset 1,430 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.tabular1K<n<10K0 likes161 downloads2mo agoHugging Face23neptun-org /neptun.scraper Data in this dataset Docker & NPM Scraped using crawl4ai. The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl. The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps. GitHub Scraped using firecrawl. The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.textquestion-answering100K<n<1M1 likes144 downloads2y agoHugging Face24reapxdev /ashby-jobs-scraper Ashby Jobs Scraper · Job Board Postings, Teams & Compensation Scrape Ashby job board listings across tech and high-growth companies. Extract job titles, departments, teams, employment types, locations, remote status, published dates, application URLs, and full salary/compensation tiers. Export clean JSON or CSV. Rows in this dataset 3,194 Fields 21 Collector runs behind it 52 Most recent observation 2026-08-03 What this is Every row here… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/ashby-jobs-scraper.1K<n<10K0 likes141 downloads2mo agoHugging Face25ryang2 /linkedin-job-scrape LinkedIn DS/ML Job Postings Daily snapshots of data science / machine learning job postings scraped from LinkedIn, one parquet file per scrape run. All splits share one schema (2023 → today); historical splits were migrated in 2026-07 — the original 8-column data is preserved at revision v1-schema. The same job_id recurs across splits (a posting stays live for days) — that's the longitudinal signal. For a unique-jobs view, dedup on job_id keeping the row with max scrape_dt.… See the full description on the dataset page: https://huggingface.co/datasets/ryang2/linkedin-job-scrape.text100K<n<1M0 likes139 downloads1mo agoHugging Face26logiover /defillama-protocols-scraper-sample-data DefiLlama Protocols Scraper Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape. What the actor scrapes 🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.tabularn<1K0 likes137 downloads5mo agoHugging Face27mkd-minju /korean_data_scraper_wikipedia Korean Data Scraper — Wikipedia Dump Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 wikipedia_dump 소스가 생성한 코퍼스입니다. 한국어 위키백과(ko.wikipedia.org) 덤프를 파싱하여 문서 본문 텍스트를 추출한 결과입니다. 스키마 파일당 1개 레코드(JSONL)이며, korean_data_scraper_kakaotalk와 동일한 공통 스키마를 따릅니다. {"id": "kowiki:70773", "source": "wikipedia_dump", "text": "..."} id는 kowiki:<문서 ID> 형태이며, text는 해당 위키백과 문서의 본문 텍스트입니다. 데이터 규모 및 한계 전체 32개 샤드(shard-00000~00031)로 구성되며, 총 용량은 약 2.24GB입니다. 이 중… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_wikipedia.text-generation1M<n<10M0 likes131 downloads1mo agoHugging Face28reapxdev /coingecko-scraper CoinGecko Scraper Scrape CoinGecko coin prices, market caps, volumes, supply and all-time highs in any of 63 currencies, plus exchange trust scores and volumes. No login, no API key. Rows in this dataset 9,713 Fields 68 Collector runs behind it 63 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/coingecko-scraper/ — 4,880 entity pages Run the collector yourself https://apify.com/reapx/coingecko-scraper What… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/coingecko-scraper.image1K<n<10K0 likes128 downloads2mo agoHugging Face29reapxdev /tvmaze-scraper TVmaze Scraper · TV Shows, Episodes, Casts & Networks Scrape TV shows, episode details, cast members, ratings, genres, and network broadcast data from TVmaze's public database. Pay-per-event pricing per show record. Rows in this dataset 1,205 Fields 32 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/tvmaze-scraper.image1K<n<10K1 likes128 downloads2mo agoHugging Face30sayurio /jugantor.com-scrape-bangla Jugantor News Archive (Bangla) Overview This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content. Purpose and Usage This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.imagetext-generation10K<n<100K1 likes112 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.