Team Ai
20 results

scraper

Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes542 downloads21h agoHugging FaceTaventix /proxy-scraper-snapshots Verified free proxies – daily snapshots One Parquet file per UTC day with every public HTTP, SOCKS4 and SOCKS5 proxy that passed all checks in the first hourly run of that day. Collected by proxy-scraper, which pulls from 700+ public lists and keeps only proxies that actually relay traffic. from datasets import load_dataset ds = load_dataset("Taventix/proxy-scraper-snapshots", split="train") df = ds.to_pandas() # residential HTTPS exits that got through to Reddit, by country… See the full description on the dataset page: https://huggingface.co/datasets/Taventix/proxy-scraper-snapshots.tabular10K<n<100K0 likes383 downloads16h agoHugging Facerbtrprjkt /scrape_residueThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue.tabularrobotics10K<n<100K0 likes348 downloads16d agoHugging Faceendomorphosis /legal_scrapers JusticeDAO Legal Scrapers Research collectors that download official legislative sources only (national gazettes and official open-data portals). Intended for building reproducible legal-text corpora, not for production legal research products. Not legal advice. The official gazette of each jurisdiction prevails. License AGPL-3.0 for the collector scripts in this repository. Contents scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.texttext-retrievaln<1K0 likes244 downloads1mo agoHugging Facerbtrprjkt /scrape_residue_20260916_110935This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue_20260916_110935.tabularrobotics1K<n<10K0 likes195 downloads24d agoHugging Facemkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes187 downloads1mo agoHugging Face