Fox-user/crawler-tool
crawler-tool
crawler-tool 是一个高中生 Python 长期学习项目的最小可运行版本(MVP)。
当前阶段的目标不是做复杂爬虫框架,而是先练习:
- Python 项目结构设计
- API 接口分析
- JSON 配置读取
- requests 请求接口
- JSON 数据解析和字段提取
- 命令行工具开发
- Streamlit 可视化小工具开发
核心思路:以后尽量通过修改 sites/*.json 配置文件适配不同网站/API,而不是每次都修改主程序代码。
项目结构
crawler-tool/
├─ app.py
├─ main.py
├─ requirements.txt
├─ README.md
├─ .gitignore
├─ sites/
│ ├─ example.json
│ └─ html_example.json
├─ samples/
│ └─ example_posts.json
├─ output/
│ └─ .gitkeep
└─ crawler_tool/
├─ __init__.py
├─ adaptive.py
├─ analyzer.py
├─ api_detector.py
├─ config.py
├─ data_discovery.py
├─ direct_extract.py
├─ downloader.py
├─ embedded_data.py
├─ extractor.py
├─ html_extractor.py
├─ runner.py
├─ selector_suggester.py
├─ smart_analyzer.py
└─ storage.py安装依赖
建议先创建虚拟环境:
python -m venv .venv
source .venv/bin/activateWindows PowerShell 可以使用:
python -m venv .venv
.venv\Scripts\Activate.ps1安装依赖:
pip install -r requirements.txt命令行运行
查看已有网站配置:
python main.py --list-sites抓取示例 API,并导出 TXT:
python main.py --site example --limit 5 --format txt抓取示例 API,并导出 JSON:
python main.py --site example --limit 5 --format json指定输出目录:
python main.py --site example --limit 5 --format json --output-dir output追加请求参数(必须是 JSON 对象字符串):
python main.py --site example --limit 5 --format json --extra-params '{"userId": 1}'运行 HTML 模式示例:
python main.py --site html_example --limit 3 --format json命令行智能分析公开网页并输出配置草稿:
python main.py --analyze-url https://example.comScrapling-like 轻量能力
在不破坏原有配置驱动结构的前提下,项目新增了一个轻量的 Scrapling-like 解析层,适合公开、静态 HTML 的学习型提取:
crawler_tool.page.Page/Element/Selection:支持 CSS、XPath、::text、::attr(href)、正则提取、文本搜索、相似元素匹配、父子兄弟节点导航。crawler_tool.fetchers.Fetcher/FetcherSession:用普通requests会话抓取公开页面并返回Page对象,支持批量抓取和礼貌延迟。- 现有
--extract-url命令扩展了--text-search、--regex-search、--similar-to,仍然保留原来的 CSS/XPath/输出文件参数。
示例:
from crawler_tool.fetchers import Fetcher
page = Fetcher.fetch("https://example.com")
print(page.title)
print(page.css("a::attr(href)").get_all())
print(page.find_by_text("Example").get())命令行示例:
python main.py --extract-url https://example.com --css-selector "h1::text"
python main.py --extract-url https://example.com --text-search Example
python main.py --extract-url https://example.com --similar-to h1说明:这里复刻的是 Scrapling 风格的安全解析和便捷查询能力,不包含登录绕过、验证码/反爬绕过、代理池、高频抓取或访问非公开数据等能力。
Streamlit 网页运行
启动网页:
streamlit run app.py页面支持:
- 下拉选择网站配置
- 输入
limit - 选择导出格式
txt/json - 输入可选的
extra params JSON - 点击“开始运行”
- 查看数据预览表格
- 下载导出文件
Smart Analyzer
Streamlit 页面中提供“智能网站分析器”区域。输入一个公开网页 URL 后,工具会请求页面并生成学习分析报告,包括:
- HTTP 状态码、最终 URL、网页标题、Meta description、页面大小
- 页面中的链接、图片、脚本、表格、
h1/h2/h3标题结构 script[type="application/ld+json"]、#__NEXT_DATA__、window.__INITIAL_STATE__、script[type="application/json"]等嵌入 JSON / 状态数据- 从已公开嵌入 JSON 中识别重复 list-of-dict 数据候选,给出路径、字段和样例预览
- 疑似 API 线索,例如
/api/、api.、.json、fetch(、axios、XMLHttpRequest、graphql、endpoint - 重复卡片、列表项、表格行、正文容器等 CSS selector 建议
api/html/manual推荐模式、页面类型判断、HTML 提取预览和sites/*.json配置草稿- 配置草稿下载按钮,文件名使用
{site_name}.json
Smart Analyzer 只用于公开网页学习分析。它不会处理登录、验证码、绕过限制、反爬保护、付费墙或高频抓取;也不会尝试获取非公开、未授权或受访问控制保护的数据。生成的配置草稿也不保证可以直接运行,仍然需要人工确认 api_url、data_path、fields 或 selectors。
HTML 模式
sites/*.json 支持 source_type 字段;如果不写,默认是 API 模式。
- API 模式适合已经找到的 JSON 接口,通过
api_url、data_path、fields提取记录。 - HTML 模式适合静态网页、列表页、表格页,通过
page_url和 CSSselectors提取标题、文本、链接和图片。 - HTML 模式的
selectors需要用户自己确认;Smart Analyzer 只给出建议。
HTML 模式示例位于 sites/html_example.json,使用 https://example.com:
{
"site_name": "html_example",
"display_name": "HTML 示例页面",
"source_type": "html",
"page_url": "https://example.com",
"method": "GET",
"headers": {
"User-Agent": "Mozilla/5.0"
},
"selectors": {
"items": "body",
"title": "h1",
"links": "a",
"images": "img"
},
"output_name": "html_example_result"
}示例配置
当前示例配置位于 sites/example.json,使用公开测试 API:
https://jsonplaceholder.typicode.com/posts配置内容说明:
{
"site_name": "example",
"display_name": "示例 API:JSONPlaceholder Posts",
"source_type": "api",
"api_url": "https://jsonplaceholder.typicode.com/posts",
"method": "GET",
"headers": {
"User-Agent": "Mozilla/5.0"
},
"params": {},
"sample_file": "samples/example_posts.json",
"data_path": "",
"fields": {
"id": "id",
"title": "title",
"body": "body"
},
"output_name": "example_posts"
}sample_file 离线兜底
网站配置可以填写可选字段 sample_file,指向仓库中的本地 JSON 样例文件:
{
"api_url": "https://jsonplaceholder.typicode.com/posts",
"sample_file": "samples/example_posts.json"
}运行时工具会先请求真实 api_url。如果外部 API 请求失败,并且配置了 sample_file,程序会读取本地样例 JSON 继续执行后续流程,因此仍然可以测试 JSON 解析、字段提取和 TXT/JSON 文件导出。使用离线样例时,命令行和 Streamlit 页面都会提示:
外部 API 请求失败,已使用离线示例数据。如果真实 API 请求失败,并且 sample_file 不存在、不是合法 JSON 或读取失败,程序会抛出包含真实 API 错误和离线样例错误的清楚错误信息。
怎么新增一个网站配置
- 在
sites/目录中新建一个 JSON 文件,例如sites/books.json。 - 复制
sites/example.json的结构。 - 修改这些字段:
site_name:配置名,建议和文件名一致。display_name:页面中显示的中文名称。source_type:可选,api或html;不写时默认api。api_url:API 模式接口地址。page_url:HTML 模式页面地址。method:目前支持GET和POST。headers:请求头。params:默认请求参数。sample_file:可选,本地离线样例 JSON 路径;真实 API 请求失败时用它兜底。data_path:接口返回 JSON 中,记录列表所在的位置。fields:要提取的字段。selectors:HTML 模式 CSS 选择器配置。output_name:导出文件名,不需要写后缀。
data_path 和 fields 的点路径
工具支持简单的点路径,例如:
data.items
result.list.0.title如果接口直接返回列表,data_path 可以写成空字符串 ""。
如果接口返回:
{
"data": {
"items": [
{"title": "第一条", "author": {"name": "小明"}}
]
}
}那么可以这样配置:
{
"data_path": "data.items",
"fields": {
"title": "title",
"author_name": "author.name"
}
}当前阶段说明
这个项目现在只做 MVP:
- 不引入 Scrapy
- 不引入 FastAPI
- 不使用数据库
- 不做复杂前端
先保证“配置 -> 请求 API/HTML -> 解析 JSON 或 HTML -> 提取字段 -> 导出文件”这条主流程能跑通。后续再逐步扩展分页、更多导出格式、任务历史记录等功能。
Scrapling 启发的安全功能子集
本项目现在加入了参考 Scrapling 思路实现的学习版功能,但刻意不实现高频访问、并发压测、代理池轮换、登录绕过、验证码/反爬绕过等能力。新增能力聚焦于“单次、低频、公开页面”的解析体验:
- 自适应 selector 指纹:HTML selector 找到元素后可通过
adaptive: true/auto_save: true保存轻量指纹;页面结构小幅变化后可用adaptive_key做相似元素兜底匹配。 - 更灵活的选择方式:字段 selector 支持 CSS、
fallback_selectors、XPath、文本包含、正则匹配和similar_to相似元素查找。 - 字段级提取配置:
selectors.fields可为每个字段配置selector、attr、all、extract_regex等选项。 - 单次 CLI 提取:无需创建
sites/*.json,可以像 Scrapling CLI 一样对一个公开 URL 做一次性正文、HTML 或 Markdown 风格文本提取。
示例 HTML 配置片段:
{
"source_type": "html",
"selectors": {
"items": {
"selector": ".product",
"fallback_selectors": ["article", "li"],
"adaptive": true,
"auto_save": true,
"adaptive_key": "demo:products"
},
"title": {
"selector": "h2",
"fallback_selectors": ["h3", ".title"],
"adaptive": true,
"adaptive_key": "demo:title"
},
"links": {"selector": "a", "attr": "href"},
"fields": {
"price": {"selector": ".price", "extract_regex": "([0-9.]+)"},
"tags": {"selector": ".tag", "all": true}
}
}
}单次命令行提取示例:
python main.py --extract-url https://example.com --css-selector body --extract-mode text
python main.py --extract-url https://example.com --css-selector body --extract-mode markdown --extract-output output/example.md
python main.py --extract-url https://example.com --xpath '//h1' --extract-mode text注意:这些功能只用于公开页面的学习分析和低频提取。请遵守目标网站的 robots.txt、服务条款和当地法律法规。
