evalstate/diffusers-pr-api
0
1Metadata-Version: 2.42Name: slop-farmer3Version: 0.1.14Summary: GitHub-to-Hub data pipeline for transformers issue and PR triage research.5Requires-Python: >=3.13.56Description-Content-Type: text/markdown7Requires-Dist: duckdb>=1.2.28Requires-Dist: pyarrow>=18.0.09Requires-Dist: fastapi>=0.115.010Requires-Dist: huggingface_hub>=1.11.011Requires-Dist: pydantic>=2.1112Requires-Dist: PyYAML>=6.0.213Requires-Dist: rank-bm25>=0.2.214Requires-Dist: fast-agent-mcp>=0.6.1715Requires-Dist: uvicorn>=0.34.016Provides-Extra: dev17Requires-Dist: httpx>=0.28.0; extra == "dev"18Requires-Dist: pytest>=8.3.0; extra == "dev"19Requires-Dist: ruff>=0.11; extra == "dev"20Requires-Dist: ty>=0.0.23; extra == "dev"21Provides-Extra: llm22Requires-Dist: fast-agent-mcp>=0.6.16; python_full_version >= "3.13.5" and extra == "llm"23 24# slop-farmer25 26Pipeline for managing PR's in high volume GitHub repositories. 27 28Scrapes PR, Issue and Contributor data in to a dataset, performs analysis and publishes a dashboard.29 30The pipeline stages are:31 1. Scrape - Collect data from the Github Repository32 1. Contributor Report - Look at contributors recent history.33 1. Analyze - Cluster PRs and Issues on 34 1. Scope - Cluster PRs on overlapping repository areas.35 1. Dashboard Export - Export data in JSON format to populate a browsing dashboard36 1. Publish Dashboard - Build a dashboard and deploy it in a Hugging Face Space.37 38 39 40## Scrape41 42To run a scrape you need to configure:43 441. The GitHub Repository ID451. A valid GitHub PAT with API access.46 47`uv run slop-farmer scrape --repo huggingface/diffusers --output-dir runs/diffusers/data`48 49## Contributor Report50 51This scans the dataset for Contributors and provides a short profile of their recent public commit history and merged PR rate.52 53## Analyze54 55Cluster PRs and Issue Content. Choice of deterministic or LLM supplemented algorithm.56 57When `ranking_backend=hybrid`, analysis writes reusable LLM review cache entries under58`<snapshot>/analysis-state/`. If you enable YAML config setting59`analysis.cached_analysis: true`, `analyze` will automatically copy `analysis-state/`60forward from the previous snapshot when the new snapshot does not already have it, then61log a cache-hit summary for the run. This is useful for incremental scrapes where many62review units are unchanged and can safely reuse cached hybrid decisions.63 64## Scope65 66Cluster PRs by touched repository areas.67 68## Dashboard Export / Publish69 70Export the report, and publish a dashboard.71 72 73 74## Quickstart75 76```bash77uv run slop-farmer scrape \78 --repo huggingface/transformers \79 --output-dir data \80 --max-issues 200 \81 --max-prs 5082```83 84To publish a snapshot to the Hub:85 86```bash87uv run slop-farmer scrape \88 --repo huggingface/transformers \89 --output-dir data \90 --hf-repo-id burtenshaw/transformers-pr-slop-dataset \91 --publish92```93 94When `--publish` is used, `slop-farmer` now also generates and uploads new contributor reviewer artifacts by default:95 96- `new_contributors.parquet`97- `new-contributors-report.json`98- `new-contributors-report.md`99 100Use `--no-new-contributor-report` to skip them.101 102## Nightly incremental runs103 104The scraper now stores a local watermark at `data/state/watermark.json` and resumes from it by default when `--since` is not provided.105 106```bash107uv run slop-farmer scrape \108 --repo huggingface/transformers \109 --output-dir data \110 --fetch-timeline111```112 113On the first run, this creates a full snapshot. On later runs against the same `--output-dir`, it uses the last successful watermark, fetches only changed records, merges them into the previous snapshot locally, and writes a new full latest snapshot.114 115To ignore the watermark and force a fresh full run:116 117```bash118uv run slop-farmer scrape \119 --repo huggingface/transformers \120 --output-dir data \121 --no-resume122```123 124Authentication defaults:125 126- GitHub: `GITHUB_TOKEN`, then `gh auth token`127- Hugging Face: `HF_TOKEN`, otherwise existing `hf auth` login128 129## Canonical dataset upkeep130 131`dataset_id` is the canonical latest dataset repo.132 133Use the remote-first writer:134 135```bash136uv run slop-farmer --config configs/transformers.yaml refresh-dataset137```138 139Or submit the generic HF Job wrapper:140 141```bash142scripts/submit_dataset_job.sh143```144 145By default this creates a scheduled HF Job that:146 147- reads `CONFIG_PATH` (defaults to `configs/transformers.yaml`)148- refreshes `dataset_id` incrementally against the current Hub dataset state149- regenerates the new contributor report150- uploads the updated snapshot back to the dataset repo151 152Useful overrides:153 154```bash155# fire once immediately instead of creating a schedule156MODE=run scripts/submit_dataset_job.sh157 158# change the cron schedule159SCHEDULE="0 */6 * * *" scripts/submit_dataset_job.sh160 161# optionally mount a writable HF bucket for temp files162SCRATCH_BUCKET=evalstate/slop-farmer-scratch \163 scripts/submit_dataset_job.sh164```165 166Buckets are best treated here as optional scratch space via `TMPDIR`, not as the canonical167published dataset. The repo's local analysis and PR-scope tooling already knows how to168materialize versioned Hub **dataset repos**; it does not currently read HF buckets directly.169 170Compatibility wrappers remain available:171 172- `scripts/submit_transformers_dataset_job.sh`173- `scripts/submit_openclaw_dataset_job.sh`174 175For the current storage model and recommended modes, see176[`docs/data-architecture.md`](docs/data-architecture.md).177 178## Analyze a Hub dataset179 180You can analyze the published Hugging Face dataset directly without scraping GitHub again:181 182```bash183uv run slop-farmer analyze \ 184 --snapshot-dir eval_data/snapshots/gh-live-latest-1000x1000 \185 --ranking-backend hybrid \186 --model "gpt-5-mini?reasoning=low" \187 --output /tmp/gh-live-latest-1000x1000-hybrid.json188```189 190This materializes the dataset-viewer parquet export into a local snapshot cache under `eval_data/snapshots/` and writes `analysis-report.json` next to it.191 192Repo-local defaults for `analyze` can be stored in `pyproject.toml` under `[tool.slop-farmer.analyze]`. This repo currently defaults to:193 194- `dashboard-data.output-dir = "web/public/data"`195 196For repo-specific remote-first analysis, prefer a YAML config with `dataset_id`, e.g.:197 198```bash199uv run slop-farmer --config configs/openclaw.yaml analyze200```201 202## Cluster open PRs by code scope203 204You can also build holistic PR scope clusters from an existing snapshot:205 206```bash207uv run slop-farmer pr-scope \208 --snapshot-dir data/snapshots/20260324T150154Z209```210 211By default this writes `pr-scope-clusters.json` next to the snapshot.212 213## Merge duplicate PR clusters214 215List only the duplicate PR clusters that pass the mergeability gate:216 217```bash218uv run slop-farmer duplicate-prs list \219 --report eval_data/snapshots/gh-live-latest-1000x1000/analysis-report-hybrid.json220```221 222Then synthesize and publish one minimal upstream PR from the top-ranked mergeable cluster:223 224```bash225uv run slop-farmer duplicate-prs merge \226 --report eval_data/snapshots/gh-live-latest-1000x1000/analysis-report-hybrid.json \227 --repo-dir /path/to/transformers228```229 230If your local checkout uses a fork as `origin`, point the merge flow at the upstream remote explicitly and relax the file policy when needed:231 232```bash233uv run slop-farmer duplicate-prs merge \234 --report eval_data/snapshots/gh-live-latest-1000x1000/analysis-report-hybrid.json \235 --repo-dir /path/to/transformers \236 --upstream-repo huggingface/transformers \237 --upstream-remote upstream \238 --fork-repo YOURNAME/transformers-minimal \239 --fork-remote origin \240 --file-policy allow-docs241```242 243## Import a historical HF checkpoint as a clean local snapshot244 245If an older dataset keeps its richest data under `_checkpoints/<snapshot_id>/`,246you can promote one of those checkpoints into a normal local snapshot:247 248```bash249uv run slop-farmer import-hf-checkpoint \250 --source-repo-id burtenshaw/transformers-pr-slop-dataset \251 --output-dir eval_data252```253 254By default this selects the latest viable checkpoint, writes a clean snapshot255under `eval_data/snapshots/`, and regenerates `links.parquet`,256`issue_comments.parquet`, and `pr_comments.parquet`.257 258## Render markdown from an analysis JSON259 260You can turn an existing analysis report into a human-readable markdown file without rerunning clustering:261 262```bash263uv run slop-farmer markdown-report \264 --input eval_data/snapshots/hf-latest-100x100/analysis-report-hybrid.json265```266 267By default this writes `analysis-report-hybrid.md` next to the JSON and uses the JSON parent directory as the snapshot source for issue and PR titles, links, and latest-activity ordering.268 269## Render a new contributor report270 271You can also render a reviewer-facing markdown report for contributors who are still new to the repo snapshot:272 273```bash274uv run slop-farmer new-contributor-report \275 --snapshot-dir data/snapshots/20260324T000000Z276```277 278By default this writes:279 280- `new_contributors.parquet`281- `new-contributors-report.md`282- `new-contributors-report.json`283 284next to the snapshot, including GitHub profile links, repo issue/PR search links, and example authored artifacts.285 286## Full end-to-end workflow287 288You can run scrape + publish + analyze + markdown + dashboard export in one command:289 290```bash291uv run slop-farmer full-pipeline \292 --repo huggingface/transformers \293 --dataset YOURNAME/transformers-pr-slop-dataset \294 --model "gpt-5-mini?reasoning=low"295```296 297This writes outputs under a repo-anchored workspace directory, for example:298 299- `runs/transformers/data/`300- `runs/transformers/web/public/data/`301 302Optional age caps are based on `created_at`:303 304```bash305 --issue-max-age-days 30 \306 --pr-max-age-days 14307```308 309## Validation checks310 311Before committing or wiring new package moves into automation, run:312 313```bash314uv run python scripts/enforce_packaging.py315uv run --extra dev ruff format --check src tests scripts jobs316uv run --extra dev ruff check src tests scripts jobs317uv run --extra dev ty check src tests scripts jobs318uv run --extra dev pytest -q319```320 321`scripts/enforce_packaging.py` verifies the coarse package boundaries:322 323- `data` must not import `app`324- `data` must not import `reports`325- `reports` must not import `app`326 327## YAML config-driven runs328 329You can keep repo-specific pipeline defaults in a YAML file and apply them to all330commands with `--config`.331 332Example: `configs/diffusers.yaml`333 334```yaml335repo: huggingface/diffusers336workspace: runs/diffusers337dataset_id: evalstate/diffusers-pr338 339pull-requests:340 template_cleanup:341 mode: merge_defaults342 line_patterns:343 - '^d(?:o not merge|ontmerge)\.?$'344 cluster_suppression_rules:345 - id: diffusers_post_release346 title_patterns:347 - '\bpost[- ]release\b'348 349dashboard:350 space_id: evalstate/diffusers-dashboard351 title: Diffusers Dashboard352 window_days: 60353 contributor_window_days: 60354 contributor_max_authors: 0355 356analysis:357 model: gpt-5.4-mini358 ranking_backend: hybrid359 cached_analysis: true360 361scrape:362 fetch-timeline: true363```364 365Then commands stay aligned without repeating repo/workspace/window settings:366 367```bash368uv run slop-farmer --config configs/diffusers.yaml refresh-dataset369uv run slop-farmer --config configs/diffusers.yaml analyze370uv run slop-farmer --config configs/diffusers.yaml pr-scope371uv run slop-farmer --config configs/diffusers.yaml pr-search refresh372uv run slop-farmer --config configs/diffusers.yaml new-contributor-report373uv run slop-farmer --config configs/diffusers.yaml dashboard-data374uv run slop-farmer --config configs/diffusers.yaml deploy-dashboard --refresh-contributors375uv run slop-farmer --config configs/diffusers.yaml dataset-status376```377 378Those reader commands default to `dataset_id` when configured. Pass `--snapshot-dir` to force379an explicit local snapshot instead.380 381If you run `analyze` before `publish-snapshot`, the uploaded snapshot will also include382`analysis-state/`, which makes the hybrid cache portable across machines and reusable in383later snapshots when `analysis.cached_analysis: true` is enabled.384 385## Export static dashboard data386 387You can export a slim JSON bundle for the React dashboard:388 389```bash390uv run slop-farmer dashboard-data \391 --snapshot-dir data/snapshots/20260324T150154Z \392 --output-dir web/public/data \393 --window-days 14394```395 396This writes:397 398- `summary.json`399- `clusters.json`400- `prs.json`401- `contributors.json`402 403The dashboard is intentionally summary-first and links out to GitHub for deep detail.404 405## Deploy a dashboard to a Hugging Face Space406 407Use the generic deploy script:408 409```bash410SPACE_ID=evalstate/openclaw-pr-report \411PIPELINE_DATA_DIR=runs/openclaw/data \412SNAPSHOT_DIR=runs/openclaw/data/snapshots/20260324T233649Z \413SPACE_TITLE="OpenClaw PR Report" \414DATASET_ID=evalstate/openclaw-pr \415scripts/deploy_dashboard_space.sh416```417 418Repo-specific wrappers are also available:419 420- `scripts/deploy_transformers_dashboard_space.sh`421- `scripts/deploy_openclaw_dashboard_space.sh`422 423Or use the CLI wrapper with a YAML config:424 425```bash426uv run slop-farmer --config configs/diffusers.yaml deploy-dashboard --refresh-contributors427```428 429## Deploy the PR similarity API to a Hugging Face Docker Space430 431The repo includes the FastAPI service for the read-oriented PR similarity surface.432The standalone `pr-search` client now lives in the downstream `pr-search-cli`433package.434 435Deploy the OpenClaw API Space with:436 437```bash438scripts/update_openclaw_pr_search_api.sh439```440 441Or use the generic deploy script directly:442 443```bash444SPACE_ID=evalstate/openclaw-pr-api \445SPACE_TITLE="OpenClaw PR API" \446DEFAULT_REPO=openclaw/openclaw \447GHR_BASE_URL=https://ghreplica.dutiful.dev \448HF_REPO_ID=evalstate/openclaw-pr \449BUCKET_ID=evalstate/openclaw-pr-api-data \450scripts/deploy_pr_search_space.sh451```452 453This deploy flow:454 455- creates or updates a Docker Space456- uploads a minimal app bundle with a generated Space `README.md`457- sets runtime variables for the API458- mounts the configured HF bucket at `/data`459 460After the Space is live, you can query it either through the in-repo admin CLI:461 462```bash463uv run slop-farmer pr-search status --repo openclaw/openclaw464uv run slop-farmer pr-search similar 67096 --repo openclaw/openclaw465```466 467Or through the downstream `pr-search-cli` package, which owns the standalone468`pr-search` executable.469 