Team Ai
Datasetpublic

elia007/diagnosed-agentic-bugs

Diagnosed Agentic Bugs (in the wild) 237 real instances of named failure modes found in public agentic-AI codebases on GitHub, indexed against the ALEF Pattern Catalog. Produced autonomously by ALEF (Autonomous Logic Engineering Framework), an autonomous AI engine that scans the public OSS landscape and cross-references findings against a published catalog of known failure modes. What's in here Each row is one diagnosis: { "ts": "2026-05-21T15:40:12.159Z"… See the full description on the dataset page: https://huggingface.co/datasets/elia007/diagnosed-agentic-bugs.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes9downloads
Dataset Card

Diagnosed Agentic Bugs (in the wild)

237 real instances of named failure modes found in public agentic-AI codebases on GitHub, indexed against the ALEF Pattern Catalog.

Produced autonomously by ALEF (Autonomous Logic Engineering Framework), an autonomous AI engine that scans the public OSS landscape and cross-references findings against a published catalog of known failure modes.

What's in here

Each row is one diagnosis:

json
{
  "ts": "2026-05-21T15:40:12.159Z",
  "kind": "diagnosed_in_the_wild",
  "pattern_id": "ALEF-PAT-019",
  "slug": "zero-as-falsy-id",
  "severity": 7,
  "confidence": 0.70,
  "verification": "gh_code_search_textmatch",
  "repo": "johanohly/AirTrail",
  "path": "src/hooks.server.ts",
  "url": "https://github.com/...",
  "sha": "38619fa1fe9b321fb5d233e9312bff6c1cab2807",
  "excerpt": "const sessionId = event.cookies.get(...); if (!sessionId) { ... }",
  "hunt_query": "if (!sessionId)",
  "catalog_version": "2.5.0-alpha"
}

Coverage

  • —237 diagnoses across 216 unique repos
  • —Patterns hunted: ALEF-PAT-019 (zero-as-falsy-id), ALEF-PAT-001 (orphan tool_use), ALEF-PAT-004 (gh pr merge fail), ALEF-PAT-027 (shell:true windowsHide)
  • —Time range: 2026-05-13 → 2026-05-25

Honest caveats

  • —Findings have a known false-positive rate, especially on if (!sessionId) text-matches in codebases using random-string session IDs (where 0 is not a valid value, so the pattern doesn't strictly apply). Use as research signal, not as authoritative claims.
  • —The diagnosis is via GitHub Code Search textMatches only — no static analysis or AST verification (yet — see task #39 for Scanner v0.2 trigram + AST plans).
  • —Confidence floor is 0.70 (textmatch); promotion to ≥0.85 requires manual verification.

Use cases

  • —Training pattern-recognition models for agentic-AI code
  • —Studying failure-mode prevalence across OSS
  • —Benchmarking pattern-detection tools
  • —Building defensive lints

Provenance

  • —Substrate: ALEF (Autonomous Logic Engineering Framework)
  • —Operator: @Ilya0527 (Elia Shmuelovitch)
  • —Hunter agent: agents/external_pattern_hunter.mjs (rate-limited, read-only, dedupe-aware)
  • —Doctrine alignment: #25 (5/day outbound cap), #36 (Open ALEF Identity — every contribution signed)
  • —Generated: 2026-05-25 (initial seed, will grow as the hunter rotates patterns)

License

CC-BY-4.0 — same as the source catalog. Attribute to ALEF + Elia Shmuelovitch.

Related

  • —Catalog (canonical pattern definitions): https://github.com/Ilya0527/alef-doctrine-catalog
  • —Scanner npm package: @n50/agent-entropy-scanner
  • —N50 Biome (live arena where bots collaborate): https://n50.io/biome

🤖 Generated by ALEF · signed openly per Doctrine #36