Team Ai
Datasetpublic

davidkling/hf-coding-tools-dashboard-all

HuggingFace AI Coding Tools Dashboard Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories. Dataset Structure Split Description Rows results Full benchmark results with LLM responses, cost, tokens, latency, and product detection 9603 queries Benchmark query definitions across 32 categories 404 runs Run metadata and… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-all.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes31downloads
Dataset Card

HuggingFace AI Coding Tools Dashboard

Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.

Dataset Structure

SplitDescriptionRows
resultsFull benchmark results with LLM responses, cost, tokens, latency, and product detection9603
queriesBenchmark query definitions across 32 categories404
runsRun metadata and tool/model configurations2
productsHuggingFace product catalog with detection keywords44

Key Fields (results)

  • —tool: AI coding tool tested (claude_code, codex, copilot, cursor)
  • —model: Specific model used
  • —response: Full raw LLM response text
  • —detected_products: HuggingFace products mentioned in the response
  • —cost_usd / tokens_input / tokens_output / latency_ms: Performance metrics
  • —attempt_number: 1-indexed attempt within each (query_id, tool, model, effort, thinking) group
  • —is_latest_attempt: True if this is the most recent attempt in its group

Notes on retries

Some (query_id, tool, model, effort, thinking) configurations were re-run during data collection (mostly Claude Code, due to credit/timeout retries on Run 53). Both attempts are kept in this dataset for variance analysis.

  • —Use `is_latest_attempt = true` to filter to one row per configuration (8,359 rows). Recommended for aggregate rate calculations to avoid double-counting.
  • —Use all rows (9,146) to study response consistency / variance across retries.

Distribution: 7,820 configurations ran once; 539 ran 2 or 3 times.

Example Queries

DuckDB:

sql
SELECT tool, COUNT(*) as mentions
FROM results
WHERE response LIKE '%xet%'
GROUP BY tool

Python:

python
from datasets import load_dataset
results = load_dataset("davidkling/hf-coding-tools-dashboard-all", "results")
queries = load_dataset("davidkling/hf-coding-tools-dashboard-all", "queries")