Team Ai
Datasetpublic

sauravsingla08/configreach-validation

ConfigReach Curated 50K Benchmark The ConfigReach Curated 50K Benchmark is the committed controlled benchmark used to evaluate ConfigReach configuration-input detection across supported programming languages and configuration formats. It contains 50,000 scenarios with deterministic ground-truth labels: 25,000 positive and 25,000 negative cases across 24 language/configuration groups. The benchmark is scored through the production configreach.engine.scan entry point. Project hub:… See the full description on the dataset page: https://huggingface.co/datasets/sauravsingla08/configreach-validation.

sourceHugging Facemitupdated 9d agoView on Hugging Face
0likes129downloads
Dataset Card

ConfigReach Curated 50K Benchmark

The ConfigReach Curated 50K Benchmark is the committed controlled benchmark used to evaluate ConfigReach configuration-input detection across supported programming languages and configuration formats.

It contains 50,000 scenarios with deterministic ground-truth labels: 25,000 positive and 25,000 negative cases across 24 language/configuration groups. The benchmark is scored through the production configreach.engine.scan entry point.

Project hub: ConfigReach — Configuration Coverage collection

Measured result

MetricResult
Scenarios50,000
True positives25,000
False positives0
True negatives25,000
False negatives0
Precision100.0000%
Recall100.0000%
F1100.0000%
Accuracy100.0000%
  • —Repository package metadata: v0.9.5
  • —Semantic engine: v0.9
  • —Dataset SHA-256: b2948a4cbf4d8d6263c3347ca4bd4c5d37d645a17cad88b2859cabe0a7b2392e
Scope: this is the measured result on the committed controlled curated benchmark. It is not a claim of universal real-world accuracy.

What is in the dataset?

Each row is one independent detection scenario with these fields:

  • —scenario_id — stable benchmark case identifier
  • —group — programming language or configuration-format family
  • —variant — scenario pattern within the group
  • —expected_detect — deterministic ground-truth detection label
  • —expected_key — expected configuration key when detection is positive
  • —suggested_filename — filename/extension used when materializing the scenario
  • —snippet — source/configuration snippet presented to the scanner

The benchmark includes true configuration accesses/declarations as well as adversarial negatives such as comments, inert strings, custom look-alike APIs, dynamically computed names, quoted shell literals, and structured-configuration value-only cases.

The corpus is curated around explicit, version-controlled scenario families with deterministic ground-truth labels. This dataset is not described as independently human-labelled.

Reproducibility evidence

This Hugging Face repository is generated from evidence committed in the ConfigReach GitHub repository. In addition to the 50K JSONL corpus, the published snapshot includes:

  • —metadata/manifest.json — corpus counts, per-group counts, byte size and SHA-256
  • —metadata/results.json — aggregate, per-group and per-variant benchmark results
  • —metadata/summary.json — compact publication summary and scanner provenance
  • —evidence/configreach_50k_predictions.csv — row-level predictions used to derive the aggregate metrics

Load with 🤗 Datasets

python
from datasets import load_dataset

ds = load_dataset("sauravsingla08/configreach-validation", split="validation")
print(ds)
print(ds[0])

A simple analysis example:

python
from datasets import load_dataset

rows = load_dataset("sauravsingla08/configreach-validation", split="validation")
positives = rows.filter(lambda row: row["expected_detect"])
negatives = rows.filter(lambda row: not row["expected_detect"])
print(len(rows), len(positives), len(negatives))

Suitable uses

This benchmark is intended for:

  • —reproducible regression testing of configuration-detection tooling;
  • —precision/recall experiments on a controlled labelled corpus;
  • —cross-language and cross-format static-analysis research;
  • —evaluation of false-positive suppression for comments, strings and look-alike APIs;
  • —teaching and examples for environment-variable, feature-flag and configuration analysis.

Limitations

The benchmark is controlled and curated around explicit scenario families. It is useful for deterministic regression evidence, but it should not be treated as a representative sample of all real-world software or as evidence of universal ConfigReach accuracy.

Project links

  • —Hugging Face Collection: https://huggingface.co/collections/sauravsingla08/configreach-configuration-coverage
  • —Hugging Face Space: https://huggingface.co/spaces/sauravsingla08/ConfigReach
  • —GitHub: https://github.com/sauravsingla/ConfigReach
  • —50K benchmark source: https://github.com/sauravsingla/ConfigReach/tree/main/validation/curated_50k
  • —Measured result: https://github.com/sauravsingla/ConfigReach/blob/main/validation/results/curated_50k.md
  • —PyPI: https://pypi.org/project/configreach/