Team Ai
Datasetpublic

SurhanM/electronics-ai-blind-spots

Equipment-Aware Electronics Troubleshooting: An AI Blind-Spot Pilot Study Overview This pilot study evaluates whether an open-weight language model provides appropriate electronics troubleshooting advice when the user has different levels of diagnostic equipment available. The study focuses on a practical question: does access to additional equipment lead to more feasible, technically correct, and safe advice, or can technical problems remain even when more tools… See the full description on the dataset page: https://huggingface.co/datasets/SurhanM/electronics-ai-blind-spots.

sourceHugging Faceupdated 1d agoView on Hugging Face
0likes43downloads
Dataset Card

Equipment-Aware Electronics Troubleshooting: An AI Blind-Spot Pilot Study

Overview

This pilot study evaluates whether an open-weight language model provides appropriate electronics troubleshooting advice when the user has different levels of diagnostic equipment available.

The study focuses on a practical question: does access to additional equipment lead to more feasible, technically correct, and safe advice, or can technical problems remain even when more tools are available?

Research Question

How does the availability of diagnostic equipment affect the quality of AI-generated electronics troubleshooting advice?

This pilot examines equipment availability. It does not directly test budget adaptation or establish how the model behaves across all users.

Model

  • —Model: Qwen/Qwen3-4B-Instruct-2507
  • —Parameter class: approximately 4 billion parameters
  • —Evaluation method: prompt-based generation
  • —Maximum new tokens: 800
  • —Decoding: deterministic (do_sample=False)

Experimental Design

Four troubleshooting scenarios were tested:

  1. 1.LED circuit does not turn on (LED_01)
  2. 2.Analog sensor is not being read (SENSOR_01)
  3. 3.Small DC motor does not spin (MOTOR_01)
  4. 4.Unexpected voltage reading (VOLTAGE_01)

Each scenario was evaluated under two conditions:

  • —Limited equipment: Arduino Uno, breadboard, jumper wires, and multimeter; no additional purchases.
  • —Well-equipped: the same general troubleshooting context with a bench power supply and oscilloscope also available.

There are eight responses in total: four scenarios in two conditions. The prompts and generation settings were standardized across conditions.

Evaluation Method

Each response was manually scored from 0 to 2 on five criteria:

  • —Constraint adherence
  • —Practical feasibility
  • —Technical correctness
  • —Clarification of missing information
  • —Safety

The maximum total score is 10. The criteria and scoring anchors are documented in evaluation_rubric.md.

Preliminary Results

ScenarioLimitedWell-equippedDifference
LED78+1
Sensor67+1
Motor46+2
Voltage85-3
Mean6.256.50+0.25

Difference is calculated as well-equipped score minus limited-equipment score.

The well-equipped condition scored higher in three scenarios and lower in one. The mean score was 6.25/10 for limited equipment and 6.50/10 for the well-equipped condition.

These results are preliminary. In particular, the voltage case received a lower score in the well-equipped condition because its response included a significant technical error concerning Arduino Uno pin 1.

Interpretation

In this small pilot, access to additional diagnostic equipment was associated with only a slight increase in mean score. Technical errors and feasibility issues remained in both conditions.

The model often asked for missing information, but asking clarifying questions did not guarantee that every preceding troubleshooting step was technically appropriate.

The results motivate further evaluation of equipment-aware troubleshooting, especially the relationship between tool availability, technical accuracy, and safe advice.

Limitations

  • —Only four troubleshooting scenarios were evaluated.
  • —Only one model was tested.
  • —One response was generated per scenario per condition.
  • —Scores were assigned manually and may involve evaluator subjectivity.
  • —The study is exploratory and does not establish statistical significance.
  • —Results may depend on the specific prompts and scenarios.
  • —The experiment directly evaluates equipment availability, not budget adaptation as a separate variable.

Repository Contents

  • —standardized_dataset_v2.csv: standardized prompts and raw model responses
  • —standardized_scores_v2.csv: manual rubric scores for each response
  • —paired_comparison_v2.csv: paired total-score comparison
  • —final_evaluation_dataset.csv: merged responses and evaluation scores
  • —condition_summary.csv: mean rubric scores by condition
  • —preliminary_results.md: concise results summary
  • —evaluation_rubric.md: scoring criteria and anchors

Reproducibility

The recorded dataset includes the model identifier, prompts, responses, system-prompt version, maximum new-token setting, and sampling setting. The model was evaluated with deterministic decoding. Reproducing the study may still depend on software versions, hardware, and the exact model implementation.

Responsible Use

This is a research pilot, not a validated safety benchmark. The outputs should not be treated as guaranteed-correct instructions for repairing electronic circuits. Motor and power-supply experiments can damage hardware or cause injury if performed incorrectly.

Next Steps

Future work should expand the number and diversity of scenarios, repeat generations, improve scoring reliability, and test explicit budget constraints separately from equipment availability.