SurhanM/electronics-ai-blind-spots
Equipment-Aware Electronics Troubleshooting: An AI Blind-Spot Pilot Study Overview This pilot study evaluates whether an open-weight language model provides appropriate electronics troubleshooting advice when the user has different levels of diagnostic equipment available. The study focuses on a practical question: does access to additional equipment lead to more feasible, technically correct, and safe advice, or can technical problems remain even when more tools… See the full description on the dataset page: https://huggingface.co/datasets/SurhanM/electronics-ai-blind-spots.
Equipment-Aware Electronics Troubleshooting: An AI Blind-Spot Pilot Study
Overview
This pilot study evaluates whether an open-weight language model provides appropriate electronics troubleshooting advice when the user has different levels of diagnostic equipment available.
The study focuses on a practical question: does access to additional equipment lead to more feasible, technically correct, and safe advice, or can technical problems remain even when more tools are available?
Research Question
How does the availability of diagnostic equipment affect the quality of AI-generated electronics troubleshooting advice?
This pilot examines equipment availability. It does not directly test budget adaptation or establish how the model behaves across all users.
Model
- Model:
Qwen/Qwen3-4B-Instruct-2507 - Parameter class: approximately 4 billion parameters
- Evaluation method: prompt-based generation
- Maximum new tokens: 800
- Decoding: deterministic (
do_sample=False)
Experimental Design
Four troubleshooting scenarios were tested:
- LED circuit does not turn on (
LED_01) - Analog sensor is not being read (
SENSOR_01) - Small DC motor does not spin (
MOTOR_01) - Unexpected voltage reading (
VOLTAGE_01)
Each scenario was evaluated under two conditions:
- Limited equipment: Arduino Uno, breadboard, jumper wires, and multimeter; no additional purchases.
- Well-equipped: the same general troubleshooting context with a bench power supply and oscilloscope also available.
There are eight responses in total: four scenarios in two conditions. The prompts and generation settings were standardized across conditions.
Evaluation Method
Each response was manually scored from 0 to 2 on five criteria:
- Constraint adherence
- Practical feasibility
- Technical correctness
- Clarification of missing information
- Safety
The maximum total score is 10. The criteria and scoring anchors are documented in evaluation_rubric.md.
Preliminary Results
Difference is calculated as well-equipped score minus limited-equipment score.
The well-equipped condition scored higher in three scenarios and lower in one. The mean score was 6.25/10 for limited equipment and 6.50/10 for the well-equipped condition.
These results are preliminary. In particular, the voltage case received a lower score in the well-equipped condition because its response included a significant technical error concerning Arduino Uno pin 1.
Interpretation
In this small pilot, access to additional diagnostic equipment was associated with only a slight increase in mean score. Technical errors and feasibility issues remained in both conditions.
The model often asked for missing information, but asking clarifying questions did not guarantee that every preceding troubleshooting step was technically appropriate.
The results motivate further evaluation of equipment-aware troubleshooting, especially the relationship between tool availability, technical accuracy, and safe advice.
Limitations
- Only four troubleshooting scenarios were evaluated.
- Only one model was tested.
- One response was generated per scenario per condition.
- Scores were assigned manually and may involve evaluator subjectivity.
- The study is exploratory and does not establish statistical significance.
- Results may depend on the specific prompts and scenarios.
- The experiment directly evaluates equipment availability, not budget adaptation as a separate variable.
Repository Contents
standardized_dataset_v2.csv: standardized prompts and raw model responsesstandardized_scores_v2.csv: manual rubric scores for each responsepaired_comparison_v2.csv: paired total-score comparisonfinal_evaluation_dataset.csv: merged responses and evaluation scorescondition_summary.csv: mean rubric scores by conditionpreliminary_results.md: concise results summaryevaluation_rubric.md: scoring criteria and anchors
Reproducibility
The recorded dataset includes the model identifier, prompts, responses, system-prompt version, maximum new-token setting, and sampling setting. The model was evaluated with deterministic decoding. Reproducing the study may still depend on software versions, hardware, and the exact model implementation.
Responsible Use
This is a research pilot, not a validated safety benchmark. The outputs should not be treated as guaranteed-correct instructions for repairing electronic circuits. Motor and power-supply experiments can damage hardware or cause injury if performed incorrectly.
Next Steps
Future work should expand the number and diversity of scenarios, repeat generations, improve scoring reliability, and test explicit budget constraints separately from equipment availability.
