Team Ai
Datasetpublic

Okyanus/agri-telemetry-sanity-bench

Agri Telemetry Sanity Bench Can a small language model check greenhouse sensor data as reliably as a few lines of rules? This benchmark asks exactly that, and ships the answer for the models we trained: not yet. Two tracks, both with exact ground truth: Track Cases Task sensor_quality 123 Is this sensor packet usable? Flag missing, impossible, broken, stale or future-dated readings. tomato_risk 153 Which tomato greenhouse risk labels apply (high pH, heat stress… See the full description on the dataset page: https://huggingface.co/datasets/Okyanus/agri-telemetry-sanity-bench.

sourceHugging Faceapache-2.0updated 13d agoView on Hugging Face
0likes70downloads
Dataset Card

Agri Telemetry Sanity Bench

Can a small language model check greenhouse sensor data as reliably as a few lines of rules? This benchmark asks exactly that, and ships the answer for the models we trained: not yet.

Two tracks, both with exact ground truth:

TrackCasesTask
sensor_quality123Is this sensor packet usable? Flag missing, impossible, broken, stale or future-dated readings.
tomato_risk153Which tomato greenhouse risk labels apply (high pH, heat stress, fungal pressure…)?

Ground truth is defined by Pomona's deterministic rules (services/model-router/app/sensor_quality.py and tomato_reasoner.py). The benchmark measures whether a model reproduces those checks. It does not measure agronomic truth.

Results

sensor_quality (123 cases)

SystemTypeParamsLabel F1Missing-field F1Suspect-field F1Review matchExact
pomona_rulesdeterministic rules01.0001.0001.0001.0001.000
always_normaltrivial baseline00.1630.7890.3250.1630.163
pomona_sq_v0.1.4_qwen3.5-0.8b_loraQLoRA fine-tune (unpublished candidate)0.8B0.4050.5530.5490.8370.358
pomona_sq_v0.1.5_qwen3.5-0.8b_loraQLoRA fine-tune (unpublished candidate)0.8B0.6840.9920.6570.7560.618

tomato_risk (153 cases)

SystemTypeParamsLabel F1Exact
pomona_rulesdeterministic rules01.0001.000
always_normaltrivial baseline00.3270.327
pomona_tomato_v0.1.7_gguf_eval_script_promptLoRA fine-tune, F16 GGUF via Ollama0.5B0.3420.333
pomona_tomato_v0.1.7_gguf_platform_promptLoRA fine-tune, F16 GGUF via Ollama0.5B0.3150.307

Per-case predictions for every row are in predictions/, and full per-bucket scores are in results.json.

What we learned

  1. 1.Exact arithmetic is the wall. Better training data fixed pattern problems (missing fields 0.55 → 0.99 missing-field F1; broken values from 0 % to 60–100 % exact, depending on the field), but the fine-tuned 0.8B model still got 0 % of stale and future timestamps right (even on its own test split), and missed pH 2.999 in 4 of 4 cases. Those are subtraction and comparison, which rules do perfectly.
  2. 2.In-distribution scores can be template memorisation. The tomato model scores 0.89 on its own generator's test split but 0.34 here, level with the always-normal baseline. It answers [] to almost everything new. Changing only an irrelevant field (growth_stage) flips a correct high_ph to [], because its training data used only a handful of distinct threshold values.
  3. 3.Always keep the rules in the loop. In Pomona the models are advisory and the rules decide. This benchmark is the evidence for that design.

Data format

sensor_quality:

FieldMeaning
id, bucketCase id and the case type (e.g. future_time, invalid_ph, boundary_humidity_pct)
input_jsonThe packet as a JSON string: farm_context, sensor, expected_fields, current_time
expected_labels, expected_missing_fields, expected_suspect_fieldsSorted lists
expected_human_reviewBoolean

tomato_risk: id, bucket, input_json, expected_risk_labels.

Inputs are stored as strings on purpose. Several sensor values are deliberately the wrong type ("broken", true, {}) or deliberately left out, and columnar loaders such as datasets or pandas can silently rewrite them (null-filling missing keys, reformatting timestamps). Parse each input_json with json.loads before building a prompt.

How to evaluate your model

  1. 1.Build prompts. Reference prompts are in prompts/: sensor_quality_render.py renders the exact sensor-quality prompt; tomato_risk_eval_prompt.txt is the tomato prompt used here. You may use any prompt you like; report it.
  2. 2.Write predictions as JSONL: {"id": "...", "output": "<raw model text>"}.
  3. 3.Score (standard library only):
bash
python evaluate.py sensor_quality my_predictions.jsonl --by-bucket
python evaluate.py tomato_risk my_predictions.jsonl --by-bucket

The first JSON object (sensor quality) or JSON list (tomato) in each output is scored. Unparseable answers score zero. Open a discussion with your results and we'll add them to the table.

How the cases were built

  • —`sensor_quality`: 123 explicit cases from Pomona's release-candidate suite (normal packets, missing values as null and as omitted keys, invalid values, exact range edges on both sides, stale and future timestamps at ±1 s of the limit, unit mismatches, drift and conflicts). None of these inputs appear in any Pomona training set (identity/time-stripped exact-match check).
  • —`tomato_risk`: 153 fresh packets with values on both sides of every rule threshold, plus normal, missing-data, actuator-conflict and combined cases. Checked for zero exact overlap with the tomato model's train, validation, test and golden files.
  • —All packets are synthetic. They are not field measurements.

Limitations

  • —Ground truth is one project's rule set. Other agronomic thresholds are equally valid for other crops or systems.
  • —Small: good for spotting failure modes, not for fine-grained ranking.
  • —The sensor-quality models listed are unpublished research candidates. The tomato model is published as a research preview.

Links

Citation

bibtex
@misc{pomona_agri_telemetry_sanity_bench_2026,
  title  = {Agri Telemetry Sanity Bench},
  author = {Okyanus},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Okyanus/agri-telemetry-sanity-bench}
}