Okyanus/agri-telemetry-sanity-bench
Agri Telemetry Sanity Bench Can a small language model check greenhouse sensor data as reliably as a few lines of rules? This benchmark asks exactly that, and ships the answer for the models we trained: not yet. Two tracks, both with exact ground truth: Track Cases Task sensor_quality 123 Is this sensor packet usable? Flag missing, impossible, broken, stale or future-dated readings. tomato_risk 153 Which tomato greenhouse risk labels apply (high pH, heat stress… See the full description on the dataset page: https://huggingface.co/datasets/Okyanus/agri-telemetry-sanity-bench.
Agri Telemetry Sanity Bench
Can a small language model check greenhouse sensor data as reliably as a few lines of rules? This benchmark asks exactly that, and ships the answer for the models we trained: not yet.
Two tracks, both with exact ground truth:
Ground truth is defined by Pomona's deterministic rules (services/model-router/app/sensor_quality.py and tomato_reasoner.py). The benchmark measures whether a model reproduces those checks. It does not measure agronomic truth.
Results
sensor_quality (123 cases)
tomato_risk (153 cases)
Per-case predictions for every row are in predictions/, and full per-bucket scores are in results.json.
What we learned
- Exact arithmetic is the wall. Better training data fixed pattern problems (missing fields 0.55 → 0.99 missing-field F1; broken values from 0 % to 60–100 % exact, depending on the field), but the fine-tuned 0.8B model still got 0 % of stale and future timestamps right (even on its own test split), and missed pH 2.999 in 4 of 4 cases. Those are subtraction and comparison, which rules do perfectly.
- In-distribution scores can be template memorisation. The tomato model scores 0.89 on its own generator's test split but 0.34 here, level with the always-normal baseline. It answers
[]to almost everything new. Changing only an irrelevant field (growth_stage) flips a correcthigh_phto[], because its training data used only a handful of distinct threshold values. - Always keep the rules in the loop. In Pomona the models are advisory and the rules decide. This benchmark is the evidence for that design.
Data format
sensor_quality:
tomato_risk: id, bucket, input_json, expected_risk_labels.
Inputs are stored as strings on purpose. Several sensor values are deliberately the wrong type ("broken", true, {}) or deliberately left out, and columnar loaders such as datasets or pandas can silently rewrite them (null-filling missing keys, reformatting timestamps). Parse each input_json with json.loads before building a prompt.
How to evaluate your model
- Build prompts. Reference prompts are in
prompts/:sensor_quality_render.pyrenders the exact sensor-quality prompt;tomato_risk_eval_prompt.txtis the tomato prompt used here. You may use any prompt you like; report it. - Write predictions as JSONL:
{"id": "...", "output": "<raw model text>"}. - Score (standard library only):
python evaluate.py sensor_quality my_predictions.jsonl --by-bucket
python evaluate.py tomato_risk my_predictions.jsonl --by-bucketThe first JSON object (sensor quality) or JSON list (tomato) in each output is scored. Unparseable answers score zero. Open a discussion with your results and we'll add them to the table.
How the cases were built
- `sensor_quality`: 123 explicit cases from Pomona's release-candidate suite (normal packets, missing values as
nulland as omitted keys, invalid values, exact range edges on both sides, stale and future timestamps at ±1 s of the limit, unit mismatches, drift and conflicts). None of these inputs appear in any Pomona training set (identity/time-stripped exact-match check). - `tomato_risk`: 153 fresh packets with values on both sides of every rule threshold, plus normal, missing-data, actuator-conflict and combined cases. Checked for zero exact overlap with the tomato model's train, validation, test and golden files.
- All packets are synthetic. They are not field measurements.
Limitations
- Ground truth is one project's rule set. Other agronomic thresholds are equally valid for other crops or systems.
- Small: good for spotting failure modes, not for fine-grained ranking.
- The sensor-quality models listed are unpublished research candidates. The tomato model is published as a research preview.
Links
- Platform and rules: okyanu/pomona
- Collection: Pomona — Local AI for Safer Greenhouse Decision Support
Citation
@misc{pomona_agri_telemetry_sanity_bench_2026,
title = {Agri Telemetry Sanity Bench},
author = {Okyanus},
year = {2026},
url = {https://huggingface.co/datasets/Okyanus/agri-telemetry-sanity-bench}
}