datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-assistant-hvac-picks
AI assistant picks for HVAC, 2,383 cities
When a homeowner asks an AI assistant "My furnace won't turn on. Who can fix it this week in Plano?", which company does it name, which does it pass over, and why?
This dataset records one AI assistant's answers to six heating and cooling questions in each of 2,383 cities (2,068 US, 88 Canada, 227 UK), September 29 to October 1, 2026: 34,907 picks with the reason given and the page it read, 22,056 companies it considered and passed over… See the full description on the dataset page: https://huggingface.co/datasets/GregM/ai-assistant-hvac-picks.AssistantEval-results
AssistantEval results
AssistantEval results are aggregate measurements of assistants completing synthetic, stateful tasks. This repository contains the published leaderboard CSV: one row per assistant, with task outcomes, behavior measures, grading counts, confidence intervals, and timing fields.
The CSV is the original public file from assistanteval.com/data/results.csv. The accompanying source index and methodology describe the published snapshot and grading process.… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/AssistantEval-results.ctrlpotato-ai-interview-assistant-benchmark
CTRLpotato AI Interview Assistant Cross-review Evidence Matrix (2026)
A citation-ready snapshot of hands-on desktop evidence for six AI interview assistants: Cluely, Interview Coder, LockedIn AI, ULTRACODE AI, Parakeet AI, and Final Round AI.
The package contains 66 assessments across 6 products and 11 shared criteria. Product versions and test dates are preserved in every row.
Important scope
This is a cross-review evidence matrix, not a statistically controlled… See the full description on the dataset page: https://huggingface.co/datasets/ae0j/ctrlpotato-ai-interview-assistant-benchmark.llava-llama3_8b_hierarchical_resultsSoVTP_pipeline_completed_qwen3-vl-8b-instructllava-phi3_3.8b_discrete_standardizedqwen3-vl_4b-instruct_continuous_resultsqwen3-vl_4b-instruct_hierarchical_standardizedgemma3_4b-it-q4_K_M_hierarchical_standardizedgemma3_4b-it-q4_K_M_discrete_resultsllava_13b_discrete_standardizedgemma3_12b-it-q4_K_M_hierarchical_standardizedministral-3_3b-instruct-2512-q4_K_M_discrete_standardizedSoVTP_sameimg_gpt-5-nano_complete_backupqwen3-vl_4b-instruct_discrete_standardizedqwen3-vl_30b-a3b-instruct_discrete_resultsall_models_resultsllava_13b_hierarchical_resultsgemma3_4b-it-q4_K_M_hierarchical_resultsall_models_results_v1results_qwen-vl8bresults_all_qwen_inst-vl8bresults_qwen_inst-vl8bminicpm-v_8b_discrete_standardizedminicpm-v_8b_discrete_resultsmoondream_1.8b_discrete_standardizedqwen3-vl_30b-a3b-instruct_discrete_standardizedqwen3-vl_4b-instruct_discrete_resultsqwen3-vl_32b-instruct_discrete_resultsllava_7b_continuous_results
