datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
on-device-auction-audit
On-Device Auction Audit
Data for When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions, Dipankar Sarkar, Skelf Research.
Paper: https://arxiv.org/abs/2609.33312
Code and full artifact: github.com/sarkar-dipankar/on-device-auction-audit (this package is built from commit 1d1a05b; the data files are unchanged since 5a6ad86).
The question
Privacy pushes ad decisions onto the device. The auction runs locally, next to… See the full description on the dataset page: https://huggingface.co/datasets/skelfresearch/on-device-auction-audit.healthbench
THE CODE IS CURRENTLY BROKEN BUT THE DATASET IS GOOD!!
HealthBench Implementation for using Opensource Judges
Easy-to-use implementation of OpenAI's HealthBench evaluation benchmark with support for any OpenAI API-compatible model as both the system under test and the judge.
Developed by: Nisten Tahiraj / OnDeviceMednotes
License: MIT
Paper: HealthBench: Evaluating Large Language Models Towards Improved Human Health
Overview
This repository contains tools… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/healthbench.on-device-latency
On-Device Latency Benchmark
Real-world inference latency data for mobile-optimized LLMs, measured on actual phone hardware.
Hardware
Spec
Value
Device
Samsung S20 FE 5G
SoC
Snapdragon 865
RAM
8GB
OS
Android 13
Runtime
llama.cpp (4 threads)
Metrics
tokens_per_sec — Generation speed during inference
latency_ms_per_token — Time per generated token
ram_usage_mb — Peak RAM during inference
file_size_mb — GGUF model file size… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/on-device-latency.
