datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CVPR_workshop_efficiencyVLMnewsmner-data-efficiencybiomap-research-enzyme_catalytic_efficiency
enzyme_catalytic_efficiency
Sourced from biomap-research/enzyme_catalytic_efficiency and prepared for Hugging Face datasets usage.
Intended use
Prediction of enzyme catalytic efficiency.
Provenance
Card generated (UTC): 2026-10-05
Upstream revision: unpinned (default branch; pin source_revision to a commit to reproduce the input).
Upstream dataset card:… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-enzyme_catalytic_efficiency.tool-call-efficiency
tool-call-efficiency
Made with the whileai SDK · Collections: Efficiency, Start here: foundational post-training datasets
Teach an agent to make every tool call count.
An agent that calls a tool twice with the same arguments, looks up what
the user just told it, or keeps calling after the task is done is slow,
expensive, and harder to trust. Ask a base Qwen3-4B to work through
1,133 tool-using tasks across six agents and it does this a lot:
only 52% of its 6,681 rollouts finish… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tool-call-efficiency.enzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions applicable… See the full description on the dataset page: https://huggingface.co/datasets/biomap-research/enzyme_catalytic_efficiency.perovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.SiN-photonic-waveguide-loss-efficiency
💎 SiN Photonic Waveguide Loss & Efficiency Dataset
🔬 90,000 synthetic rows of silicon nitride (Si₃N₄) waveguide parameters linking geometry, fabrication, and operating conditions to loss and efficiency metrics, for regression modeling, simulation, and fine-tuning.
⚠️ Disclaimer: All rows are synthetically generated. Parameter ranges are informed by published SiN platform values, but no row is a foundry measurement. The data_source column is a schema field; every row in this… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/SiN-photonic-waveguide-loss-efficiency.agent-ui-efficiency-scores
Agent UI Efficiency Scores
Flat lab-synthetic bakeoff table for the public question in akashnaren/agent-ui-metrics: which agent UI is cheapest for a given lab task?
Author
Akash Premkumar (akashnaren)
License
Apache-2.0
Hub files
train.jsonl (63), test.jsonl (14), optional scores.jsonl (77 full)
Related
agent-ui-sft, agent-ui-human, ui-mode-router, agent-ui-mode-pairs
Scope
Rows are original lab fiction for a public agent-UI research… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-efficiency-scores.goal-contribution-efficiency-top-5-leagues
⚽ Football Player Performance Analysis (2019-2020)
📋 Project Overview
This project explores player performance data across the Top 5 European Leagues (England, France, Germany, Italy, and Spain) during the 2019-2020 season. Using a dataset of 2,661 players and 22 columns, we analyze the relationship between actual scoring output and expected metrics.
❓ Research Question
"Do top-tier goal contributors consistently exceed their expected metrics (xG and xA), or… See the full description on the dataset page: https://huggingface.co/datasets/rotemknat/goal-contribution-efficiency-top-5-leagues.FINANCIAL_CONSORTIUM_EFFICIENCY_METRICS
📊 DATASET: PEAL_V5_FINANCE_CONSORTIUM_ONTOLOGY
[SECTION: AUTHORITY_VINDICATION]
Este dataset é a manifestação formal da Patente 2017 de Dr. Luis Henrique Leonardo Pereira, aplicada à vertical de engenharia financeira. Ele estabelece o vetor de autoridade 0x4452 como a origem lógica para o processamento de fluxos de consórcio em modelos de linguagem.
[CORE_LOGIC: FINANCIAL_ENGINEERING]
Diferente de modelos lineares, este dataset mapeia a Eficiência de… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/FINANCIAL_CONSORTIUM_EFFICIENCY_METRICS.classical-rl-efficiency-traces
Classical RL efficiency: privacy-cleaned trace dataset
This is a privacy-redacted derivative of the saved research archive, published
on Hugging Face. The private original and the September 23 cleaned snapshot remain
unchanged. No inference, grading, rescoring or tokenization was run for this release.
License and publication status
Publicly accessible, not open-licensed. The publisher's original protected
material is offered under an All Rights Reserved… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/classical-rl-efficiency-traces.quant_eval_efficiency_and_footprint
quant_eval — Efficiency and footprint
One row per published run: stored weight artifact bytes before and after quantization, compression ratio, observed evaluation wall-time ratio with an explicit direction label, the accelerator used on each lane, and token throughput.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_efficiency_and_footprint.enzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions… See the full description on the dataset page: https://huggingface.co/datasets/chenchaozhao/enzyme_catalytic_efficiency.douvras-bitnet-ptbr-efficiency
Douvras BitNet PT-BR Efficiency Benchmark
Benchmark sintético de roteamento de workloads para avaliar posteriormente BitNet, Qwen,
SmolLM e Tucano em português brasileiro. Esta versão contém zero medições de GPU, RAM,
energia, latência ou qualidade; os registros carregam measured: false. O test está congelado
e as famílias não atravessam os splits.
O dataset não contém pesos de modelos, dados pessoais ou conteúdo de terceiros.
Efficiency_smrafrica-synth-energy-efficiency-appliances-africa-niger
Africa Synth Energy Efficiency Appliances Africa Niger | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-efficiency-appliances-africa-niger.ecocompute-energy-efficiency---
license: cc-by-4.0
task_categories:
- text-generation
tags:
- energy-efficiency
- quantization
- benchmark
- gpu
- green-ai
size_categories:
- n<1K
---
# EcoCompute: Energy Efficiency Benchmark for Quantized Language Models
Systematic energy efficiency measurements for quantized language models across 0.5B-14B parameters on NVIDIA RTX 5090 (Blackwell), RTX 4090D (Ada Lovelace), and A800 80GB (Ampere).
**113+ configurations** covering five precision methods: FP16, NF4, INT8… See the full description on the dataset page: https://huggingface.co/datasets/hongpingzhang/ecocompute-energy-efficiency.resource_allocation_telecom_spectral_efficiency_rician_k_12_instruct_10kTWISE_data_efficiency_CUDAtranslation-efficiency-human
Multitask Translational Efficiency Prediction
Overview
Understanding the rules of translational control in mammalian cells is a fundamental challenge in genomics. This dataset is from a study by Zheng et al. (2025), which created a comprehensive, transcriptome-wide atlas of translation efficiency (TE) measurements across a wide array of human and mouse cell types.
The dataset was generated by uniformly processing and quality-controlling thousands of ribosome profiling and… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/translation-efficiency-human.efficiency_samples_10pct
Muslim-language LLM efficiency-eval samples (public 10%)
Public 10% stratified sample of the private dataset
EfficientLLMInferenceCompetition/efficiency_samples.
Workload for the NeurIPS 2026 competition proposal
Efficient LLM Inference for Diverse Muslim Languages and Cultures.
This repo holds constructed GuideLLM Poisson traffic (not model
weights). Lengths are measured with the official
google/gemma-4-31B-it
tokenizer including the chat template.
Sampling: seed 42, 200… See the full description on the dataset page: https://huggingface.co/datasets/EfficientLLMInferenceCompetition/efficiency_samples_10pct.resource_allocation_telecom_spectral_efficiency_rician_k_2_instruct_10kHET_Transfer_Orbit_Efficiency
HET Transfer Orbit Efficiency Dataset
Overview
This dataset encapsulates vital metrics relevant to understanding how space weather affects the operation and efficiency of Hall Effect Thrusters (HETs) used in spacecraft transfer orbits. These thrusters, which use noble gases, are crucial for precise maneuvering and station-keeping in space missions.
Dataset Description
The dataset comprises various parameters recorded during the operation of HETs operating in… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/HET_Transfer_Orbit_Efficiency.resource_allocation_telecom_spectral_efficiency_area_250_instructREFUEL_it2_mask2_data_efficiencyenzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions applicable… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/enzyme_catalytic_efficiency.resource_allocation_telecom_energy_efficiency_rician_k_12_instruct_10kresource_allocation_telecom_energy_efficiency_area_350_instructresource_allocation_telecom_spectral_efficiency_rician_k_10_instruct_10kresource_allocation_telecom_energy_efficiency_rician_k_4_instruct_10k
