datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biomap-research-enzyme_catalytic_efficiency
enzyme_catalytic_efficiency
Sourced from biomap-research/enzyme_catalytic_efficiency and prepared for Hugging Face datasets usage.
Intended use
Prediction of enzyme catalytic efficiency.
Provenance
Card generated (UTC): 2026-10-06
Upstream revision: unpinned (default branch; pin source_revision to a commit to reproduce the input).
Upstream dataset card:… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-enzyme_catalytic_efficiency.tool-call-efficiency
tool-call-efficiency
Made with the whileai SDK · Collections: Efficiency, Start here: foundational post-training datasets
Teach an agent to make every tool call count.
An agent that calls a tool twice with the same arguments, looks up what
the user just told it, or keeps calling after the task is done is slow,
expensive, and harder to trust. Ask a base Qwen3-4B to work through
1,133 tool-using tasks across six agents and it does this a lot:
only 52% of its 6,681 rollouts finish… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tool-call-efficiency.newsmner-data-efficiencyperovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.enzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions applicable… See the full description on the dataset page: https://huggingface.co/datasets/biomap-research/enzyme_catalytic_efficiency.SiN-photonic-waveguide-loss-efficiency
💎 SiN Photonic Waveguide Loss & Efficiency Dataset
🔬 90,000 synthetic rows of silicon nitride (Si₃N₄) waveguide parameters linking geometry, fabrication, and operating conditions to loss and efficiency metrics, for regression modeling, simulation, and fine-tuning.
⚠️ Disclaimer: All rows are synthetically generated. Parameter ranges are informed by published SiN platform values, but no row is a foundry measurement. The data_source column is a schema field; every row in this… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/SiN-photonic-waveguide-loss-efficiency.agent-ui-efficiency-scores
Agent UI Efficiency Scores
Flat lab-synthetic bakeoff table for the public question: which agent UI is cheapest for a given lab task?
Author
Akash Premkumar (akashnaren)
License
Apache-2.0
Hub files
train.jsonl (63), test.jsonl (14), optional scores.jsonl (77 full)
Related
agent-ui-sft, agent-ui-human, ui-mode-router, agent-ui-mode-pairs
Scope
Rows are original lab fiction for a public agent-UI research question. Identifiers are invented.… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-efficiency-scores.goal-contribution-efficiency-top-5-leagues
⚽ Football Player Performance Analysis (2019-2020)
📋 Project Overview
This project explores player performance data across the Top 5 European Leagues (England, France, Germany, Italy, and Spain) during the 2019-2020 season. Using a dataset of 2,661 players and 22 columns, we analyze the relationship between actual scoring output and expected metrics.
❓ Research Question
"Do top-tier goal contributors consistently exceed their expected metrics (xG and xA), or… See the full description on the dataset page: https://huggingface.co/datasets/rotemknat/goal-contribution-efficiency-top-5-leagues.quant_eval_efficiency_and_footprint
quant_eval — Efficiency and footprint
One row per published run: stored weight artifact bytes before and after quantization, compression ratio, observed evaluation wall-time ratio with an explicit direction label, the accelerator used on each lane, and token throughput.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_efficiency_and_footprint.douvras-bitnet-ptbr-efficiency
Douvras BitNet PT-BR Efficiency Benchmark
Benchmark sintético de roteamento de workloads para avaliar posteriormente BitNet, Qwen,
SmolLM e Tucano em português brasileiro. Esta versão contém zero medições de GPU, RAM,
energia, latência ou qualidade; os registros carregam measured: false. O test está congelado
e as famílias não atravessam os splits.
O dataset não contém pesos de modelos, dados pessoais ou conteúdo de terceiros.
enzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions… See the full description on the dataset page: https://huggingface.co/datasets/chenchaozhao/enzyme_catalytic_efficiency.translation-efficiency-human
Multitask Translational Efficiency Prediction
Overview
Understanding the rules of translational control in mammalian cells is a fundamental challenge in genomics. This dataset is from a study by Zheng et al. (2025), which created a comprehensive, transcriptome-wide atlas of translation efficiency (TE) measurements across a wide array of human and mouse cell types.
The dataset was generated by uniformly processing and quality-controlling thousands of ribosome profiling and… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/translation-efficiency-human.Giant-Freshwater-Prawn-Growth-And-Feed-Efficiency-Causal-Reasoning-QA-Dataset
Giant Freshwater Prawn Growth and Feed Efficiency Causal Reasoning QA Dataset
This dataset provides question-and-answer text about growth performance and feed efficiency in giant freshwater prawns, covering feed, feeding methods, stocking density, water temperature, molting stage, and other farming factors. Each record includes a question, background and evidence, an answer, causal reasoning, a causal caveat, and a management implication, supporting evidence-based analysis of… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Giant-Freshwater-Prawn-Growth-And-Feed-Efficiency-Causal-Reasoning-QA-Dataset.resource_allocation_telecom_spectral_efficiency_rician_k_12_instruct_10kREFUEL_it2_mask2_data_efficiencyafrica-synth-energy-efficiency-appliances-africa-niger
Africa Synth Energy Efficiency Appliances Africa Niger | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-efficiency-appliances-africa-niger.resource_allocation_telecom_spectral_efficiency_area_250_instructbiomap-research-enzyme_catalytic_efficiency
enzyme_catalytic_efficiency
Sourced from biomap-research/enzyme_catalytic_efficiency and prepared for Hugging Face datasets usage.
Intended use
Prediction of enzyme catalytic efficiency.
Provenance
Card generated (UTC): 2026-10-05
Upstream revision: unpinned (default branch; pin source_revision to a commit to reproduce the input).
Upstream dataset card:… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/biomap-research-enzyme_catalytic_efficiency.resource_allocation_telecom_energy_efficiency_area_350_instructF1-driver-car-harmonic-efficiency-and-energy-waste-mapping-v0.1What this dataset tests
Whether a system can detectharmonic inefficiency in driver-car coupling.
Focus
Overcorrection loopsoscillation signaturesenergy leaksegment efficiency rank
Required outputs
harmonic waste index
correction loop density
oscillation signature type
energy leak score
efficiency rank by segment
All scores0 to 1
Highermeans more waste.
africa-synth-agriculture-irrigation-access-efficiency-africa-all
Africa Synth Agriculture Irrigation Access Efficiency Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-irrigation-access-efficiency-africa-all.resource_allocation_telecom_spectral_efficiency_rician_k_2_instruct_10kenzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions applicable… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/enzyme_catalytic_efficiency.reasoning_efficiency
Reasoning Efficiency Evaluation Artifact
Anonymous review dataset accompanying the NeurIPS 2026 Evaluations & Datasets submission
“Diagnosing Reasoning Efficiency with Trace-Optional Evaluation”.
The artifact contains benchmark instances, raw visible model outputs, token/count metadata,
correctness and truncation flags, native workload metadata, derived model-level metrics,
and decomposition tables used by the paper.
Files
instances/*.jsonl.gz: benchmark prompts, gold… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-efficiency-authors/reasoning_efficiency.resource_allocation_telecom_energy_efficiency_rician_k_12_instruct_10kresource_allocation_telecom_energy_efficiency_rician_k_4_instruct_10kresource_allocation_telecom_spectral_efficiency_rician_k_10_instruct_10kEfficiency_smrresource_allocation_telecom_energy_efficiency_area_150_instructresource_allocation_telecom_energy_efficiency_rician_k_4_instruct_1k
