datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodonTransformer
CodonTransformer Dataset
A comprehensive compilation of 1,001,197 DNA and protein sequence pairs, sourced from 164 organisms across Eukaryotes, Bacteria, and Archaea.
This dataset provides a rich resource for various computational biology and bioinformatics applications such as studying gene sequences, codon usage, and protein expression across diverse species.
Dataset Contents
1,001,197 DNA-protein sequence pairs
Sequences from 164 organisms, including:
Eukaryotes:… See the full description on the dataset page: https://huggingface.co/datasets/adibvafa/CodonTransformer.TinySFT
TinySFT
TinySFT contains 848,063 supervised fine-tuning samples for SFT of small models. Each sample includes a system message, preserves the native tool-calling structure, and every assistant turn carries a reasoning_effort label (none / low / high / max) indicating how much thinking that turn should require.
The dataset has 848,063 rows / 980,750 assistant turns / 1 parquet file, about 1.4 GiB in total. Rows are globally shuffled; each file is a random sample of the full… See the full description on the dataset page: https://huggingface.co/datasets/CodonProject/TinySFT.
