datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodonTranslator-data
CodonTranslator Data
This repository contains the final public training-data release used for CodonTranslator.
Contents
train/: representative-only training shards
val/: representative-only validation shards
test/: representative-only held-out test shards
embeddings_v2/: precomputed species conditioning embeddings used in training
_work/final_representative_counts.json: final released split sizes
_work/split_report.json: split audit report… See the full description on the dataset page: https://huggingface.co/datasets/alegendaryfish/CodonTranslator-data.TinySFT
TinySFT
TinySFT contains 848,063 supervised fine-tuning samples for SFT of small models. Each sample includes a system message, preserves the native tool-calling structure, and every assistant turn carries a reasoning_effort label (none / low / high / max) indicating how much thinking that turn should require.
The dataset has 848,063 rows / 980,750 assistant turns / 1 parquet file, about 1.4 GiB in total. Rows are globally shuffled; each file is a random sample of the full… See the full description on the dataset page: https://huggingface.co/datasets/CodonProject/TinySFT.
