AINovice2005/carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.
carbon-cpu-enriched-sequences
A CPU-enriched subset of the [carbon pretraining corpus (`eukaryote_generator`)](https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
The dataset contains both the original Carbon fields and derived enrichment features.
Dataset Summary
carbon-cpu-enriched-sequences provides a reusable CPU-side canonical representation of the Carbon genomic corpus before computationally expensive GPU enrichment.
The original Carbon source corpus contains approximately 46.3 million records. This release contains 32,410,000 records, corresponding to approximately 70% of the source row population. The release was selected as a practical balance between corpus representation, CPU processing cost, storage requirements, and downstream GPU compute availability.
The dataset preserves the original Carbon source representation while adding CPU-derived enrichment features:
Original Carbon columns
+
CPU-derived enrichment columns
↓
CPU-enriched Parquet datasetThe current materialized release contains approximately 25 columns, consisting of approximately 14 source/raw columns and 11 derived enrichment columns. The exact schema should be inspected from the Parquet files, as future enrichment revisions may add or modify derived fields.
CPU enrichment includes sequence-level and metadata-derived information such as sequence characteristics, coding status, strand/orientation information, and taxonomy depth. Coding sequences are identified using the observed Carbon representation <cds>.
The dataset is stored as Apache Parquet with Zstandard compression, using approximately 50,000 rows per shard.
Intended Uses
This dataset is intended for research and data-engineering workflows involving genomic sequence corpora, including:
- Genomic data quality and validation analysis
- Sequence-distribution and statistical analysis
- CPU-side preprocessing research
- Dataset sampling and cohort construction
- GPU enrichment experiments
- Embedding generation
- Nearest-neighbor and clustering analysis
- Genomic foundation-model data engineering
- Benchmarking large-scale data-processing pipelines
- Construction of downstream enriched datasets
For large-scale datasets, streaming access is recommended rather than loading the entire dataset into memory.
