Team Ai
Datasetpublic

AINovice2005/carbon-cpu-enriched-sequences

carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
0likes562downloads
Dataset Card

carbon-cpu-enriched-sequences

A CPU-enriched subset of the [carbon pretraining corpus (`eukaryote_generator`)](https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation.


Information of Features

FeatureTypeDescription
record_idstringNCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequencestringBeginning-of-sequence special token marking the start of a sequence record, e.g. <s>.
end_of_sequencestringEnd-of-sequence special token marking the end of a sequence record, e.g. </s>.
begin_of_genestringMarker identifying the beginning of a gene/feature segment, e.g. <bog>.
end_of_genestringMarker identifying the end of a gene/feature segment, e.g. <eog>.
gene_typestringAnnotation describing the type of genomic feature. In the observed data, <cds> represents a coding sequence and <fng> represents another gene-feature category.
species_typestringBroad biological category of the source sequence. The current data is represented as Eukaryota.
strandstringDNA strand/orientation associated with the sequence, represented by + or -.
sequencestringThe actual nucleotide sequence represented by the record. This is the primary biological sequence field.
molecule_typestringMolecular type of the sequence; the current data is represented as DNA.
topologystringStructural topology of the molecule, such as linear.
taxonomystringHierarchical taxonomic lineage of the sequence, with ranks represented as a semicolon-separated lineage, e.g. Eukaryota;Fungi;...;Agaricus.
startint64Starting position of the record within the pre-trained Carbon corpus. It is a corpus coordinate, not necessarily a genomic coordinate.
endint64Ending position of the record within the processed/tokenized Carbon corpus.
gc_contentfloat64Fraction of nucleotides that are G or C: (G + C) / sequence_length. It measures GC composition.
gc_skewfloat64Measures asymmetry between G and C: (G - C) / (G + C). Positive values indicate relative G enrichment; negative values indicate relative C enrichment.
sequence_lengthint64Number of nucleotide bases in the sequence.
gene_lengthint64Length of the associated gene/feature sequence. For the current records it commonly corresponds to sequence_length.
shannon_entropyfloat64Sequence-complexity measure based on nucleotide-frequency distribution.
kmer_frequency_vectorlistA 64-element vector containing normalized frequencies of all possible DNA 3-mers (4³ = 64). It captures local sequence-composition patterns.
strand_normalized_sequencestringSequence transformed into a consistent orientation so that strand direction does not create artificial differences in downstream analysis.
taxonomy_domainstringBroadest taxonomic/domain-level classification extracted from the taxonomy hierarchy.
taxonomy_depthint64Number of taxonomic levels represented in the lineage. It provides a compact measure of how deeply the sequence is classified.
is_coding_regionboolBoolean indicator of whether the record is classified as a coding sequence. The pipeline identifies <cds> as the coding gene type.
qc_flagstringQuality-control status assigned to the record, indicating whether it passes the pipeline's sequence-level QC criteria.

The dataset contains both the original Carbon fields and derived enrichment features.


Dataset Summary

carbon-cpu-enriched-sequences provides a reusable CPU-side canonical representation of the Carbon genomic corpus before computationally expensive GPU enrichment.

The original Carbon source corpus contains approximately 46.3 million records. This release contains 32,410,000 records, corresponding to approximately 70% of the source row population. The release was selected as a practical balance between corpus representation, CPU processing cost, storage requirements, and downstream GPU compute availability.

The dataset preserves the original Carbon source representation while adding CPU-derived enrichment features:

text
Original Carbon columns
        +
CPU-derived enrichment columns
        ↓
CPU-enriched Parquet dataset

The current materialized release contains approximately 25 columns, consisting of approximately 14 source/raw columns and 11 derived enrichment columns. The exact schema should be inspected from the Parquet files, as future enrichment revisions may add or modify derived fields.

CPU enrichment includes sequence-level and metadata-derived information such as sequence characteristics, coding status, strand/orientation information, and taxonomy depth. Coding sequences are identified using the observed Carbon representation <cds>.

The dataset is stored as Apache Parquet with Zstandard compression, using approximately 50,000 rows per shard.

Intended Uses

This dataset is intended for research and data-engineering workflows involving genomic sequence corpora, including:

  • —Genomic data quality and validation analysis
  • —Sequence-distribution and statistical analysis
  • —CPU-side preprocessing research
  • —Dataset sampling and cohort construction
  • —GPU enrichment experiments
  • —Embedding generation
  • —Nearest-neighbor and clustering analysis
  • —Genomic foundation-model data engineering
  • —Benchmarking large-scale data-processing pipelines
  • —Construction of downstream enriched datasets

For large-scale datasets, streaming access is recommended rather than loading the entire dataset into memory.