Team Ai
Datasetpublic

shanexf/packaged-snack-nutrition-data

24-679 (Fall 2026): Packaged Snack Nutrition shanexf/packaged-snack-nutrition-data Nutrition Facts values for 30 packaged snack products, hand-collected from package labels, plus explicitly marked synthetic variants. The classroom task is multiclass classification: predict the snack category (chips, crackers, cookies, candy, granola_bars) from eight per-serving nutrition numbers. Purpose Built for the 24-679 (Fall 2026, Carnegie Mellon University) assignment on… See the full description on the dataset page: https://huggingface.co/datasets/shanexf/packaged-snack-nutrition-data.

sourceHugging Facecc-by-4.0updated 26d agoView on Hugging Face
0likes67downloads
Dataset Card

24-679 (Fall 2026): Packaged Snack Nutrition

shanexf/packaged-snack-nutrition-data

Nutrition Facts values for 30 packaged snack products, hand-collected from package labels, plus explicitly marked synthetic variants. The classroom task is multiclass classification: predict the snack category (chips, crackers, cookies, candy, granola_bars) from eight per-serving nutrition numbers.

Purpose

Built for the 24-679 (Fall 2026, Carnegie Mellon University) assignment on building and augmenting a small tabular dataset. It exists to practice data contracts, provenance tracking, label-preserving augmentation, and honest documentation on a dataset small enough to inspect row by row. It is a teaching artifact, not a nutrition reference.

Composition

Each row is one packaged snack product (or a synthetic variant of one). There are 30 unique real products, six in each of five shelf categories (chips, crackers, cookies, candy, granola_bars). Every row carries eight numeric features, the category target, and provenance fields. No text, images, or personal data are included.

Stored field / groupMeaning and modeling role
serving_size_g, servings_per_containerContinuous features from the label; strictly positive.
calories, total_fat_g, sodium_mg, carbs_g, sugar_g, protein_gContinuous per-serving nutrition features; nonnegative; sugar never exceeds carbs.
categoryNominal classification target with five classes. Synthetic rows inherit it from their parents.
source_id, parent_id, second_parent_idUnique example key and original source keys; provenance only.
augmentation, is_augmented, mix_weightMethod, synthetic flag, and primary-parent weight; exclude from predictors.

Numeric features are stored as floats in every split. The machine-readable feature metadata at the top of this card preserves every exact column name and storage type.

Collection

The 30 products were selected and measured by the author in September 2026. All eight features were transcribed directly from each product's printed Nutrition Facts panel (serving size in grams, servings per container, and per-serving calories, total fat, sodium, total carbohydrate, sugar, and protein). The category is the type of snack the product is sold as, recorded at collection time. The preparation notebook reads Package food nutritional value.csv, keeps the collected sample_id as source_id, and checks that every value is finite and nonnegative, that serving fields are positive, that sugar never exceeds total carbohydrate. These checks establish valid domains, not the accuracy of a manufacturer's label.

Labels

category is the prediction target, with five balanced classes in the original split (six products each). It is a property of the product type, assigned at collection, not a threshold applied to any nutrition value; the model's job is to recover it from nutrition alone. Synthetic rows always inherit the label of their parent(s), so no augmentation method creates, flips, or blends a label.

Preprocessing and augmentation

Original rows were randomly partitioned into train, validation, and test before augmentation (holdout fraction 30%, stratified by category, seed 24679; test receives 50% of that holdout, seed 24680). Only training rows serve as parents. This run requests 5 copies per method and training row before filtering; unchanged rows and repeated feature combinations for a given primary parent are removed. Every synthetic value is rounded to the resolution printed on a Nutrition Facts label (1 g serving size, 0.5 servings, 5 kcal, 0.5 g fat/sugar/protein, 5 mg sodium, 1 g carbohydrate) and clipped to its domain.

  • —Additive numeric jitter: Gaussian noise on every feature with standard deviation max(5% of the training IQR, one label step). Justification: mimics label rounding and small formulation differences between batches or brands.
  • —Multiplicative numeric scaling: each feature multiplied by an independent factor in 0.90–1.10. Justification: mimics a slightly different serving size or a near-identical product from another brand.
  • —Within-class Mixup: two distinct training parents from the same category blended with one primary weight in 0.60–0.90 across all eight features. Justification: a convex combination of two cookies is a plausible cookie, and the shared class makes the label unambiguous.
  • —Within-class nearest-neighbor interpolation (SMOTE-style): a child placed at a uniform step between a training product and one of its 3 nearest same-category neighbors in standardized feature space. Justification: fills locally dense regions of each class rather than spanning the whole class.

All four methods are label-preserving by construction and every child records its parent(s) and weight. Original and single-parent rows repeat the primary key in second_parent_id and use mix_weight=1.0. Categorical perturbation and CTGAN are not used.

Augmentation methodStored rows
additive_numeric_jitter105
multiplicative_numeric_scale105
within_class_mixup99
within_class_neighbor_interpolation90

Splits

These counts are computed from the packaged splits for this run.

SplitOriginal rowsSynthetic rowsTotal rows
original30030
augmented0399399
train21399420
validation404
test505
SplitcategoryRows
originalcandy6
originalchips6
originalcookies6
originalcrackers6
originalgranola_bars6
augmentedcandy76
augmentedchips74
augmentedcookies76
augmentedcrackers96
augmentedgranola_bars77
traincandy80
trainchips78
traincookies80
traincrackers101
traingranola_bars81
validationcandy1
validationchips1
validationcookies1
validationgranola_bars1
testcandy1
testchips1
testcookies1
testcrackers1
testgranola_bars1

original holds the 30 collected rows unchanged and augmented holds only synthetic rows; these two splits are the assignment deliverable. train = training originals + every synthetic row; validation and test = untouched holdout originals. Every parent of every synthetic row is a training original, so no holdout product has a synthetic child anywhere in the dataset, and augmented is exactly the synthetic portion of train. Keep these boundaries fixed for downstream comparisons; changing or reordering the source CSV changes them.

Intended use and limitations

Use for teaching tabular data contracts, provenance, augmentation, and small-sample multiclass classification. Compare training on the original training rows alone against the train split, using the same validation and test rows; report accuracy and macro-F1 against a majority-class baseline. With four or five holdout rows per partition, scores will be noisy.

Limitations: thirty products from a small number of brands and one shopper's selection do not represent the packaged-food market. Categories overlap in nutrition space (a sweet granola bar can resemble a cookie), and the category boundary is a shelf convention, not a nutrition fact. Synthetic rows add spread around real products but not new independent products, and a domain-valid synthetic row may still describe a product nobody sells. Nutrition label values are manufacturer-reported and rounded.

Ethical notes

The data describe products, not people; no personal information, purchase history, or consumption data is included, and the collector is not identifiable from the rows. Values were transcribed from public package labels; brand and product names are deliberately omitted so no manufacturer is singled out. The category label is a descriptive shelf category and carries no health judgment. Do not use this dataset for nutrition guidance, dietary advice, product ranking, or any consequential decision.

License

Released under CC BY 4.0. Nutrition Facts values are factual data; the compilation, labels, synthetic rows, and this documentation are the author's contribution and may be reused with attribution.

AI usage disclosure

The preparation notebook and this card were written with assistance from Claude (Anthropic), using the 24-679 Tabular Data class notebook as the template. The product selection, all 30 measurements, and the category labels were collected by the author without AI. The author reviewed, ran, and verified every step of the notebook and is responsible for the dataset and its documentation.

Load

python
from datasets import load_dataset
ds = load_dataset("shanexf/packaged-snack-nutrition-data")
original, augmented = ds["original"], ds["augmented"]
# Train with ds["train"], choose settings with ds["validation"], then score ds["test"].

Regenerate this card with the preparation notebook after changing the data; its counts are calculated from the actual packaged splits. The YAML schema and split configuration are preserved from the upload.