small-code
testing_codealpaca_small
Dataset Card for "testing_codealpaca_small"
More Information needed
code-comments-small
Comment Dataset
Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit.
Files are grouped as <dataset>/<language>/part-*.parquet.
The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language.
Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata.
For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.motherlode-code-small-en-v0.1-evidence
Evidence for motherlode-code-small-en-v0.1
Every figure on the card of RiverRider/motherlode-code-small-en-v0.1
comes from a file here. Each figure also has a row in Sunstone North Labs' claims ledger, named in the first column. The
table maps each row to its file and field. The per-instance and per-query files let you recompute a paired test without
our code or our hardware.
"The soup" is the model: the parameter-wise mean of two fine-tunes of bge-small-en-v1.5, a2 and d1, at… See the full description on the dataset page: https://huggingface.co/datasets/RiverRider/motherlode-code-small-en-v0.1-evidence.hardware_code_and_sec_smalltokenized-codeparrot-ds-smallagent-trace-privacy-scrubber-codex-traces
