datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodeLlama-2-20k
CodeLlama-2-20k: A Llama 2 Version of CodeAlpaca
This dataset is the sahil2801/CodeAlpaca-20k dataset with the Llama 2 prompt format described here.
Here is the code I used to format it:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset('sahil2801/CodeAlpaca-20k')
# Define a function to merge the three columns into one
def merge_columns(example):
if example['input']:
merged = f"<s>[INST] <<SYS>>\nBelow is an instruction that describes a task… See the full description on the dataset page: https://huggingface.co/datasets/mlabonne/CodeLlama-2-20k.llama-q-a-code
llama-q-a-code
Description
This dataset consists of synthetic coding-focused question-and-answer pairs
designed for training or fine-tuning models to become better coding assistants.
The dataset covers a broad range of programming topics,
including Python, JavaScript, SQL, C++, and Git.
Disclaimer
Limitations
To ensure efficient generation,
responses were generated with a max_new_tokens limit of 512.
Consequently, some complex coding answers
or long… See the full description on the dataset page: https://huggingface.co/datasets/Fu01978/llama-q-a-code.local-code-arena-mbpp-codellama_7b
Local Code Arena Telemetry: MBPP Benchmark on Code Llama 7B
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against Meta's Code Llama 7B model.
This specific partition documents the baseline performance of early-generation specialized code engines, establishing a vital chronological anchor point to measure modern post-training alignment improvements.… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-codellama_7b.
