datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swallow-code-v2
SwallowCode-v2
Resources
๐ arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
๐ค Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
๐ป What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited toโฆ See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.swallow-code
SwallowCode
Notice
May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility.
May 21, 2025: ClamAV has flagged โWin.Trojan.MSShellcode-88โ inโฆ See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.llm0to1-pt-tokenized-code
LLM0to1 ์ฌ์ ํ์ต ํ ํฐํ๋ณธ โ ์ฝ๋
10B ๊ท๋ชจ ํ/์ ์ด์ค์ธ์ด LLM LLM0to1-10b ๋ฅผ ๋ฐ๋ฅ๋ถํฐ ํ์ตํ ๋ ์ค์ ๋ก ํฌ์
๋
์ฝ๋ ์ฝํผ์ค์ ํ ํฐํ๋ณธ. ์ด 27์ข
/ 59.1B ํ ํฐ.
์ ์๋ฌธ ํ
์คํธ๊ฐ ์๋๋ผ ํ ํฐํ๋ณธ์ธ๊ฐ
์ด ์ฝํผ์ค๋ค์ ๊ณต๊ฐ ๋ฐ์ดํฐ์
์ ์คํธ๋ฆฌ๋ฐ์ผ๋ก ๋ฐ์ ๊ณง๋ฐ๋ก ํ ํฐํํ๊ณ ์ค๊ฐ ํ
์คํธ๋ฅผ ๋ณด๊ดํ์ง ์์๋ค.
๋ฐ๋ผ์ ์ด .ds ํ์ผ์ด ํ์ต์ ๋ค์ด๊ฐ ๋ฐ์ดํฐ์ ์ ์ผํ ์ฌ๋ณธ์ด๋ค.
๋จ์ ๋ง ์๋ ๊ฑด ์๋๋ค. ํ ํฐํ๋ณธ์ ํ์ต ์
๋ ฅ ๊ทธ ์์ฒด์ด๋ฏ๋ก,
์ฌํ ํฐํ ๊ณผ์ ์์ ์๊ธธ ์ ์๋ ์ฐจ์ด ์์ด ํ์ต์ ๊ทธ๋๋ก ์ฌํํ ์ ์๋ค.
์๋ณธ ์ถ์ฒ
bigcode/starcoderdata(์ธ์ด๋ณ ์๋ธ์
) ยท bigcode/commitpackft ยท bigcode/jupyter-code-text-pairs ยท deepmind/code_contests
code_c ์ code_c2 ์ฒ๋ผ 2 ๊ฐ ๋ถ์ ๊ฒ์ ์ ์ ๊ธฐโฆ See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-code.swallow-code-v0.1
What is it?
Swallow-code-v0.1 consists of 4 staged dataset subsets and are filtered from bigcode/the-stack-v2-train-smol-ids.
What is being released?
The dataset is released in four versions:
Swallow Code v0.1 stage 1: 36B tokens, 41M documents containing Python scripts.
Swallow Code v0.1 stage 2: 31B tokens, 37M documents containing Python scripts that are syntax error-free.
Swallow Code v0.1 stage 3: 20B tokens, 24M documents containing Python scripts that are filteredโฆ See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v0.1.LLMxCPG-Code
LLMxCPG-Code
This dataset contains the raw C files used in our paper:LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language Models
๐ Models
The models LLMxCPG-Q and LLMxCPG-D are available in the following Hugging Face collection:๐ https://huggingface.co/collections/QCRI/llmxcpg-6855f80e601774b43eba2d14
๐ป Source Code
The source code for LLMxCPG can be found here:๐ https://github.com/qcri/llmxcpg
code-nomist-llm-datasetCode Nomist ้กน็ฎๅพฎ่ฐๆฐๆฎ้
็จไบๅคงๆจกๅๅพฎ่ฐไฝฟ็จ๏ผๅ
ๅซๆ ผๅผๅๅ็้ฎ้ข๏ผไปฅๅๅฏนๅบ็ญๆกๆฏไธชๅ็งฐไฝฟ็จ | ็ฌฆๅทๅๅฒใ
