Team Ai
Datasetpublic

LLM-EDA/vgen_cpp

Dataset Card for Opencores In the process of continual pre-training, we utilized the publicly available VGen dataset. VGen aggregates Verilog repositories from GitHub, systematically filters out duplicates and excessively large files, and retains only those files containing \texttt{module} and \texttt{endmodule} statements. We also incorporated the CodeSearchNet dataset \cite{codesearchnet}, which contains approximately 40MB function codes and their documentation.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-EDA/vgen_cpp.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes27downloads
5 commits on main
a1846d02y ago

Upload README.md

LLM-EDA
d3f64082y ago

Upload README.md

WANG Ning
d7e21f12y ago

Create README.md

WANG Ning
0344e2a2y ago

Upload shuffled_pretrain.jsonl

WANG Ning
05aa11d2y ago

initial commit

WANG Ning