LiteCoder/LiteCoder_SourceCode
LiteCoder Experiment Reproducing package To run the pre-train objective use the following scripts: Reproduce LiteCoder with all objectives: Navigate the folder Pre-training containing the LiteCoder.py file Then, run Python LiteCoder.py --train-tt --train-cs --train-pd The pretrained model is released on hugging face, therefore it automatically loads. To run the ablation studies: Ablation 1: Python LiteCoder.py --train-tt Ablation 2: Python LiteCoder.py --train-tt… See the full description on the dataset page: https://huggingface.co/datasets/LiteCoder/LiteCoder_SourceCode.
09
1# LiteCoder Experiment Reproducing package2 3- To run the pre-train objective use the following scripts:4 5 - Reproduce LiteCoder with all objectives:6 7 - Navigate the folder `Pre-training` containing the `LiteCoder.py` file8 - Then, run `Python LiteCoder.py --train-tt --train-cs --train-pd`9 10 - The pretrained model is released on [hugging face](https://huggingface.co/LiteCoder/LiteCoder_pretrained), therefore it automatically loads.11 12 - To run the ablation studies:13 14 - Ablation 1: `Python LiteCoder.py --train-tt`15 - Ablation 2: `Python LiteCoder.py --train-tt --train-cs`16 - Ablation 3: `Python LiteCoder.py --train-tt --train-cs --train-pd`17 18- To `Fine-tuning` LiteCoder on downstream tasks:19 20 - Navigate to the `Fine-tuning` folder and then `Downstream task` folder:21 22 - Code Clone Detection:23 - Follow the instruction of `readme.md` file.24 25 - Code Translation:26 27 - Run `setup.sh` file.28 - Navigate to the `scripts/finetune` and run `translate.sh` file.29 30- To extract the programming language features (i.e., `token type`, `code sememe`, and `code dependencies`)31 - We used open source datasets to extract language features. we released the extracted datasets on the Hugging Face:32 - `LT_Java` : [LiteCoder/LT_Java](https://huggingface.co/datasets/LiteCoder/LT_Java)33 - `LT_Python` : [LiteCoder/LT_Python](https://huggingface.co/datasets/LiteCoder/LT_Python)34 - `LT_Java_Dependency` : [LiteCoder/LT_Java_Dependency](https://huggingface.co/datasets/LiteCoder/LT_Java_Dependency)35 36 - Navigate to the utils directory:37 - Use either the `Java` or `Python` notebook file to run over your dataset.38 - Run the cells, for which, you want to extract the features.39 40- Dependencies:41 - Feature extraction dependencies:42 ```bash43 - pip install ast-comments44 - pip install ast45 - pip install javalang46 - pip install tree-sitter47 48 - Model training dependencies:49 ``` bash50 - pip install transformers 51 - pip install datasets52 - pip install pytorch_lightning53 - pip install torch54 55 - Or `pip install -r requirements.txt`