Team Ai
Datasetpublic

LiteCoder/LiteCoder_SourceCode

LiteCoder Experiment Reproducing package To run the pre-train objective use the following scripts: Reproduce LiteCoder with all objectives: Navigate the folder Pre-training containing the LiteCoder.py file Then, run Python LiteCoder.py --train-tt --train-cs --train-pd The pretrained model is released on hugging face, therefore it automatically loads. To run the ablation studies: Ablation 1: Python LiteCoder.py --train-tt Ablation 2: Python LiteCoder.py --train-tt… See the full description on the dataset page: https://huggingface.co/datasets/LiteCoder/LiteCoder_SourceCode.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes9downloads
README.md55 linesDownload Raw Back to root
1# LiteCoder Experiment Reproducing package2 3- To run the pre-train objective use the following scripts:4  5  - Reproduce LiteCoder with all objectives:6    7    - Navigate the folder `Pre-training` containing the `LiteCoder.py` file8    - Then, run `Python LiteCoder.py --train-tt --train-cs --train-pd`9      10      - The pretrained model is released on [hugging face](https://huggingface.co/LiteCoder/LiteCoder_pretrained), therefore it automatically loads.11 12  - To run the ablation studies:13    14    - Ablation 1: `Python LiteCoder.py --train-tt`15    - Ablation 2: `Python LiteCoder.py --train-tt --train-cs`16    - Ablation 3: `Python LiteCoder.py --train-tt --train-cs --train-pd`17 18- To `Fine-tuning` LiteCoder on downstream tasks:19  20  - Navigate to the `Fine-tuning` folder and then `Downstream task` folder:21     22    - Code Clone Detection:23      - Follow the instruction of `readme.md` file.24        25    - Code Translation:26      27      - Run `setup.sh` file.28      - Navigate to the `scripts/finetune` and run `translate.sh` file.29 30- To extract the programming language features (i.e., `token type`, `code sememe`, and `code dependencies`)31  - We used open source datasets to extract language features. we released the extracted datasets on the Hugging Face:32    - `LT_Java` :  [LiteCoder/LT_Java](https://huggingface.co/datasets/LiteCoder/LT_Java)33    - `LT_Python` :  [LiteCoder/LT_Python](https://huggingface.co/datasets/LiteCoder/LT_Python)34    - `LT_Java_Dependency` :  [LiteCoder/LT_Java_Dependency](https://huggingface.co/datasets/LiteCoder/LT_Java_Dependency)35 36  - Navigate to the utils directory:37    - Use either the `Java` or `Python` notebook file to run over your dataset.38    - Run the cells, for which, you want to extract the features.39 40- Dependencies:41  - Feature extraction dependencies:42    ```bash43    - pip install ast-comments44    - pip install ast45    - pip install javalang46    - pip install tree-sitter47    48  - Model training dependencies:49    ``` bash50    - pip install transformers 51    - pip install datasets52    - pip install pytorch_lightning53    - pip install torch54 55  - Or `pip install -r requirements.txt`