Team Ai
Modelpublic

MachineLearningLM/MachineLearningLM-7B-v1

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
33likes64downloads
README.md193 linesDownload Raw Back to root
1---2base_model:3- Qwen/Qwen2.5-7B-Instruct4license: apache-2.05pipeline_tag: text-generation6library_name: transformers7datasets:8- MachineLearningLM/machinelearninglm-scm-synthetic-tabularml9tags:10- Tabular Classification11---12 13# MachineLearningLM14 15This repository contains the model presented in the paper [MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining](https://huggingface.co/papers/2509.06806).16 17## Model Summary18 19Can LLMs learn from 1,000 in-context examples?20 21Introducing **MachineLearningLM** πŸ§ͺπŸ“Š β€” a model continuously pretrained on millions of synthetic tabular ML tasks, enabling robust many-shot in-context learning.22 23πŸ“ˆ **Scales from 8 to 1,024 examples**24 25πŸ“ˆ ​**​~15% improvement​**​ on unseen tabular tasks compared to o3-mini / GPT-5-mini / Qwen-2.5-7B-Instruct26 27🌲 ​**​Random-Forest–level numerical modeling robustness​**​28 29🧠 ​**​MMLU score: 75.4%​**​30 31πŸ“„ Read the paper:  https://huggingface.co/papers/2509.0680632 33   GitHub: https://github.com/HaoAreYuDong/MachineLearningLM34 35## Evaluation and Validation36 37We have developed an automated evaluation framework β€” simply configure the parameters to easily perform validation and evaluation. 38**The code is now open-sourced at our [GitHub repository](https://github.com/HaoAreYuDong/MachineLearningLM).**39 40**Quick Start**41 42```bash43pip install -r requirements.txt44python ./src/evaluation/model_pred/dl_model_pred.py \45  --input_dir ./demo_input.jsonl \46  --output_dir ./demo_output.jsonl \47  --model_name MachineLearningLM/MachineLearningLM-7B-v148```49**Pipeline**50```bash51# modify the evaluate_parameters.sh file52source evaluate_parameters.sh53 54# Option 1  End-to-End Pipeline55./scripts/evaluate_pipeline.sh56 57# Option 2  Parallel Processing58./scripts/multi_process/data_prep.sh59./scripts/multi_process/prompt_gen.sh  # For deep learning only60./scripts/multi_process/model_pred.sh61./scripts/multi_process/evaluation.sh62./scripts/multi_process/report.sh63 64# Option3   Sequential Processing65./scripts/single_process/data_prep.sh66./scripts/single_process/prompt_gen.sh  # For deep learning only67./scripts/single_process/model_pred.sh68./scripts/single_process/evaluation.sh69./scripts/single_process/report.sh70```71 72For more usage details, please visit our GitHub.73 74**Quants of Checkpoints**75 76https://huggingface.co/QuantFactory/MachineLearningLM-7B-v1-GGUF77 78 79## Tabicl Evaluation80 81**This part of the code needs to run in an environment with the tabicl and openpyxl libraries installed.**82 83The evaluation code for tabicl is placed separately in the `./src/evaluation/tabicl_evaluate.py` file. Use `./scripts/tabicl_evaluate.sh` to obtain the evaluation results for tabicl.84 85Use --datasets to specify the datasets to be evaluated, and --sample_sizes to indicate the number of shots. 86 87If multiple datasets need to be evaluated, separate them with spaces. To evaluate all CSV files in the input folder, use **all**.88 89## Prior_data90 91MachineLearningLM uses the code from tabicl to generate prior data.92 93Use `./scripts/generate_data.sh` to generate the prior data. It generates the corresponding .pt and .csv files, and normalizes the feature values in the CSV files to the range of 0–999, as we did in the paper.94 95### Parameter Introduction(refer to the comments in the file `tabicl\src\tabicl\prior\dataset.py`οΌ‰96 97**Data Scale & Structure**98 99| Parameter      | Type | Description                                             |100| :------------- | :--- | :------------------------------------------------------ |101| `min_features` | int  | Minimum number of features per dataset                  |102| `max_features` | int  | Maximum number of features per dataset                  |103| `max_classes`  | int  | Maximum number of target classes                        |104| `min_seq_len`  | int  | Minimum samples per dataset. Uses `max_seq_len` if None |105| `max_seq_len`  | int  | Maximum samples per dataset (Not IncludeοΌ‰             |106 107**Batch Configuration**108 109| Parameter              | Type | Description                                                  |110| :--------------------- | :--- | :----------------------------------------------------------- |111| `batch_size`           | int  | Total number of datasets to generate per batch               |112| `batch_size_per_gp`    | int  | Number of datasets per group (shared characteristics)        |113| `batch_size_per_subgp` | int  | Number of datasets per subgroup (similar causal structures). Defaults to `batch_size_per_gp` if None |114 115**Sequence Length Control**116 117| Parameter        | Type | Description                                                  |118| :--------------- | :--- | :----------------------------------------------------------- |119| `log_seq_len`    | bool | Sample sequence length from log-uniform distribution if True |120| `seq_len_per_gp` | bool | Sample sequence length per group (enables variable-sized datasets) |121| `replay_small`   | bool | Occasionally sample smaller sequences for model robustness   |122 123**Train-Test Split**124 125| Parameter        | Type      | Description                                                  |126| :--------------- | :-------- | :----------------------------------------------------------- |127| `min_train_size` | int/float | Start position/ratio for train split (int: absolute, float: fractional) |128| `max_train_size` | int/float | End position/ratio for train split (int: absolute, float: fractional) |129 130**Generation Method**131 132| Parameter    | Type | Description                                                  |133| :----------- | :--- | :----------------------------------------------------------- |134| `prior_type` | str  | Prior type: 'mlp_scm', 'tree_scm', or 'mix_scm' (random selection) |135| `fixed_hp`   | dict | Fixed structural configuration parameters                    |136| `sampled_hp` | dict | Parameters sampled during generation                         |137 138**Computation Settings**139 140| Parameter                  | Type | Description                                       |141| :------------------------- | :--- | :------------------------------------------------ |142| `n_jobs`                   | int  | Number of parallel jobs (-1 = use all processors) |143| `num_threads_per_generate` | int  | Number of threads per generation job              |144| `device`                   | str  | Computation device ('cpu' or 'cuda')              |145 146## Train147 148MachineLearningLM uses the LLaMA-Factory framework for training.149 150#### Training Environment Configuration151 152```bash153cd ./third_party/LLaMA-Factory154pip install -e ".[torch,metrics]" --no-build-isolation155pip install wandb156```157 158Use `./scripts/train.sh` for training.159 160## Project Structure161 162```163MachineLearningLM/164β”œβ”€β”€src/165|   β”œβ”€β”€evaluation/166β”‚   β”‚   β”œβ”€β”€ data_prep/          # Data preprocessing and chunking utilities167β”‚   β”‚   β”œβ”€β”€ prompt_gen/         # Prompt generation for deep learning models168β”‚   β”‚   β”œβ”€β”€ model_pred/         # Model inference (ML and DL prediction engines)169β”‚   β”‚   β”œβ”€β”€ result_proc/        # 5-layer evaluation architecture and metrics processing170β”‚   β”‚   β”œβ”€β”€ zero_summary/       # Result summarization and report generation171β”‚   β”‚   └── tabicl_evaluate.py172β”‚   └──prior_data173β”‚       └── pt_to_csv.py     174β”œβ”€β”€ scripts/175β”‚   β”œβ”€β”€ single_process/         # Sequential execution shell scripts176β”‚   β”œβ”€β”€ multi_process/          # Parallel execution shell scripts (with _mp suffix)177β”‚   β”œβ”€β”€ evaluate_parameters.sh  # Global parameter configuration178|   β”œβ”€β”€ evaluate_pipeline.sh    # automated pipeline179|   β”œβ”€β”€ generate_data.sh180|   β”œβ”€β”€ tabicl_evaluate.sh181|   └── train.sh182β”œβ”€β”€ datahub_inputs/183β”‚   β”œβ”€β”€ data_demo/          # Demo datasets for testing184β”‚   └── data_raw/           # Raw input datasets185β”œβ”€β”€ third_party/186β”‚   β”œβ”€β”€ tabicl/          187β”‚   └── LLaMA-Factory/   188β”œβ”€β”€ requirements.txt        # Python dependencies for Evaluation Framework189β”œβ”€β”€ README.md190β”œβ”€β”€ README_zh.md191β”œβ”€β”€ THIRD_PARTY_NOTICES.md192└── LICENSE193```