Team Ai
Apppublic

AUXteam/Critical_Code_Agent

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes
README.md412 linesDownload Raw Back to root
1---2title: Critical Code Agent3emoji: 🦀4colorFrom: red5colorTo: blue6sdk: docker7pinned: false8app_port: 78609---10<h1 align="center">11  <a href="https://github.com/SakanaAI/AI-Scientist/blob/main/docs/logo_2.png">12    <img src="docs/logo_2.png" width="215" /></a><br>13  <b>The AI Scientist: Towards Fully Automated</b><br>14  <b>Open-Ended Scientific Discovery 🧑‍🔬</b><br>15</h1>16 17<p align="center">18  📚 <a href="https://arxiv.org/abs/2408.06292">[Paper]</a> |19  📝 <a href="https://sakana.ai/ai-scientist/">[Blog Post]</a> |20  📂 <a href="https://drive.google.com/drive/folders/1G7A0wTqfXVa-cpexjk0oaXakaSJwffEt">[Drive Folder]</a>21</p>22 23One of the grand challenges of artificial intelligence is developing agents capable of conducting scientific research and discovering new knowledge. While frontier models have already been used to aid human scientists—for example, for brainstorming ideas or writing code—they still require extensive manual supervision or are heavily constrained to specific tasks.24 25We're excited to introduce **The AI Scientist**, the first comprehensive system for fully automatic scientific discovery, enabling Foundation Models such as Large Language Models (LLMs) to perform research independently.26 27We provide all runs and data from our paper [here](https://drive.google.com/drive/folders/1G7A0wTqfXVa-cpexjk0oaXakaSJwffEt?usp=sharing), where we run each base model on each template for approximately 50 ideas. We *highly* recommend reading through some of the [Claude papers](https://drive.google.com/drive/folders/1Mmpz6M1FK4q8e-SewgZcUzdeD0Q2zC39?usp=sharing) to get a sense of the system's strengths and weaknesses. Here are some example papers generated by **The AI Scientist** 📝:28 291. [DualScale Diffusion: Adaptive Feature Balancing for Low-Dimensional Generative Models](https://github.com/SakanaAI/AI-Scientist/blob/main/example_papers/adaptive_dual_scale_denoising.pdf)302. [Multi-scale Grid Noise Adaptation: Enhancing Diffusion Models For Low-dimensional Data](https://github.com/SakanaAI/AI-Scientist/blob/main/example_papers/grid_based_noise_adaptation.pdf)313. [GAN-Enhanced Diffusion: Boosting Sample Quality and Diversity](https://github.com/SakanaAI/AI-Scientist/blob/main/example_papers/gan_diffusion.pdf)324. [DualDiff: Enhancing Mode Capture in Low-dimensional Diffusion Models via Dual-expert Denoising](https://github.com/SakanaAI/AI-Scientist/tree/main/example_papers/dual_expert_denoiser.pdf) 335. [StyleFusion: Adaptive Multi-style Generation in Character-Level Language Models](https://github.com/SakanaAI/AI-Scientist/blob/main/example_papers/multi_style_adapter.pdf)346. [Adaptive Learning Rates for Transformers via Q-Learning](https://github.com/SakanaAI/AI-Scientist/tree/main/example_papers/rl_lr_adaptation.pdf)357. [Unlocking Grokking: A Comparative Study of Weight Initialization Strategies in Transformer Models](https://github.com/SakanaAI/AI-Scientist/tree/main/example_papers/weight_initialization_grokking.pdf)368. [Grokking Accelerated: Layer-wise Learning Rates for Transformer Generalization](https://github.com/SakanaAI/AI-Scientist/tree/main/example_papers/layerwise_lr_grokking.pdf)379. [Grokking Through Compression: Unveiling Sudden Generalization via Minimal Description Length](https://github.com/SakanaAI/AI-Scientist/tree/main/example_papers/mdl_grokking_correlation.pdf)3810. [Accelerating Mathematical Insight: Boosting Grokking Through Strategic Data Augmentation](https://github.com/SakanaAI/AI-Scientist/tree/main/example_papers/data_augmentation_grokking.pdf)39 40> **Note:**  41> **Caution!** This codebase will execute LLM-written code. There are various risks and challenges associated with this autonomy, including the use of potentially dangerous packages, web access, and potential spawning of processes. Use at your own discretion. Please make sure to [containerize](#containerization) and restrict web access appropriately.42 43<p align="center">44  <a href="https://github.com/SakanaAI/AI-Scientist/blob/main/example_papers/adaptive_dual_scale_denoising/adaptive_dual_scale_denoising.pdf"><img src="https://github.com/SakanaAI/AI-Scientist/blob/main/docs/anim-ai-scientist.gif" alt="Adaptive Dual Scale Denoising" width="80%" />45</a></p>46 47## Table of Contents48 491. [Introduction](#introduction)502. [Requirements](#requirements)51   - [Installation](#installation)52   - [Supported Models and API Keys](#supported-models-and-api-keys)533. [Setting Up the Templates](#setting-up-the-templates)54   - [NanoGPT Template](#nanogpt-template)55   - [2D Diffusion Template](#2d-diffusion-template)56   - [Grokking Template](#grokking-template)574. [Run AI Scientist Paper Generation Experiments](#run-ai-scientist-paper-generation-experiments)585. [Getting an LLM-Generated Paper Review](#getting-an-llm-generated-paper-review)596. [Making Your Own Template](#making-your-own-template)60   - [Community-Contributed Templates](#community-contributed-templates)617. [Template Resources](#template-resources)628. [Citing The AI Scientist](#citing-the-ai-scientist)639. [Frequently Asked Questions](#frequently-asked-questions)6410. [Containerization](#containerization)65 66## Introduction67 68We provide three templates, which were used in our paper, covering the following domains: **NanoGPT**, **2D Diffusion**, and **Grokking**. These templates enable The AI Scientist to generate ideas and conduct experiments in these areas. We accept contributions of new templates from the community, but please note that they are not maintained by us. All other templates beyond the three provided are community contributions.69 70## Requirements71 72This code is designed to run on Linux with NVIDIA GPUs using CUDA and PyTorch. Support for other GPU architectures may be possible by following the [PyTorch guidelines](https://pytorch.org/get-started/locally/). The current templates would likely take an infeasible amount of time on CPU-only machines. Running on other operating systems may require significant adjustments.73 74### Installation75 76```bash77conda create -n ai_scientist python=3.1178conda activate ai_scientist79# Install pdflatex80sudo apt-get install texlive-full81 82# Install PyPI requirements83pip install -r requirements.txt84```85 86**Note:** Installing `texlive-full` can take a long time. You may need to [hold Enter](https://askubuntu.com/questions/956006/pregenerating-context-markiv-format-this-may-take-some-time-takes-forever) during the installation.87 88### Supported Models and API Keys89 90We support a wide variety of models, including open-weight and API-only models. In general, we recommend using only frontier models above the capability of the original GPT-4. To see a full list of supported models, see [here](https://github.com/SakanaAI/AI-Scientist/blob/main/ai_scientist/llm.py).91 92#### OpenAI API (GPT-4o, GPT-4o-mini, o1 models)93 94By default, this uses the `OPENAI_API_KEY` environment variable.95 96#### Anthropic API (Claude Sonnet 3.5)97 98By default, this uses the `ANTHROPIC_API_KEY` environment variable.99 100##### Claude Models via Bedrock101 102For Claude models provided by [Amazon Bedrock](https://aws.amazon.com/bedrock/), please install these additional packages:103 104```bash105pip install anthropic[bedrock]106```107 108Next, specify a set of valid [AWS Credentials](https://docs.aws.amazon.com/cli/v1/userguide/cli-configure-envvars.html) and the target [AWS Region](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-regions.html):109 110Set the environment variables: `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_REGION_NAME`.111 112##### Claude Models via Vertex AI113 114For Claude models provided by [Vertex AI Model Garden](https://cloud.google.com/model-garden?hl=en), please install these additional packages:115 116```bash117pip install google-cloud-aiplatform118pip install anthropic[vertex]119```120 121Next, set up valid authentication for a [Google Cloud project](https://cloud.google.com/vertex-ai/docs/authentication), for example by providing the region and project ID:122 123```bash124export CLOUD_ML_REGION="REGION"           # for Model Garden call125export ANTHROPIC_VERTEX_PROJECT_ID="PROJECT_ID"  # for Model Garden call126export VERTEXAI_LOCATION="REGION"         # for Aider/LiteLLM call127export VERTEXAI_PROJECT="PROJECT_ID"      # for Aider/LiteLLM call128```129 130#### DeepSeek API (deepseek-chat, deepseek-reasoner)131By default, this uses the `DEEPSEEK_API_KEY` environment variable.132 133#### OpenRouter API (Llama3.1)134 135By default, this uses the `OPENROUTER_API_KEY` environment variable.136 137#### Google Gemini138We support Google Gemini models (e.g., "gemini-1.5-flash", "gemini-1.5-pro") via the [google-generativeai](https://pypi.org/project/google-generativeai) Python library. By default, it uses the environment variable:139 140```bash141export GEMINI_API_KEY="YOUR GEMINI API KEY"142```143 144#### Semantic Scholar API (Literature Search)145 146Our code can also optionally use a Semantic Scholar API Key (`S2_API_KEY`) for higher throughput [if you have one](https://www.semanticscholar.org/product/api), though it should work without it in principle. If you have problems with Semantic Scholar, you can skip the literature search and citation phases of paper generation.147 148Be sure to provide the key for the model used for your runs, e.g.:149 150```bash151export OPENAI_API_KEY="YOUR KEY HERE"152export S2_API_KEY="YOUR KEY HERE"153```154 155#### OpenAlex API (Literature Search Alternative)156 157OpenAlex API can be used as an alternative if you do not have a Semantic Scholar API Key.158OpenAlex does not require API key.159 160```bash161pip install pyalex162export OPENALEX_MAIL_ADDRESS="YOUR EMAIL ADDRESS"163```164 165And specify `--engine openalex` when you execute the AI Scientist code.166 167Note that this is experimental for those who do not have a Semantic Scholar API Key.168 169## Setting Up the Templates170 171This section provides instructions for setting up each of the three templates used in our paper. Before running The AI Scientist experiments, please ensure you have completed the setup steps for the templates you are interested in.172 173### NanoGPT Template174 175**Description:** This template investigates transformer-based autoregressive next-token prediction tasks.176 177**Setup Steps:**178 1791. **Prepare the data:**180 181   ```bash182   python data/enwik8/prepare.py183   python data/shakespeare_char/prepare.py184   python data/text8/prepare.py185   ```186 1872. **Create baseline runs (machine dependent):**188 189   ```bash190   # Set up NanoGPT baseline run191   # NOTE: YOU MUST FIRST RUN THE PREPARE SCRIPTS ABOVE!192   cd templates/nanoGPT193   python experiment.py --out_dir run_0194   python plot.py195   ```196 197### 2D Diffusion Template198 199**Description:** This template studies improving the performance of diffusion generative models on low-dimensional datasets.200 201**Setup Steps:**202 2031. **Install dependencies:**204 205   ```bash206   # Set up 2D Diffusion207   git clone https://github.com/gregversteeg/NPEET.git208   cd NPEET209   pip install .210   pip install scikit-learn211   ```212 2132. **Create baseline runs:**214 215   ```bash216   # Set up 2D Diffusion baseline run217   cd templates/2d_diffusion218   python experiment.py --out_dir run_0219   python plot.py220   ```221 222### Grokking Template223 224**Description:** This template investigates questions about generalization and learning speed in deep neural networks.225 226**Setup Steps:**227 2281. **Install dependencies:**229 230   ```bash231   # Set up Grokking232   pip install einops233   ```234 2352. **Create baseline runs:**236 237   ```bash238   # Set up Grokking baseline run239   cd templates/grokking240   python experiment.py --out_dir run_0241   python plot.py242   ```243 244## Run AI Scientist Paper Generation Experiments245 246**Note:** Please ensure the setup steps above are completed before running these experiments.247 248```bash249conda activate ai_scientist250# Run the paper generation.251python launch_scientist.py --model "gpt-4o-2024-05-13" --experiment nanoGPT_lite --num-ideas 2252python launch_scientist.py --model "claude-3-5-sonnet-20241022" --experiment nanoGPT_lite --num-ideas 2253```254 255If you have more than one GPU, use the `--parallel` option to parallelize ideas across multiple GPUs.256 257## Getting an LLM-Generated Paper Review258 259```python260import openai261from ai_scientist.perform_review import load_paper, perform_review262 263client = openai.OpenAI()264model = "gpt-4o-2024-05-13"265 266# Load paper from PDF file (raw text)267paper_txt = load_paper("report.pdf")268 269# Get the review dictionary270review = perform_review(271    paper_txt,272    model,273    client,274    num_reflections=5,275    num_fs_examples=1,276    num_reviews_ensemble=5,277    temperature=0.1,278)279 280# Inspect review results281review["Overall"]    # Overall score (1-10)282review["Decision"]   # 'Accept' or 'Reject'283review["Weaknesses"] # List of weaknesses (strings)284```285 286To run batch analysis:287 288```bash289cd review_iclr_bench290python iclr_analysis.py --num_reviews 500 --batch_size 100 --num_fs_examples 1 --num_reflections 5 --temperature 0.1 --num_reviews_ensemble 5291```292 293## Making Your Own Template294 295If there is an area of study you would like **The AI Scientist** to explore, it is straightforward to create your own templates. In general, follow the structure of the existing templates, which consist of:296 297- `experiment.py` — This is the main script where the core content is. It takes an argument `--out_dir`, which specifies where it should create the folder and save the relevant information from the run.298- `plot.py` — This script takes the information from the `run` folders and creates plots. The code should be clear and easy to edit.299- `prompt.json` — Put information about your template here.300- `seed_ideas.json` — Place example ideas here. You can also try to generate ideas without any examples and then pick the best one or two to put here.301- `latex/template.tex` — We recommend using our LaTeX folder but be sure to replace the pre-loaded citations with ones that you expect to be more relevant.302 303The key to making new templates work is matching the base filenames and output JSONs to the existing format; everything else is free to change.304You should also ensure that the `template.tex` file is updated to use the correct citation style / base plots for your template.305 306### Community-Contributed Templates307 308We welcome community contributions in the form of new templates. While these are not maintained by us, we are delighted to highlight your templates to others. Below, we list community-contributed templates along with links to their pull requests (PRs):309 310- Infectious Disease Modeling (`seir`) - [PR #137](https://github.com/SakanaAI/AI-Scientist/pull/137)311- Image Classification with MobileNetV3 (`mobilenetV3`) - [PR #141](https://github.com/SakanaAI/AI-Scientist/pull/141)312- Sketch RNN (`sketch_rnn`) - [PR #143](https://github.com/SakanaAI/AI-Scientist/pull/143)313- AI in Quantum Chemistry (`MACE`) - [PR#157](https://github.com/SakanaAI/AI-Scientist/pull/157)314- Earthquake Prediction (`earthquake-prediction`) - [PR #167](https://github.com/SakanaAI/AI-Scientist/pull/167)315- Tensorial Radiance Fields (`tensorf`) - [PR #175](https://github.com/SakanaAI/AI-Scientist/pull/175)316- Large Language Model Steering / Probes (`probes`) - [PR #215](https://github.com/SakanaAI/AI-Scientist/pull/215)317 318*This section is reserved for community contributions. Please submit a pull request to add your template to the list! Please describe the template in the PR description, and also show examples of the generated papers.*319 320## Template Resources321 322We provide three templates, which heavily use code from other repositories, credited below:323 324- **NanoGPT Template** uses code from [NanoGPT](https://github.com/karpathy/nanoGPT) and this [PR](https://github.com/karpathy/nanoGPT/pull/254).325- **2D Diffusion Template** uses code from [tiny-diffusion](https://github.com/tanelp/tiny-diffusion), [ema-pytorch](https://github.com/lucidrains/ema-pytorch), and [Datasaur](https://www.research.autodesk.com/publications/same-stats-different-graphs/).326- **Grokking Template** uses code from [Sea-Snell/grokking](https://github.com/Sea-Snell/grokking) and [danielmamay/grokking](https://github.com/danielmamay/grokking).327 328We would like to thank the developers of the open-source models and packages for their contributions and for making their work available.329 330## Citing The AI Scientist331 332If you use **The AI Scientist** in your research, please cite it as follows:333 334```335@article{lu2024aiscientist,336  title={The {AI} {S}cientist: Towards Fully Automated Open-Ended Scientific Discovery},337  author={Lu, Chris and Lu, Cong and Lange, Robert Tjarko and Foerster, Jakob and Clune, Jeff and Ha, David},338  journal={arXiv preprint arXiv:2408.06292},339  year={2024}340}341```342 343## Frequently Asked Questions344 345We recommend reading our paper first for any questions you have on The AI Scientist.346 347**Why am I missing files when running The AI Scientist?**348 349Ensure you have completed all the setup and preparation steps before the main experiment script.350 351**Why has a PDF or a review not been generated?**352 353The AI Scientist finishes an idea with a success rate that depends on the template, the base foundation model, and the complexity of the idea. We advise referring to our main paper. The highest success rates are observed with Claude Sonnet 3.5. Reviews are best done with GPT-4o; all other models have issues with positivity bias or failure to conform to required outputs.354 355**What is the cost of each idea generated?**356 357Typically less than $15 per paper with Claude Sonnet 3.5. We recommend DeepSeek Coder V2 for a much more cost-effective approach. A good place to look for new models is the [Aider leaderboard](https://aider.chat/docs/leaderboards/).358 359**How do I change the base conference format associated with the write-ups?**360 361Change the base `template.tex` files contained within each template.362 363**How do I run The AI Scientist for different subject fields?**364 365Please refer to the instructions for different templates. In this current iteration, this is restricted to ideas that can be expressed in code. However, lifting this restriction would represent exciting future work! :)366 367**How do I add support for a new foundation model?**368 369You may modify `ai_scientist/llm.py` to add support for a new foundation model. We do not advise using any model that is significantly weaker than GPT-4 level for **The AI Scientist**.370 371**Why do I need to run the baseline runs myself?**372 373These appear as `run_0` and should be run per machine you execute **The AI Scientist** on for accurate run-time comparisons due to hardware differences.374 375**What if I have problems accessing the Semantic Scholar API?**376 377We use the Semantic Scholar API to check ideas for novelty and collect citations for the paper write-up. You may be able to skip these phases if you don't have an API key or the API is slow to access.378 379## Containerization380 381We include a [community-contributed](https://github.com/SakanaAI/AI-Scientist/pull/21) Docker image that may assist with your containerization efforts in `experimental/Dockerfile`.382 383You can use this image like this:384 385```bash386# Endpoint Script387docker run -e OPENAI_API_KEY=$OPENAI_API_KEY -v `pwd`/templates:/app/AI-Scientist/templates <AI_SCIENTIST_IMAGE> \388  --model gpt-4o-2024-05-13 \389  --experiment 2d_diffusion \390  --num-ideas 2391```392 393```bash394# Interactive395docker run -it -e OPENAI_API_KEY=$OPENAI_API_KEY \396  --entrypoint /bin/bash \397  <AI_SCIENTIST_IMAGE>398```399 400## ⚖️ License & Responsible Use401 402This project is licensed under **The AI Scientist Source Code License** (a derivative of the Responsible AI License). 403 404**Mandatory Disclosure:** By using this code, you are legally bound to clearly and prominently disclose the use of AI in any resulting scientific manuscripts or papers. 405 406We recommend the following attribution in your paper's Abstract or Methods section:407> "This manuscript was autonomously generated using [The AI Scientist](https://github.com/SakanaAI/AI-Scientist)."408 409## Star History410 411[![Star History Chart](https://api.star-history.com/svg?repos=SakanaAI/AI-Scientist&type=Date)](https://star-history.com/#SakanaAI/AI-Scientist&Date)412