Team Ai
Modelpublic

AnnaMats/ppo-Pyramids-Training

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes110downloads
Getting-Started.md266 linesDownload Raw Back to docs
1# Getting Started Guide2 3This guide walks through the end-to-end process of opening one of our4[example environments](Learning-Environment-Examples.md) in Unity, training an5Agent in it, and embedding the trained model into the Unity environment. After6reading this tutorial, you should be able to train any of the example7environments. If you are not familiar with the8[Unity Engine](https://unity3d.com/unity), view our9[Background: Unity](Background-Unity.md) page for helpful pointers.10Additionally, if you're not familiar with machine learning, view our11[Background: Machine Learning](Background-Machine-Learning.md) page for a brief12overview and helpful pointers.13 14![3D Balance Ball](images/balance.png)15 16For this guide, we'll use the **3D Balance Ball** environment which contains a17number of agent cubes and balls (which are all copies of each other). Each agent18cube tries to keep its ball from falling by rotating either horizontally or19vertically. In this environment, an agent cube is an **Agent** that receives a20reward for every step that it balances the ball. An agent is also penalized with21a negative reward for dropping the ball. The goal of the training process is to22have the agents learn to balance the ball on their head.23 24Let's get started!25 26## Installation27 28If you haven't already, follow the [installation instructions](Installation.md).29Afterwards, open the Unity Project that contains all the example environments:30 311. Open the Package Manager Window by navigating to `Window -> Package Manager`32   in the menu.331. Navigate to the ML-Agents Package and click on it.341. Find the `3D Ball` sample and click `Import`.351. In the **Project** window, go to the36   `Assets/ML-Agents/Examples/3DBall/Scenes` folder and open the `3DBall` scene37   file.38 39## Understanding a Unity Environment40 41An agent is an autonomous actor that observes and interacts with an42_environment_. In the context of Unity, an environment is a scene containing one43or more Agent objects, and, of course, the other entities that an agent44interacts with.45 46![Unity Editor](images/mlagents-3DBallHierarchy.png)47 48**Note:** In Unity, the base object of everything in a scene is the49_GameObject_. The GameObject is essentially a container for everything else,50including behaviors, graphics, physics, etc. To see the components that make up51a GameObject, select the GameObject in the Scene window, and open the Inspector52window. The Inspector shows every component on a GameObject.53 54The first thing you may notice after opening the 3D Balance Ball scene is that55it contains not one, but several agent cubes. Each agent cube in the scene is an56independent agent, but they all share the same Behavior. 3D Balance Ball does57this to speed up training since all twelve agents contribute to training in58parallel.59 60### Agent61 62The Agent is the actor that observes and takes actions in the environment. In63the 3D Balance Ball environment, the Agent components are placed on the twelve64"Agent" GameObjects. The base Agent object has a few properties that affect its65behavior:66 67- **Behavior Parameters** — Every Agent must have a Behavior. The Behavior68  determines how an Agent makes decisions.69- **Max Step** — Defines how many simulation steps can occur before the Agent's70  episode ends. In 3D Balance Ball, an Agent restarts after 5000 steps.71 72#### Behavior Parameters : Vector Observation Space73 74Before making a decision, an agent collects its observation about its state in75the world. The vector observation is a vector of floating point numbers which76contain relevant information for the agent to make decisions.77 78The Behavior Parameters of the 3D Balance Ball example uses a `Space Size` of 8.79This means that the feature vector containing the Agent's observations contains80eight elements: the `x` and `z` components of the agent cube's rotation and the81`x`, `y`, and `z` components of the ball's relative position and velocity.82 83#### Behavior Parameters : Actions84 85An Agent is given instructions in the form of actions.86ML-Agents Toolkit classifies actions into two types: continuous and discrete.87The 3D Balance Ball example is programmed to use continuous actions, which88are a vector of floating-point numbers that can vary continuously. More specifically,89it uses a `Space Size` of 2 to control the amount of `x` and `z` rotations to apply to90itself to keep the ball balanced on its head.91 92## Running a pre-trained model93 94We include pre-trained models for our agents (`.onnx` files) and we use the95[Unity Inference Engine](Unity-Inference-Engine.md) to run these models inside96Unity. In this section, we will use the pre-trained model for the 3D Ball97example.98 991. In the **Project** window, go to the100   `Assets/ML-Agents/Examples/3DBall/Prefabs` folder. Expand `3DBall` and click101   on the `Agent` prefab. You should see the `Agent` prefab in the **Inspector**102   window.103 104   **Note**: The platforms in the `3DBall` scene were created using the `3DBall`105   prefab. Instead of updating all 12 platforms individually, you can update the106   `3DBall` prefab instead.107 108   ![Platform Prefab](images/platform_prefab.png)109 1101. In the **Project** window, drag the **3DBall** Model located in111   `Assets/ML-Agents/Examples/3DBall/TFModels` into the `Model` property under112   `Behavior Parameters (Script)` component in the Agent GameObject113   **Inspector** window.114 115   ![3dball learning brain](images/3dball_learning_brain.png)116 1171. You should notice that each `Agent` under each `3DBall` in the **Hierarchy**118   windows now contains **3DBall** as `Model` on the `Behavior Parameters`.119   **Note** : You can modify multiple game objects in a scene by selecting them120   all at once using the search bar in the Scene Hierarchy.1211. Set the **Inference Device** to use for this model as `CPU`.1221. Click the **Play** button in the Unity Editor and you will see the platforms123   balance the balls using the pre-trained model.124 125## Training a new model with Reinforcement Learning126 127While we provide pre-trained models for the agents in this environment, any128environment you make yourself will require training agents from scratch to129generate a new model file. In this section we will demonstrate how to use the130reinforcement learning algorithms that are part of the ML-Agents Python package131to accomplish this. We have provided a convenient command `mlagents-learn` which132accepts arguments used to configure both training and inference phases.133 134### Training the environment135 1361. Open a command or terminal window.1371. Navigate to the folder where you cloned the `ml-agents` repository. **Note**:138   If you followed the default [installation](Installation.md), then you should139   be able to run `mlagents-learn` from any directory.1401. Run `mlagents-learn config/ppo/3DBall.yaml --run-id=first3DBallRun`.141   - `config/ppo/3DBall.yaml` is the path to a default training142     configuration file that we provide. The `config/ppo` folder includes training configuration143     files for all our example environments, including 3DBall.144   - `run-id` is a unique name for this training session.1451. When the message _"Start training by pressing the Play button in the Unity146   Editor"_ is displayed on the screen, you can press the **Play** button in147   Unity to start training in the Editor.148 149If `mlagents-learn` runs correctly and starts training, you should see something150like this:151 152```console153INFO:mlagents_envs:154'Ball3DAcademy' started successfully!155Unity Academy name: Ball3DAcademy156 157INFO:mlagents_envs:Connected new brain:158Unity brain name: 3DBallLearning159        Number of Visual Observations (per agent): 0160        Vector Observation space size (per agent): 8161        Number of stacked Vector Observation: 1162INFO:mlagents_envs:Hyperparameters for the PPO Trainer of brain 3DBallLearning:163        batch_size:          64164        beta:                0.001165        buffer_size:         12000166        epsilon:             0.2167        gamma:               0.995168        hidden_units:        128169        lambd:               0.99170        learning_rate:       0.0003171        max_steps:           5.0e4172        normalize:           True173        num_epoch:           3174        num_layers:          2175        time_horizon:        1000176        sequence_length:     64177        summary_freq:        1000178        use_recurrent:       False179        memory_size:         256180        use_curiosity:       False181        curiosity_strength:  0.01182        curiosity_enc_size:  128183        output_path: ./results/first3DBallRun/3DBallLearning184INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 1000. Mean Reward: 1.242. Std of Reward: 0.746. Training.185INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 2000. Mean Reward: 1.319. Std of Reward: 0.693. Training.186INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 3000. Mean Reward: 1.804. Std of Reward: 1.056. Training.187INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 4000. Mean Reward: 2.151. Std of Reward: 1.432. Training.188INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 5000. Mean Reward: 3.175. Std of Reward: 2.250. Training.189INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 6000. Mean Reward: 4.898. Std of Reward: 4.019. Training.190INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 7000. Mean Reward: 6.716. Std of Reward: 5.125. Training.191INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 8000. Mean Reward: 12.124. Std of Reward: 11.929. Training.192INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 9000. Mean Reward: 18.151. Std of Reward: 16.871. Training.193INFO:mlagents.trainers: first3DBallRun: 3DBallLearning: Step: 10000. Mean Reward: 27.284. Std of Reward: 28.667. Training.194```195 196Note how the `Mean Reward` value printed to the screen increases as training197progresses. This is a positive sign that training is succeeding.198 199**Note**: You can train using an executable rather than the Editor. To do so,200follow the instructions in201[Using an Executable](Learning-Environment-Executable.md).202 203### Observing Training Progress204 205Once you start training using `mlagents-learn` in the way described in the206previous section, the `ml-agents` directory will contain a `results`207directory. In order to observe the training process in more detail, you can use208TensorBoard. From the command line run:209 210```sh211tensorboard --logdir results212```213 214Then navigate to `localhost:6006` in your browser to view the TensorBoard215summary statistics as shown below. For the purposes of this section, the most216important statistic is `Environment/Cumulative Reward` which should increase217throughout training, eventually converging close to `100` which is the maximum218reward the agent can accumulate.219 220![Example TensorBoard Run](images/mlagents-TensorBoard.png)221 222## Embedding the model into the Unity Environment223 224Once the training process completes, and the training process saves the model225(denoted by the `Saved Model` message) you can add it to the Unity project and226use it with compatible Agents (the Agents that generated the model). **Note:**227Do not just close the Unity Window once the `Saved Model` message appears.228Either wait for the training process to close the window or press `Ctrl+C` at229the command-line prompt. If you close the window manually, the `.onnx` file230containing the trained model is not exported into the ml-agents folder.231 232If you've quit the training early using `Ctrl+C` and want to resume training,233run the same command again, appending the `--resume` flag:234 235```sh236mlagents-learn config/ppo/3DBall.yaml --run-id=first3DBallRun --resume237```238 239Your trained model will be at `results/<run-identifier>/<behavior_name>.onnx` where240`<behavior_name>` is the name of the `Behavior Name` of the agents corresponding241to the model. This file corresponds to your model's latest checkpoint. You can242now embed this trained model into your Agents by following the steps below,243which is similar to the steps described [above](#running-a-pre-trained-model).244 2451. Move your model file into246   `Project/Assets/ML-Agents/Examples/3DBall/TFModels/`.2471. Open the Unity Editor, and select the **3DBall** scene as described above.2481. Select the **3DBall** prefab Agent object.2491. Drag the `<behavior_name>.onnx` file from the Project window of the Editor to250   the **Model** placeholder in the **Ball3DAgent** inspector window.2511. Press the **Play** button at the top of the Editor.252 253## Next Steps254 255- For more information on the ML-Agents Toolkit, in addition to helpful256  background, check out the [ML-Agents Toolkit Overview](ML-Agents-Overview.md)257  page.258- For a "Hello World" introduction to creating your own Learning Environment,259  check out the260  [Making a New Learning Environment](Learning-Environment-Create-New.md) page.261- For an overview on the more complex example environments that are provided in262  this toolkit, check out the263  [Example Environments](Learning-Environment-Examples.md) page.264- For more information on the various training options available, check out the265  [Training ML-Agents](Training-ML-Agents.md) page.266