Team Ai
Modelpublic

AnnaMats/ppo-Pyramids-Training

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes110downloads
Using-Tensorboard.md137 linesDownload Raw Back to docs
1# Using TensorBoard to Observe Training2 3The ML-Agents Toolkit saves statistics during learning session that you can view4with a TensorFlow utility named,5[TensorBoard](https://www.tensorflow.org/programmers_guide/summaries_and_tensorboard).6 7The `mlagents-learn` command saves training statistics to a folder named8`results`, organized by the `run-id` value you assign to a training session.9 10In order to observe the training process, either during training or afterward,11start TensorBoard:12 131. Open a terminal or console window:141. Navigate to the directory where the ML-Agents Toolkit is installed.151. From the command line run: `tensorboard --logdir results --port 6006`161. Open a browser window and navigate to17   [localhost:6006](http://localhost:6006).18 19**Note:** The default port TensorBoard uses is 6006. If there is an existing20session running on port 6006 a new session can be launched on an open port using21the --port option.22 23**Note:** If you don't assign a `run-id` identifier, `mlagents-learn` uses the24default string, "ppo". You can delete the folders under the `results` directory25to clear out old statistics.26 27On the left side of the TensorBoard window, you can select which of the training28runs you want to display. You can select multiple run-ids to compare statistics.29The TensorBoard window also provides options for how to display and smooth30graphs.31 32## The ML-Agents Toolkit training statistics33 34The ML-Agents training program saves the following statistics:35 36![Example TensorBoard Run](images/mlagents-TensorBoard.png)37 38### Environment Statistics39 40- `Environment/Lesson` - Plots the progress from lesson to lesson. Only41  interesting when performing curriculum training.42 43- `Environment/Cumulative Reward` - The mean cumulative episode reward over all44  agents. Should increase during a successful training session.45 46- `Environment/Episode Length` - The mean length of each episode in the47  environment for all agents.48 49### Is Training50 51- `Is Training` - A boolean indicating if the agent is updating its model.52 53### Policy Statistics54 55- `Policy/Entropy` (PPO; SAC) - How random the decisions of the model are.56  Should slowly decrease during a successful training process. If it decreases57  too quickly, the `beta` hyperparameter should be increased.58 59- `Policy/Learning Rate` (PPO; SAC) - How large a step the training algorithm60  takes as it searches for the optimal policy. Should decrease over time.61 62- `Policy/Entropy Coefficient` (SAC) - Determines the relative importance of the63  entropy term. This value is adjusted automatically so that the agent retains64  some amount of randomness during training.65 66- `Policy/Extrinsic Reward` (PPO; SAC) - This corresponds to the mean cumulative67  reward received from the environment per-episode.68 69- `Policy/Value Estimate` (PPO; SAC) - The mean value estimate for all states70  visited by the agent. Should increase during a successful training session.71 72- `Policy/Curiosity Reward` (PPO/SAC+Curiosity) - This corresponds to the mean73  cumulative intrinsic reward generated per-episode.74 75- `Policy/Curiosity Value Estimate` (PPO/SAC+Curiosity) - The agent's value76  estimate for the curiosity reward.77 78- `Policy/GAIL Reward` (PPO/SAC+GAIL) - This corresponds to the mean cumulative79  discriminator-based reward generated per-episode.80 81- `Policy/GAIL Value Estimate` (PPO/SAC+GAIL) - The agent's value estimate for82  the GAIL reward.83 84- `Policy/GAIL Policy Estimate` (PPO/SAC+GAIL) - The discriminator's estimate85  for states and actions generated by the policy.86 87- `Policy/GAIL Expert Estimate` (PPO/SAC+GAIL) - The discriminator's estimate88  for states and actions drawn from expert demonstrations.89 90### Learning Loss Functions91 92- `Losses/Policy Loss` (PPO; SAC) - The mean magnitude of policy loss function.93  Correlates to how much the policy (process for deciding actions) is changing.94  The magnitude of this should decrease during a successful training session.95 96- `Losses/Value Loss` (PPO; SAC) - The mean loss of the value function update.97  Correlates to how well the model is able to predict the value of each state.98  This should increase while the agent is learning, and then decrease once the99  reward stabilizes.100 101- `Losses/Forward Loss` (PPO/SAC+Curiosity) - The mean magnitude of the forward102  model loss function. Corresponds to how well the model is able to predict the103  new observation encoding.104 105- `Losses/Inverse Loss` (PPO/SAC+Curiosity) - The mean magnitude of the inverse106  model loss function. Corresponds to how well the model is able to predict the107  action taken between two observations.108 109- `Losses/Pretraining Loss` (BC) - The mean magnitude of the behavioral cloning110  loss. Corresponds to how well the model imitates the demonstration data.111 112- `Losses/GAIL Loss` (GAIL) - The mean magnitude of the GAIL discriminator loss.113  Corresponds to how well the model imitates the demonstration data.114 115### Self-Play116 117- `Self-Play/ELO` (Self-Play) -118  [ELO](https://en.wikipedia.org/wiki/Elo_rating_system) measures the relative119  skill level between two players. In a proper training run, the ELO of the120  agent should steadily increase.121 122## Exporting Data from TensorBoard123To export timeseries data in CSV or JSON format, check the "Show data download124links" in the upper left. This will enable download links below each chart.125 126![Example TensorBoard Run](images/TensorBoard-download.png)127 128## Custom Metrics from Unity129 130To get custom metrics from a C# environment into TensorBoard, you can use the131`StatsRecorder`:132 133```csharp134var statsRecorder = Academy.Instance.StatsRecorder;135statsRecorder.Add("MyMetric", 1.0);136```137