AnnaMats/ppo-Pyramids-Training
0110
1# Unity ML-Agents PettingZoo Wrapper2 3With the increasing interest in multi-agent training with a gym-like API, we provide a4PettingZoo Wrapper around the [Petting Zoo API](https://www.pettingzoo.ml/). Our wrapper5provides interfaces on top of our `UnityEnvironment` class, which is the default way of6interfacing with a Unity environment via Python.7 8## Installation and Examples9 10The PettingZoo wrapper is part of the `mlgents_envs` package. Please refer to the11[mlagents_envs installation instructions](ML-Agents-Envs-README.md).12 13[[Colab] PettingZoo Wrapper Example](https://colab.research.google.com/github/Unity-Technologies/ml-agents/blob/develop-python-api-ga/ml-agents-envs/colabs/Colab_PettingZoo.ipynb)14 15This colab notebook demonstrates the example usage of the wrapper, including installation,16basic usages, and an example with our17[Striker vs Goalie environment](https://github.com/Unity-Technologies/ml-agents/blob/main/docs/Learning-Environment-Examples.md#strikers-vs-goalie)18which is a multi-agents environment with multiple different behavior names.19 20## API interface21 22This wrapper is compatible with PettingZoo API. Please check out23[PettingZoo API page](https://www.pettingzoo.ml/api) for more details.24Here's an example of interacting with wrapped environment:25 26```python27from mlagents_envs.environment import UnityEnvironment28from mlagents_envs.envs import UnityToPettingZooWrapper29 30unity_env = UnityEnvironment("StrikersVsGoalie")31env = UnityToPettingZooWrapper(unity_env)32env.reset()33for agent in env.agent_iter():34 observation, reward, done, info = env.last()35 action = policy(observation, agent)36 env.step(action)37```38 39## Notes40- There is support for both [AEC](https://www.pettingzoo.ml/api#interacting-with-environments)41 and [Parallel](https://www.pettingzoo.ml/api#parallel-api) PettingZoo APIs.42- The AEC wrapper is compatible with PettingZoo (PZ) API interface but works in a slightly43 different way under the hood. For the AEC API, Instead of stepping the environment in every `env.step(action)`,44 the PZ wrapper will store the action, and will only perform environment stepping when all the45 agents requesting for actions in the current step have been assigned an action. This is for46 performance, considering that the communication between Unity and python is more efficient47 when data are sent in batches.48- Since the actions for the AEC wrapper are stored without applying them to the environment until49 all the actions are queued, some components of the API might behave in unexpected way. For example, a call50 to `env.reward` should return the instantaneous reward for that particular step, but the true51 reward would only be available when an actual environment step is performed. It's recommended that52 you follow the API definition for training (access rewards from `env.last()` instead of53 `env.reward`) and the underlying mechanism shouldn't affect training results.54- The environments will automatically reset when it's done, so `env.agent_iter(max_step)` will55 keep going on until the specified max step is reached (default: `2**63`). There is no need to56 call `env.reset()` except for the very beginning of instantiating an environment.57 58 