Team Ai
Modelpublic

AnnaMats/ppo-Pyramids-Training

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes110downloads
Python-LLAPI.md355 linesDownload Raw Back to docs
1# Unity ML-Agents Python Low Level API2 3The `mlagents` Python package contains two components: a low level API which4allows you to interact directly with a Unity Environment (`mlagents_envs`) and5an entry point to train (`mlagents-learn`) which allows you to train agents in6Unity Environments using our implementations of reinforcement learning or7imitation learning. This document describes how to use the `mlagents_envs` API.8For information on using `mlagents-learn`, see [here](Training-ML-Agents.md).9For Python Low Level API documentation, see [here](Python-LLAPI-Documentation.md).10 11The Python Low Level API can be used to interact directly with your Unity12learning environment. As such, it can serve as the basis for developing and13evaluating new learning algorithms.14 15## mlagents_envs16 17The ML-Agents Toolkit Low Level API is a Python API for controlling the18simulation loop of an environment or game built with Unity. This API is used by19the training algorithms inside the ML-Agent Toolkit, but you can also write your20own Python programs using this API.21 22The key objects in the Python API include:23 24- **UnityEnvironment** — the main interface between the Unity application and25  your code. Use UnityEnvironment to start and control a simulation or training26  session.27- **BehaviorName** - is a string that identifies a behavior in the simulation.28- **AgentId** - is an `int` that serves as unique identifier for Agents in the29  simulation.30- **DecisionSteps** — contains the data from Agents belonging to the same31  "Behavior" in the simulation, such as observations and rewards. Only Agents32  that requested a decision since the last call to `env.step()` are in the33  DecisionSteps object.34- **TerminalSteps** — contains the data from Agents belonging to the same35  "Behavior" in the simulation, such as observations and rewards. Only Agents36  whose episode ended since the last call to `env.step()` are in the37  TerminalSteps object.38- **BehaviorSpec** — describes the shape of the observation data inside39  DecisionSteps and TerminalSteps as well as the expected action shapes.40 41These classes are all defined in the42[base_env](../ml-agents-envs/mlagents_envs/base_env.py) script.43 44An Agent "Behavior" is a group of Agents identified by a `BehaviorName` that45share the same observations and action types (described in their46`BehaviorSpec`). You can think about Agent Behavior as a group of agents that47will share the same policy. All Agents with the same behavior have the same goal48and reward signals.49 50To communicate with an Agent in a Unity environment from a Python program, the51Agent in the simulation must have `Behavior Parameters` set to communicate. You52must set the `Behavior Type` to `Default` and give it a `Behavior Name`.53 54_Notice: Currently communication between Unity and Python takes place over an55open socket without authentication. As such, please make sure that the network56where training takes place is secure. This will be addressed in a future57release._58 59## Loading a Unity Environment60 61Python-side communication happens through `UnityEnvironment` which is located in62[`environment.py`](../ml-agents-envs/mlagents_envs/environment.py). To load a63Unity environment from a built binary file, put the file in the same directory64as `envs`. For example, if the filename of your Unity environment is `3DBall`,65in python, run:66 67```python68from mlagents_envs.environment import UnityEnvironment69# This is a non-blocking call that only loads the environment.70env = UnityEnvironment(file_name="3DBall", seed=1, side_channels=[])71# Start interacting with the environment.72env.reset()73behavior_names = env.behavior_specs.keys()74...75```76**NOTE:** Please read [Interacting with a Unity Environment](#interacting-with-a-unity-environment)77to read more about how you can interact with the Unity environment from Python.78 79- `file_name` is the name of the environment binary (located in the root80  directory of the python project).81- `worker_id` indicates which port to use for communication with the82  environment. For use in parallel training regimes such as A3C.83- `seed` indicates the seed to use when generating random numbers during the84  training process. In environments which are stochastic, setting the seed85  enables reproducible experimentation by ensuring that the environment and86  trainers utilize the same random seed.87- `side_channels` provides a way to exchange data with the Unity simulation that88  is not related to the reinforcement learning loop. For example: configurations89  or properties. More on them in the [Side Channels](Custom-SideChannels.md) doc.90 91If you want to directly interact with the Editor, you need to use92`file_name=None`, then press the **Play** button in the Editor when the message93_"Start training by pressing the Play button in the Unity Editor"_ is displayed94on the screen95 96### Interacting with a Unity Environment97 98#### The BaseEnv interface99 100A `BaseEnv` has the following methods:101 102- **Reset : `env.reset()`** Sends a signal to reset the environment. Returns103  None.104- **Step : `env.step()`** Sends a signal to step the environment. Returns None.105  Note that a "step" for Python does not correspond to either Unity `Update` nor106  `FixedUpdate`. When `step()` or `reset()` is called, the Unity simulation will107  move forward until an Agent in the simulation needs a input from Python to108  act.109- **Close : `env.close()`** Sends a shutdown signal to the environment and110  terminates the communication.111- **Behavior Specs : `env.behavior_specs`** Returns a Mapping of112  `BehaviorName` to `BehaviorSpec` objects (read only).113  A `BehaviorSpec` contains the observation shapes and the114  `ActionSpec` (which defines the action shape). Note that115  the `BehaviorSpec` for a specific group is fixed throughout the simulation.116  The number of entries in the Mapping can change over time in the simulation117  if new Agent behaviors are created in the simulation.118- **Get Steps : `env.get_steps(behavior_name: str)`** Returns a tuple119  `DecisionSteps, TerminalSteps` corresponding to the behavior_name given as120  input. The `DecisionSteps` contains information about the state of the agents121  **that need an action this step** and have the behavior behavior_name. The122  `TerminalSteps` contains information about the state of the agents **whose123  episode ended** and have the behavior behavior_name. Both `DecisionSteps` and124  `TerminalSteps` contain information such as the observations, the rewards and125  the agent identifiers. `DecisionSteps` also contains action masks for the next126  action while `TerminalSteps` contains the reason for termination (did the127  Agent reach its maximum step and was interrupted). The data is in `np.array`128  of which the first dimension is always the number of agents note that the129  number of agents is not guaranteed to remain constant during the simulation130  and it is not unusual to have either `DecisionSteps` or `TerminalSteps`131  contain no Agents at all.132- **Set Actions :`env.set_actions(behavior_name: str, action: ActionTuple)`** Sets133  the actions for a whole agent group. `action` is an `ActionTuple`, which134  is made up of a 2D `np.array` of `dtype=np.int32` for discrete actions, and135  `dtype=np.float32` for continuous actions. The first dimension of `np.array`136  in the tuple is the number of agents that requested a decision since the137  last call to `env.step()`. The second dimension is the number of discrete or138  continuous actions for the corresponding array.139- **Set Action for Agent :140  `env.set_action_for_agent(agent_group: str, agent_id: int, action: ActionTuple)`**141  Sets the action for a specific Agent in an agent group. `agent_group` is the142  name of the group the Agent belongs to and `agent_id` is the integer143  identifier of the Agent. `action` is an `ActionTuple` as described above.144**Note:** If no action is provided for an agent group between two calls to145`env.step()` then the default action will be all zeros.146 147#### DecisionSteps and DecisionStep148 149`DecisionSteps` (with `s`) contains information about a whole batch of Agents150while `DecisionStep` (no `s`) only contains information about a single Agent.151 152A `DecisionSteps` has the following fields :153 154- `obs` is a list of numpy arrays observations collected by the group of agent.155  The first dimension of the array corresponds to the batch size of the group156  (number of agents requesting a decision since the last call to `env.step()`).157- `reward` is a float vector of length batch size. Corresponds to the rewards158  collected by each agent since the last simulation step.159- `agent_id` is an int vector of length batch size containing unique identifier160  for the corresponding Agent. This is used to track Agents across simulation161  steps.162- `action_mask` is an optional list of two dimensional arrays of booleans which is only163  available when using multi-discrete actions. Each array corresponds to an164  action branch. The first dimension of each array is the batch size and the165  second contains a mask for each action of the branch. If true, the action is166  not available for the agent during this simulation step.167 168It also has the two following methods:169 170- `len(DecisionSteps)` Returns the number of agents requesting a decision since171  the last call to `env.step()`.172- `DecisionSteps[agent_id]` Returns a `DecisionStep` for the Agent with the173  `agent_id` unique identifier.174 175A `DecisionStep` has the following fields:176 177- `obs` is a list of numpy arrays observations collected by the agent. (Each178  array has one less dimension than the arrays in `DecisionSteps`)179- `reward` is a float. Corresponds to the rewards collected by the agent since180  the last simulation step.181- `agent_id` is an int and an unique identifier for the corresponding Agent.182- `action_mask` is an optional list of one dimensional arrays of booleans which is only183  available when using multi-discrete actions. Each array corresponds to an184  action branch. Each array contains a mask for each action of the branch. If185  true, the action is not available for the agent during this simulation step.186 187#### TerminalSteps and TerminalStep188 189Similarly to `DecisionSteps` and `DecisionStep`, `TerminalSteps` (with `s`)190contains information about a whole batch of Agents while `TerminalStep` (no `s`)191only contains information about a single Agent.192 193A `TerminalSteps` has the following fields :194 195- `obs` is a list of numpy arrays observations collected by the group of agent.196  The first dimension of the array corresponds to the batch size of the group197  (number of agents requesting a decision since the last call to `env.step()`).198- `reward` is a float vector of length batch size. Corresponds to the rewards199  collected by each agent since the last simulation step.200- `agent_id` is an int vector of length batch size containing unique identifier201  for the corresponding Agent. This is used to track Agents across simulation202  steps.203 - `interrupted` is an array of booleans of length batch size. Is true if the204 associated Agent was interrupted since the last decision step. For example,205 if the Agent reached the maximum number of steps for the episode.206 207It also has the two following methods:208 209- `len(TerminalSteps)` Returns the number of agents requesting a decision since210  the last call to `env.step()`.211- `TerminalSteps[agent_id]` Returns a `TerminalStep` for the Agent with the212  `agent_id` unique identifier.213 214A `TerminalStep` has the following fields:215 216- `obs` is a list of numpy arrays observations collected by the agent. (Each217  array has one less dimension than the arrays in `TerminalSteps`)218- `reward` is a float. Corresponds to the rewards collected by the agent since219  the last simulation step.220- `agent_id` is an int and an unique identifier for the corresponding Agent.221 - `interrupted` is a bool. Is true if the Agent was interrupted since the last222 decision step. For example, if the Agent reached the maximum number of steps for223 the episode.224 225#### BehaviorSpec226 227A `BehaviorSpec` has the following fields :228 229- `observation_specs` is a List of `ObservationSpec` objects : Each `ObservationSpec`230  corresponds to an observation's properties: `shape` is a tuple of ints that231  corresponds to the shape of the observation (without the number of agents dimension).232  `dimension_property` is a tuple of flags containing extra information about how the233  data should be processed in the corresponding dimension. `observation_type` is an enum234  corresponding to what type of observation is generating the data (i.e., default, goal,235  etc). Note that the `ObservationSpec` have the same ordering as the ordering of observations236  in the DecisionSteps, DecisionStep, TerminalSteps and TerminalStep.237- `action_spec` is an `ActionSpec` namedtuple that defines the number and types238  of actions for the Agent.239 240An `ActionSpec` has the following fields and properties:241- `continuous_size` is the number of floats that constitute the continuous actions.242- `discrete_size` is the number of branches (the number of independent actions) that243  constitute the multi-discrete actions.244- `discrete_branches` is a Tuple of ints. Each int corresponds to the number of245  different options for each branch of the action. For example:246  In a game direction input (no movement, left, right) and247  jump input (no jump, jump) there will be two branches (direction and jump),248  the first one with 3 options and the second with 2 options. (`discrete_size = 2`249  and `discrete_action_branches = (3,2,)`)250 251 252### Communicating additional information with the Environment253 254In addition to the means of communicating between Unity and python described255above, we also provide methods for sharing agent-agnostic information. These256additional methods are referred to as side channels. ML-Agents includes two257ready-made side channels, described below. It is also possible to create custom258side channels to communicate any additional data between a Unity environment and259Python. Instructions for creating custom side channels can be found260[here](Custom-SideChannels.md).261 262Side channels exist as separate classes which are instantiated, and then passed263as list to the `side_channels` argument of the constructor of the264`UnityEnvironment` class.265 266```python267channel = MyChannel()268 269env = UnityEnvironment(side_channels = [channel])270```271 272**Note** : A side channel will only send/receive messages when `env.step` or273`env.reset()` is called.274 275#### EngineConfigurationChannel276 277The `EngineConfiguration` side channel allows you to modify the time-scale,278resolution, and graphics quality of the environment. This can be useful for279adjusting the environment to perform better during training, or be more280interpretable during inference.281 282`EngineConfigurationChannel` has two methods :283 284- `set_configuration_parameters` which takes the following arguments:285  - `width`: Defines the width of the display. (Must be set alongside height)286  - `height`: Defines the height of the display. (Must be set alongside width)287  - `quality_level`: Defines the quality level of the simulation.288  - `time_scale`: Defines the multiplier for the deltatime in the simulation. If289    set to a higher value, time will pass faster in the simulation but the290    physics may perform unpredictably.291  - `target_frame_rate`: Instructs simulation to try to render at a specified292    frame rate.293  - `capture_frame_rate` Instructs the simulation to consider time between294    updates to always be constant, regardless of the actual frame rate.295- `set_configuration` with argument config which is an `EngineConfig` NamedTuple296  object.297 298For example, the following code would adjust the time-scale of the simulation to299be 2x realtime.300 301```python302from mlagents_envs.environment import UnityEnvironment303from mlagents_envs.side_channel.engine_configuration_channel import EngineConfigurationChannel304 305channel = EngineConfigurationChannel()306 307env = UnityEnvironment(side_channels=[channel])308 309channel.set_configuration_parameters(time_scale = 2.0)310 311i = env.reset()312...313```314 315#### EnvironmentParameters316 317The `EnvironmentParameters` will allow you to get and set pre-defined numerical318values in the environment. This can be useful for adjusting environment-specific319settings, or for reading non-agent related information from the environment. You320can call `get_property` and `set_property` on the side channel to read and write321properties.322 323`EnvironmentParametersChannel` has one methods:324 325- `set_float_parameter` Sets a float parameter in the Unity Environment.326  - key: The string identifier of the property.327  - value: The float value of the property.328 329```python330from mlagents_envs.environment import UnityEnvironment331from mlagents_envs.side_channel.environment_parameters_channel import EnvironmentParametersChannel332 333channel = EnvironmentParametersChannel()334 335env = UnityEnvironment(side_channels=[channel])336 337channel.set_float_parameter("parameter_1", 2.0)338 339i = env.reset()340...341```342 343Once a property has been modified in Python, you can access it in C# after the344next call to `step` as follows:345 346```csharp347var envParameters = Academy.Instance.EnvironmentParameters;348float property1 = envParameters.GetWithDefault("parameter_1", 0.0f);349```350 351#### Custom side channels352 353For information on how to make custom side channels for sending additional data354types, see the documentation [here](Custom-SideChannels.md).355