Team Ai
Modelpublic

AnnaMats/ppo-Pyramids-Training

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes110downloads
Python-LLAPI-Documentation.md1353 linesDownload Raw Back to docs
1# Table of Contents2 3* [mlagents\_envs.base\_env](#mlagents_envs.base_env)4  * [DecisionStep](#mlagents_envs.base_env.DecisionStep)5  * [DecisionSteps](#mlagents_envs.base_env.DecisionSteps)6    * [agent\_id\_to\_index](#mlagents_envs.base_env.DecisionSteps.agent_id_to_index)7    * [\_\_getitem\_\_](#mlagents_envs.base_env.DecisionSteps.__getitem__)8    * [empty](#mlagents_envs.base_env.DecisionSteps.empty)9  * [TerminalStep](#mlagents_envs.base_env.TerminalStep)10  * [TerminalSteps](#mlagents_envs.base_env.TerminalSteps)11    * [agent\_id\_to\_index](#mlagents_envs.base_env.TerminalSteps.agent_id_to_index)12    * [\_\_getitem\_\_](#mlagents_envs.base_env.TerminalSteps.__getitem__)13    * [empty](#mlagents_envs.base_env.TerminalSteps.empty)14  * [ActionTuple](#mlagents_envs.base_env.ActionTuple)15    * [discrete\_dtype](#mlagents_envs.base_env.ActionTuple.discrete_dtype)16  * [ActionSpec](#mlagents_envs.base_env.ActionSpec)17    * [is\_discrete](#mlagents_envs.base_env.ActionSpec.is_discrete)18    * [is\_continuous](#mlagents_envs.base_env.ActionSpec.is_continuous)19    * [discrete\_size](#mlagents_envs.base_env.ActionSpec.discrete_size)20    * [empty\_action](#mlagents_envs.base_env.ActionSpec.empty_action)21    * [random\_action](#mlagents_envs.base_env.ActionSpec.random_action)22    * [create\_continuous](#mlagents_envs.base_env.ActionSpec.create_continuous)23    * [create\_discrete](#mlagents_envs.base_env.ActionSpec.create_discrete)24  * [DimensionProperty](#mlagents_envs.base_env.DimensionProperty)25    * [UNSPECIFIED](#mlagents_envs.base_env.DimensionProperty.UNSPECIFIED)26    * [NONE](#mlagents_envs.base_env.DimensionProperty.NONE)27    * [TRANSLATIONAL\_EQUIVARIANCE](#mlagents_envs.base_env.DimensionProperty.TRANSLATIONAL_EQUIVARIANCE)28    * [VARIABLE\_SIZE](#mlagents_envs.base_env.DimensionProperty.VARIABLE_SIZE)29  * [ObservationType](#mlagents_envs.base_env.ObservationType)30    * [DEFAULT](#mlagents_envs.base_env.ObservationType.DEFAULT)31    * [GOAL\_SIGNAL](#mlagents_envs.base_env.ObservationType.GOAL_SIGNAL)32  * [ObservationSpec](#mlagents_envs.base_env.ObservationSpec)33  * [BehaviorSpec](#mlagents_envs.base_env.BehaviorSpec)34  * [BaseEnv](#mlagents_envs.base_env.BaseEnv)35    * [step](#mlagents_envs.base_env.BaseEnv.step)36    * [reset](#mlagents_envs.base_env.BaseEnv.reset)37    * [close](#mlagents_envs.base_env.BaseEnv.close)38    * [behavior\_specs](#mlagents_envs.base_env.BaseEnv.behavior_specs)39    * [set\_actions](#mlagents_envs.base_env.BaseEnv.set_actions)40    * [set\_action\_for\_agent](#mlagents_envs.base_env.BaseEnv.set_action_for_agent)41    * [get\_steps](#mlagents_envs.base_env.BaseEnv.get_steps)42* [mlagents\_envs.environment](#mlagents_envs.environment)43  * [UnityEnvironment](#mlagents_envs.environment.UnityEnvironment)44    * [\_\_init\_\_](#mlagents_envs.environment.UnityEnvironment.__init__)45    * [close](#mlagents_envs.environment.UnityEnvironment.close)46* [mlagents\_envs.registry](#mlagents_envs.registry)47* [mlagents\_envs.registry.unity\_env\_registry](#mlagents_envs.registry.unity_env_registry)48  * [UnityEnvRegistry](#mlagents_envs.registry.unity_env_registry.UnityEnvRegistry)49    * [register](#mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.register)50    * [register\_from\_yaml](#mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.register_from_yaml)51    * [clear](#mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.clear)52    * [\_\_getitem\_\_](#mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.__getitem__)53* [mlagents\_envs.side\_channel](#mlagents_envs.side_channel)54* [mlagents\_envs.side\_channel.raw\_bytes\_channel](#mlagents_envs.side_channel.raw_bytes_channel)55  * [RawBytesChannel](#mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel)56    * [on\_message\_received](#mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel.on_message_received)57    * [get\_and\_clear\_received\_messages](#mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel.get_and_clear_received_messages)58    * [send\_raw\_data](#mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel.send_raw_data)59* [mlagents\_envs.side\_channel.outgoing\_message](#mlagents_envs.side_channel.outgoing_message)60  * [OutgoingMessage](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage)61    * [\_\_init\_\_](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.__init__)62    * [write\_bool](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_bool)63    * [write\_int32](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_int32)64    * [write\_float32](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_float32)65    * [write\_float32\_list](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_float32_list)66    * [write\_string](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_string)67    * [set\_raw\_bytes](#mlagents_envs.side_channel.outgoing_message.OutgoingMessage.set_raw_bytes)68* [mlagents\_envs.side\_channel.engine\_configuration\_channel](#mlagents_envs.side_channel.engine_configuration_channel)69  * [EngineConfigurationChannel](#mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel)70    * [on\_message\_received](#mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel.on_message_received)71    * [set\_configuration\_parameters](#mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel.set_configuration_parameters)72    * [set\_configuration](#mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel.set_configuration)73* [mlagents\_envs.side\_channel.side\_channel\_manager](#mlagents_envs.side_channel.side_channel_manager)74  * [SideChannelManager](#mlagents_envs.side_channel.side_channel_manager.SideChannelManager)75    * [process\_side\_channel\_message](#mlagents_envs.side_channel.side_channel_manager.SideChannelManager.process_side_channel_message)76    * [generate\_side\_channel\_messages](#mlagents_envs.side_channel.side_channel_manager.SideChannelManager.generate_side_channel_messages)77* [mlagents\_envs.side\_channel.stats\_side\_channel](#mlagents_envs.side_channel.stats_side_channel)78  * [StatsSideChannel](#mlagents_envs.side_channel.stats_side_channel.StatsSideChannel)79    * [on\_message\_received](#mlagents_envs.side_channel.stats_side_channel.StatsSideChannel.on_message_received)80    * [get\_and\_reset\_stats](#mlagents_envs.side_channel.stats_side_channel.StatsSideChannel.get_and_reset_stats)81* [mlagents\_envs.side\_channel.incoming\_message](#mlagents_envs.side_channel.incoming_message)82  * [IncomingMessage](#mlagents_envs.side_channel.incoming_message.IncomingMessage)83    * [\_\_init\_\_](#mlagents_envs.side_channel.incoming_message.IncomingMessage.__init__)84    * [read\_bool](#mlagents_envs.side_channel.incoming_message.IncomingMessage.read_bool)85    * [read\_int32](#mlagents_envs.side_channel.incoming_message.IncomingMessage.read_int32)86    * [read\_float32](#mlagents_envs.side_channel.incoming_message.IncomingMessage.read_float32)87    * [read\_float32\_list](#mlagents_envs.side_channel.incoming_message.IncomingMessage.read_float32_list)88    * [read\_string](#mlagents_envs.side_channel.incoming_message.IncomingMessage.read_string)89    * [get\_raw\_bytes](#mlagents_envs.side_channel.incoming_message.IncomingMessage.get_raw_bytes)90* [mlagents\_envs.side\_channel.float\_properties\_channel](#mlagents_envs.side_channel.float_properties_channel)91  * [FloatPropertiesChannel](#mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel)92    * [on\_message\_received](#mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.on_message_received)93    * [set\_property](#mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.set_property)94    * [get\_property](#mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.get_property)95    * [list\_properties](#mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.list_properties)96    * [get\_property\_dict\_copy](#mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.get_property_dict_copy)97* [mlagents\_envs.side\_channel.environment\_parameters\_channel](#mlagents_envs.side_channel.environment_parameters_channel)98  * [EnvironmentParametersChannel](#mlagents_envs.side_channel.environment_parameters_channel.EnvironmentParametersChannel)99    * [set\_float\_parameter](#mlagents_envs.side_channel.environment_parameters_channel.EnvironmentParametersChannel.set_float_parameter)100    * [set\_uniform\_sampler\_parameters](#mlagents_envs.side_channel.environment_parameters_channel.EnvironmentParametersChannel.set_uniform_sampler_parameters)101    * [set\_gaussian\_sampler\_parameters](#mlagents_envs.side_channel.environment_parameters_channel.EnvironmentParametersChannel.set_gaussian_sampler_parameters)102    * [set\_multirangeuniform\_sampler\_parameters](#mlagents_envs.side_channel.environment_parameters_channel.EnvironmentParametersChannel.set_multirangeuniform_sampler_parameters)103* [mlagents\_envs.side\_channel.side\_channel](#mlagents_envs.side_channel.side_channel)104  * [SideChannel](#mlagents_envs.side_channel.side_channel.SideChannel)105    * [queue\_message\_to\_send](#mlagents_envs.side_channel.side_channel.SideChannel.queue_message_to_send)106    * [on\_message\_received](#mlagents_envs.side_channel.side_channel.SideChannel.on_message_received)107    * [channel\_id](#mlagents_envs.side_channel.side_channel.SideChannel.channel_id)108 109<a name="mlagents_envs.base_env"></a>110# mlagents\_envs.base\_env111 112Python Environment API for the ML-Agents Toolkit113The aim of this API is to expose Agents evolving in a simulation114to perform reinforcement learning on.115This API supports multi-agent scenarios and groups similar Agents (same116observations, actions spaces and behavior) together. These groups of Agents are117identified by their BehaviorName.118For performance reasons, the data of each group of agents is processed in a119batched manner. Agents are identified by a unique AgentId identifier that120allows tracking of Agents across simulation steps. Note that there is no121guarantee that the number or order of the Agents in the state will be122consistent across simulation steps.123A simulation steps corresponds to moving the simulation forward until at least124one agent in the simulation sends its observations to Python again. Since125Agents can request decisions at different frequencies, a simulation step does126not necessarily correspond to a fixed simulation time increment.127 128<a name="mlagents_envs.base_env.DecisionStep"></a>129## DecisionStep Objects130 131```python132class DecisionStep(NamedTuple)133```134 135Contains the data a single Agent collected since the last136simulation step.137 - obs is a list of numpy arrays observations collected by the agent.138 - reward is a float. Corresponds to the rewards collected by the agent139 since the last simulation step.140 - agent_id is an int and an unique identifier for the corresponding Agent.141 - action_mask is an optional list of one dimensional array of booleans.142 Only available when using multi-discrete actions.143 Each array corresponds to an action branch. Each array contains a mask144 for each action of the branch. If true, the action is not available for145 the agent during this simulation step.146 147<a name="mlagents_envs.base_env.DecisionSteps"></a>148## DecisionSteps Objects149 150```python151class DecisionSteps(Mapping)152```153 154Contains the data a batch of similar Agents collected since the last155simulation step. Note that all Agents do not necessarily have new156information to send at each simulation step. Therefore, the ordering of157agents and the batch size of the DecisionSteps are not fixed across158simulation steps.159 - obs is a list of numpy arrays observations collected by the batch of160 agent. Each obs has one extra dimension compared to DecisionStep: the161 first dimension of the array corresponds to the batch size of the batch.162 - reward is a float vector of length batch size. Corresponds to the163 rewards collected by each agent since the last simulation step.164 - agent_id is an int vector of length batch size containing unique165 identifier for the corresponding Agent. This is used to track Agents166 across simulation steps.167 - action_mask is an optional list of two dimensional array of booleans.168 Only available when using multi-discrete actions.169 Each array corresponds to an action branch. The first dimension of each170 array is the batch size and the second contains a mask for each action of171 the branch. If true, the action is not available for the agent during172 this simulation step.173 174<a name="mlagents_envs.base_env.DecisionSteps.agent_id_to_index"></a>175#### agent\_id\_to\_index176 177```python178 | @property179 | agent_id_to_index() -> Dict[AgentId, int]180```181 182**Returns**:183 184A Dict that maps agent_id to the index of those agents in185this DecisionSteps.186 187<a name="mlagents_envs.base_env.DecisionSteps.__getitem__"></a>188#### \_\_getitem\_\_189 190```python191 | __getitem__(agent_id: AgentId) -> DecisionStep192```193 194returns the DecisionStep for a specific agent.195 196**Arguments**:197 198- `agent_id`: The id of the agent199 200**Returns**:201 202The DecisionStep203 204<a name="mlagents_envs.base_env.DecisionSteps.empty"></a>205#### empty206 207```python208 | @staticmethod209 | empty(spec: "BehaviorSpec") -> "DecisionSteps"210```211 212Returns an empty DecisionSteps.213 214**Arguments**:215 216- `spec`: The BehaviorSpec for the DecisionSteps217 218<a name="mlagents_envs.base_env.TerminalStep"></a>219## TerminalStep Objects220 221```python222class TerminalStep(NamedTuple)223```224 225Contains the data a single Agent collected when its episode ended.226 - obs is a list of numpy arrays observations collected by the agent.227 - reward is a float. Corresponds to the rewards collected by the agent228 since the last simulation step.229 - interrupted is a bool. Is true if the Agent was interrupted since the last230 decision step. For example, if the Agent reached the maximum number of steps for231 the episode.232 - agent_id is an int and an unique identifier for the corresponding Agent.233 234<a name="mlagents_envs.base_env.TerminalSteps"></a>235## TerminalSteps Objects236 237```python238class TerminalSteps(Mapping)239```240 241Contains the data a batch of Agents collected when their episode242terminated. All Agents present in the TerminalSteps have ended their243episode.244 - obs is a list of numpy arrays observations collected by the batch of245 agent. Each obs has one extra dimension compared to DecisionStep: the246 first dimension of the array corresponds to the batch size of the batch.247 - reward is a float vector of length batch size. Corresponds to the248 rewards collected by each agent since the last simulation step.249 - interrupted is an array of booleans of length batch size. Is true if the250 associated Agent was interrupted since the last decision step. For example, if the251 Agent reached the maximum number of steps for the episode.252 - agent_id is an int vector of length batch size containing unique253 identifier for the corresponding Agent. This is used to track Agents254 across simulation steps.255 256<a name="mlagents_envs.base_env.TerminalSteps.agent_id_to_index"></a>257#### agent\_id\_to\_index258 259```python260 | @property261 | agent_id_to_index() -> Dict[AgentId, int]262```263 264**Returns**:265 266A Dict that maps agent_id to the index of those agents in267this TerminalSteps.268 269<a name="mlagents_envs.base_env.TerminalSteps.__getitem__"></a>270#### \_\_getitem\_\_271 272```python273 | __getitem__(agent_id: AgentId) -> TerminalStep274```275 276returns the TerminalStep for a specific agent.277 278**Arguments**:279 280- `agent_id`: The id of the agent281 282**Returns**:283 284obs, reward, done, agent_id and optional action mask for a285specific agent286 287<a name="mlagents_envs.base_env.TerminalSteps.empty"></a>288#### empty289 290```python291 | @staticmethod292 | empty(spec: "BehaviorSpec") -> "TerminalSteps"293```294 295Returns an empty TerminalSteps.296 297**Arguments**:298 299- `spec`: The BehaviorSpec for the TerminalSteps300 301<a name="mlagents_envs.base_env.ActionTuple"></a>302## ActionTuple Objects303 304```python305class ActionTuple(_ActionTupleBase)306```307 308An object whose fields correspond to actions of different types.309Continuous and discrete actions are numpy arrays of type float32 and310int32, respectively and are type checked on construction.311Dimensions are of (n_agents, continuous_size) and (n_agents, discrete_size),312respectively. Note, this also holds when continuous or discrete size is313zero.314 315<a name="mlagents_envs.base_env.ActionTuple.discrete_dtype"></a>316#### discrete\_dtype317 318```python319 | @property320 | discrete_dtype() -> np.dtype321```322 323The dtype of a discrete action.324 325<a name="mlagents_envs.base_env.ActionSpec"></a>326## ActionSpec Objects327 328```python329class ActionSpec(NamedTuple)330```331 332A NamedTuple containing utility functions and information about the action spaces333for a group of Agents under the same behavior.334- num_continuous_actions is an int corresponding to the number of floats which335constitute the action.336- discrete_branch_sizes is a Tuple of int where each int corresponds to337the number of discrete actions available to the agent on an independent action branch.338 339<a name="mlagents_envs.base_env.ActionSpec.is_discrete"></a>340#### is\_discrete341 342```python343 | is_discrete() -> bool344```345 346Returns true if this Behavior uses discrete actions347 348<a name="mlagents_envs.base_env.ActionSpec.is_continuous"></a>349#### is\_continuous350 351```python352 | is_continuous() -> bool353```354 355Returns true if this Behavior uses continuous actions356 357<a name="mlagents_envs.base_env.ActionSpec.discrete_size"></a>358#### discrete\_size359 360```python361 | @property362 | discrete_size() -> int363```364 365Returns a an int corresponding to the number of discrete branches.366 367<a name="mlagents_envs.base_env.ActionSpec.empty_action"></a>368#### empty\_action369 370```python371 | empty_action(n_agents: int) -> ActionTuple372```373 374Generates ActionTuple corresponding to an empty action (all zeros)375for a number of agents.376 377**Arguments**:378 379- `n_agents`: The number of agents that will have actions generated380 381<a name="mlagents_envs.base_env.ActionSpec.random_action"></a>382#### random\_action383 384```python385 | random_action(n_agents: int) -> ActionTuple386```387 388Generates ActionTuple corresponding to a random action (either discrete389or continuous) for a number of agents.390 391**Arguments**:392 393- `n_agents`: The number of agents that will have actions generated394 395<a name="mlagents_envs.base_env.ActionSpec.create_continuous"></a>396#### create\_continuous397 398```python399 | @staticmethod400 | create_continuous(continuous_size: int) -> "ActionSpec"401```402 403Creates an ActionSpec that is homogenously continuous404 405<a name="mlagents_envs.base_env.ActionSpec.create_discrete"></a>406#### create\_discrete407 408```python409 | @staticmethod410 | create_discrete(discrete_branches: Tuple[int]) -> "ActionSpec"411```412 413Creates an ActionSpec that is homogenously discrete414 415<a name="mlagents_envs.base_env.DimensionProperty"></a>416## DimensionProperty Objects417 418```python419class DimensionProperty(IntFlag)420```421 422The dimension property of a dimension of an observation.423 424<a name="mlagents_envs.base_env.DimensionProperty.UNSPECIFIED"></a>425#### UNSPECIFIED426 427No properties specified.428 429<a name="mlagents_envs.base_env.DimensionProperty.NONE"></a>430#### NONE431 432No Property of the observation in that dimension. Observation can be processed with433Fully connected networks.434 435<a name="mlagents_envs.base_env.DimensionProperty.TRANSLATIONAL_EQUIVARIANCE"></a>436#### TRANSLATIONAL\_EQUIVARIANCE437 438Means it is suitable to do a convolution in this dimension.439 440<a name="mlagents_envs.base_env.DimensionProperty.VARIABLE_SIZE"></a>441#### VARIABLE\_SIZE442 443Means that there can be a variable number of observations in this dimension.444The observations are unordered.445 446<a name="mlagents_envs.base_env.ObservationType"></a>447## ObservationType Objects448 449```python450class ObservationType(Enum)451```452 453An Enum which defines the type of information carried in the observation454of the agent.455 456<a name="mlagents_envs.base_env.ObservationType.DEFAULT"></a>457#### DEFAULT458 459Observation information is generic.460 461<a name="mlagents_envs.base_env.ObservationType.GOAL_SIGNAL"></a>462#### GOAL\_SIGNAL463 464Observation contains goal information for current task.465 466<a name="mlagents_envs.base_env.ObservationSpec"></a>467## ObservationSpec Objects468 469```python470class ObservationSpec(NamedTuple)471```472 473A NamedTuple containing information about the observation of Agents.474- shape is a Tuple of int : It corresponds to the shape of475an observation's dimensions.476- dimension_property is a Tuple of DimensionProperties flag, one flag for each477dimension.478- observation_type is an enum of ObservationType.479 480<a name="mlagents_envs.base_env.BehaviorSpec"></a>481## BehaviorSpec Objects482 483```python484class BehaviorSpec(NamedTuple)485```486 487A NamedTuple containing information about the observation and action488spaces for a group of Agents under the same behavior.489- observation_specs is a List of ObservationSpec NamedTuple containing490information about the information of the Agent's observations such as their shapes.491The order of the ObservationSpec is the same as the order of the observations of an492agent.493- action_spec is an ActionSpec NamedTuple.494 495<a name="mlagents_envs.base_env.BaseEnv"></a>496## BaseEnv Objects497 498```python499class BaseEnv(ABC)500```501 502<a name="mlagents_envs.base_env.BaseEnv.step"></a>503#### step504 505```python506 | @abstractmethod507 | step() -> None508```509 510Signals the environment that it must move the simulation forward511by one step.512 513<a name="mlagents_envs.base_env.BaseEnv.reset"></a>514#### reset515 516```python517 | @abstractmethod518 | reset() -> None519```520 521Signals the environment that it must reset the simulation.522 523<a name="mlagents_envs.base_env.BaseEnv.close"></a>524#### close525 526```python527 | @abstractmethod528 | close() -> None529```530 531Signals the environment that it must close.532 533<a name="mlagents_envs.base_env.BaseEnv.behavior_specs"></a>534#### behavior\_specs535 536```python537 | @property538 | @abstractmethod539 | behavior_specs() -> MappingType[str, BehaviorSpec]540```541 542Returns a Mapping from behavior names to behavior specs.543Agents grouped under the same behavior name have the same action and544observation specs, and are expected to behave similarly in the545environment.546Note that new keys can be added to this mapping as new policies are instantiated.547 548<a name="mlagents_envs.base_env.BaseEnv.set_actions"></a>549#### set\_actions550 551```python552 | @abstractmethod553 | set_actions(behavior_name: BehaviorName, action: ActionTuple) -> None554```555 556Sets the action for all of the agents in the simulation for the next557step. The Actions must be in the same order as the order received in558the DecisionSteps.559 560**Arguments**:561 562- `behavior_name`: The name of the behavior the agents are part of563- `action`: ActionTuple tuple of continuous and/or discrete action.564Actions are np.arrays with dimensions  (n_agents, continuous_size) and565(n_agents, discrete_size), respectively.566 567<a name="mlagents_envs.base_env.BaseEnv.set_action_for_agent"></a>568#### set\_action\_for\_agent569 570```python571 | @abstractmethod572 | set_action_for_agent(behavior_name: BehaviorName, agent_id: AgentId, action: ActionTuple) -> None573```574 575Sets the action for one of the agents in the simulation for the next576step.577 578**Arguments**:579 580- `behavior_name`: The name of the behavior the agent is part of581- `agent_id`: The id of the agent the action is set for582- `action`: ActionTuple tuple of continuous and/or discrete action583Actions are np.arrays with dimensions  (1, continuous_size) and584(1, discrete_size), respectively. Note, this initial dimensions of 1 is because585this action is meant for a single agent.586 587<a name="mlagents_envs.base_env.BaseEnv.get_steps"></a>588#### get\_steps589 590```python591 | @abstractmethod592 | get_steps(behavior_name: BehaviorName) -> Tuple[DecisionSteps, TerminalSteps]593```594 595Retrieves the steps of the agents that requested a step in the596simulation.597 598**Arguments**:599 600- `behavior_name`: The name of the behavior the agents are part of601 602**Returns**:603 604A tuple containing :605- A DecisionSteps NamedTuple containing the observations,606the rewards, the agent ids and the action masks for the Agents607of the specified behavior. These Agents need an action this step.608- A TerminalSteps NamedTuple containing the observations,609rewards, agent ids and interrupted flags of the agents that had their610episode terminated last step.611 612<a name="mlagents_envs.environment"></a>613# mlagents\_envs.environment614 615<a name="mlagents_envs.environment.UnityEnvironment"></a>616## UnityEnvironment Objects617 618```python619class UnityEnvironment(BaseEnv)620```621 622<a name="mlagents_envs.environment.UnityEnvironment.__init__"></a>623#### \_\_init\_\_624 625```python626 | __init__(file_name: Optional[str] = None, worker_id: int = 0, base_port: Optional[int] = None, seed: int = 0, no_graphics: bool = False, timeout_wait: int = 60, additional_args: Optional[List[str]] = None, side_channels: Optional[List[SideChannel]] = None, log_folder: Optional[str] = None, num_areas: int = 1)627```628 629Starts a new unity environment and establishes a connection with the environment.630Notice: Currently communication between Unity and Python takes place over an open socket without authentication.631Ensure that the network where training takes place is secure.632 633:string file_name: Name of Unity environment binary.634:int base_port: Baseline port number to connect to Unity environment over. worker_id increments over this.635If no environment is specified (i.e. file_name is None), the DEFAULT_EDITOR_PORT will be used.636:int worker_id: Offset from base_port. Used for training multiple environments simultaneously.637:bool no_graphics: Whether to run the Unity simulator in no-graphics mode638:int timeout_wait: Time (in seconds) to wait for connection from environment.639:list args: Addition Unity command line arguments640:list side_channels: Additional side channel for no-rl communication with Unity641:str log_folder: Optional folder to write the Unity Player log file into.  Requires absolute path.642 643<a name="mlagents_envs.environment.UnityEnvironment.close"></a>644#### close645 646```python647 | close()648```649 650Sends a shutdown signal to the unity environment, and closes the socket connection.651 652<a name="mlagents_envs.registry"></a>653# mlagents\_envs.registry654 655<a name="mlagents_envs.registry.unity_env_registry"></a>656# mlagents\_envs.registry.unity\_env\_registry657 658<a name="mlagents_envs.registry.unity_env_registry.UnityEnvRegistry"></a>659## UnityEnvRegistry Objects660 661```python662class UnityEnvRegistry(Mapping)663```664 665### UnityEnvRegistry666Provides a library of Unity environments that can be launched without the need667of downloading the Unity Editor.668The UnityEnvRegistry implements a Map, to access an entry of the Registry, use:669```python670registry = UnityEnvRegistry()671entry = registry[<environment_identifyier>]672```673An entry has the following properties :674 * `identifier` : Uniquely identifies this environment675 * `expected_reward` : Corresponds to the reward an agent must obtained for the task676 to be considered completed.677 * `description` : A human readable description of the environment.678 679To launch a Unity environment from a registry entry, use the `make` method:680```python681registry = UnityEnvRegistry()682env = registry[<environment_identifyier>].make()683```684 685<a name="mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.register"></a>686#### register687 688```python689 | register(new_entry: BaseRegistryEntry) -> None690```691 692Registers a new BaseRegistryEntry to the registry. The693BaseRegistryEntry.identifier value will be used as indexing key.694If two are more environments are registered under the same key, the most695recentry added will replace the others.696 697<a name="mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.register_from_yaml"></a>698#### register\_from\_yaml699 700```python701 | register_from_yaml(path_to_yaml: str) -> None702```703 704Registers the environments listed in a yaml file (either local or remote). Note705that the entries are registered lazily: the registration will only happen when706an environment is accessed.707The yaml file must have the following format :708```yaml709environments:710- <identifier of the first environment>:711    expected_reward: <expected reward of the environment>712    description: | <a multi line description of the environment>713      <continued multi line description>714    linux_url: <The url for the Linux executable zip file>715    darwin_url: <The url for the OSX executable zip file>716    win_url: <The url for the Windows executable zip file>717 718- <identifier of the second environment>:719    expected_reward: <expected reward of the environment>720    description: | <a multi line description of the environment>721      <continued multi line description>722    linux_url: <The url for the Linux executable zip file>723    darwin_url: <The url for the OSX executable zip file>724    win_url: <The url for the Windows executable zip file>725 726- ...727```728 729**Arguments**:730 731- `path_to_yaml`: A local path or url to the yaml file732 733<a name="mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.clear"></a>734#### clear735 736```python737 | clear() -> None738```739 740Deletes all entries in the registry.741 742<a name="mlagents_envs.registry.unity_env_registry.UnityEnvRegistry.__getitem__"></a>743#### \_\_getitem\_\_744 745```python746 | __getitem__(identifier: str) -> BaseRegistryEntry747```748 749Returns the BaseRegistryEntry with the provided identifier. BaseRegistryEntry750can then be used to make a Unity Environment.751 752**Arguments**:753 754- `identifier`: The identifier of the BaseRegistryEntry755 756**Returns**:757 758The associated BaseRegistryEntry759 760<a name="mlagents_envs.side_channel"></a>761# mlagents\_envs.side\_channel762 763<a name="mlagents_envs.side_channel.raw_bytes_channel"></a>764# mlagents\_envs.side\_channel.raw\_bytes\_channel765 766<a name="mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel"></a>767## RawBytesChannel Objects768 769```python770class RawBytesChannel(SideChannel)771```772 773This is an example of what the SideChannel for raw bytes exchange would774look like. Is meant to be used for general research purpose.775 776<a name="mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel.on_message_received"></a>777#### on\_message\_received778 779```python780 | on_message_received(msg: IncomingMessage) -> None781```782 783Is called by the environment to the side channel. Can be called784multiple times per step if multiple messages are meant for that785SideChannel.786 787<a name="mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel.get_and_clear_received_messages"></a>788#### get\_and\_clear\_received\_messages789 790```python791 | get_and_clear_received_messages() -> List[bytes]792```793 794returns a list of bytearray received from the environment.795 796<a name="mlagents_envs.side_channel.raw_bytes_channel.RawBytesChannel.send_raw_data"></a>797#### send\_raw\_data798 799```python800 | send_raw_data(data: bytearray) -> None801```802 803Queues a message to be sent by the environment at the next call to804step.805 806<a name="mlagents_envs.side_channel.outgoing_message"></a>807# mlagents\_envs.side\_channel.outgoing\_message808 809<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage"></a>810## OutgoingMessage Objects811 812```python813class OutgoingMessage()814```815 816Utility class for forming the message that is written to a SideChannel.817All data is written in little-endian format using the struct module.818 819<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.__init__"></a>820#### \_\_init\_\_821 822```python823 | __init__()824```825 826Create an OutgoingMessage with an empty buffer.827 828<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_bool"></a>829#### write\_bool830 831```python832 | write_bool(b: bool) -> None833```834 835Append a boolean value.836 837<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_int32"></a>838#### write\_int32839 840```python841 | write_int32(i: int) -> None842```843 844Append an integer value.845 846<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_float32"></a>847#### write\_float32848 849```python850 | write_float32(f: float) -> None851```852 853Append a float value. It will be truncated to 32-bit precision.854 855<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_float32_list"></a>856#### write\_float32\_list857 858```python859 | write_float32_list(float_list: List[float]) -> None860```861 862Append a list of float values. They will be truncated to 32-bit precision.863 864<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.write_string"></a>865#### write\_string866 867```python868 | write_string(s: str) -> None869```870 871Append a string value. Internally, it will be encoded to ascii, and the872encoded length will also be written to the message.873 874<a name="mlagents_envs.side_channel.outgoing_message.OutgoingMessage.set_raw_bytes"></a>875#### set\_raw\_bytes876 877```python878 | set_raw_bytes(buffer: bytearray) -> None879```880 881Set the internal buffer to a new bytearray. This will overwrite any existing data.882 883**Arguments**:884 885- `buffer`:886 887**Returns**:888 889 890 891<a name="mlagents_envs.side_channel.engine_configuration_channel"></a>892# mlagents\_envs.side\_channel.engine\_configuration\_channel893 894<a name="mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel"></a>895## EngineConfigurationChannel Objects896 897```python898class EngineConfigurationChannel(SideChannel)899```900 901This is the SideChannel for engine configuration exchange. The data in the902engine configuration is as follows :903 - int width;904 - int height;905 - int qualityLevel;906 - float timeScale;907 - int targetFrameRate;908 - int captureFrameRate;909 910<a name="mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel.on_message_received"></a>911#### on\_message\_received912 913```python914 | on_message_received(msg: IncomingMessage) -> None915```916 917Is called by the environment to the side channel. Can be called918multiple times per step if multiple messages are meant for that919SideChannel.920Note that Python should never receive an engine configuration from921Unity922 923<a name="mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel.set_configuration_parameters"></a>924#### set\_configuration\_parameters925 926```python927 | set_configuration_parameters(width: Optional[int] = None, height: Optional[int] = None, quality_level: Optional[int] = None, time_scale: Optional[float] = None, target_frame_rate: Optional[int] = None, capture_frame_rate: Optional[int] = None) -> None928```929 930Sets the engine configuration. Takes as input the configurations of the931engine.932 933**Arguments**:934 935- `width`: Defines the width of the display. (Must be set alongside height)936- `height`: Defines the height of the display. (Must be set alongside width)937- `quality_level`: Defines the quality level of the simulation.938- `time_scale`: Defines the multiplier for the deltatime in the939simulation. If set to a higher value, time will pass faster in the940simulation but the physics might break.941- `target_frame_rate`: Instructs simulation to try to render at a942specified frame rate.943- `capture_frame_rate`: Instructs the simulation to consider time between944updates to always be constant, regardless of the actual frame rate.945 946<a name="mlagents_envs.side_channel.engine_configuration_channel.EngineConfigurationChannel.set_configuration"></a>947#### set\_configuration948 949```python950 | set_configuration(config: EngineConfig) -> None951```952 953Sets the engine configuration. Takes as input an EngineConfig.954 955<a name="mlagents_envs.side_channel.side_channel_manager"></a>956# mlagents\_envs.side\_channel.side\_channel\_manager957 958<a name="mlagents_envs.side_channel.side_channel_manager.SideChannelManager"></a>959## SideChannelManager Objects960 961```python962class SideChannelManager()963```964 965<a name="mlagents_envs.side_channel.side_channel_manager.SideChannelManager.process_side_channel_message"></a>966#### process\_side\_channel\_message967 968```python969 | process_side_channel_message(data: bytes) -> None970```971 972Separates the data received from Python into individual messages for each973registered side channel and calls on_message_received on them.974 975**Arguments**:976 977- `data`: The packed message sent by Unity978 979<a name="mlagents_envs.side_channel.side_channel_manager.SideChannelManager.generate_side_channel_messages"></a>980#### generate\_side\_channel\_messages981 982```python983 | generate_side_channel_messages() -> bytearray984```985 986Gathers the messages that the registered side channels will send to Unity987and combines them into a single message ready to be sent.988 989<a name="mlagents_envs.side_channel.stats_side_channel"></a>990# mlagents\_envs.side\_channel.stats\_side\_channel991 992<a name="mlagents_envs.side_channel.stats_side_channel.StatsSideChannel"></a>993## StatsSideChannel Objects994 995```python996class StatsSideChannel(SideChannel)997```998 999Side channel that receives (string, float) pairs from the environment, so that they can eventually1000be passed to a StatsReporter.1001 1002<a name="mlagents_envs.side_channel.stats_side_channel.StatsSideChannel.on_message_received"></a>1003#### on\_message\_received1004 1005```python1006 | on_message_received(msg: IncomingMessage) -> None1007```1008 1009Receive the message from the environment, and save it for later retrieval.1010 1011**Arguments**:1012 1013- `msg`:1014 1015**Returns**:1016 1017 1018 1019<a name="mlagents_envs.side_channel.stats_side_channel.StatsSideChannel.get_and_reset_stats"></a>1020#### get\_and\_reset\_stats1021 1022```python1023 | get_and_reset_stats() -> EnvironmentStats1024```1025 1026Returns the current stats, and resets the internal storage of the stats.1027 1028**Returns**:1029 1030 1031 1032<a name="mlagents_envs.side_channel.incoming_message"></a>1033# mlagents\_envs.side\_channel.incoming\_message1034 1035<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage"></a>1036## IncomingMessage Objects1037 1038```python1039class IncomingMessage()1040```1041 1042Utility class for reading the message written to a SideChannel.1043Values must be read in the order they were written.1044 1045<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.__init__"></a>1046#### \_\_init\_\_1047 1048```python1049 | __init__(buffer: bytes, offset: int = 0)1050```1051 1052Create a new IncomingMessage from the bytes.1053 1054<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.read_bool"></a>1055#### read\_bool1056 1057```python1058 | read_bool(default_value: bool = False) -> bool1059```1060 1061Read a boolean value from the message buffer.1062 1063**Arguments**:1064 1065- `default_value`: Default value to use if the end of the message is reached.1066 1067**Returns**:1068 1069The value read from the message, or the default value if the end was reached.1070 1071<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.read_int32"></a>1072#### read\_int321073 1074```python1075 | read_int32(default_value: int = 0) -> int1076```1077 1078Read an integer value from the message buffer.1079 1080**Arguments**:1081 1082- `default_value`: Default value to use if the end of the message is reached.1083 1084**Returns**:1085 1086The value read from the message, or the default value if the end was reached.1087 1088<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.read_float32"></a>1089#### read\_float321090 1091```python1092 | read_float32(default_value: float = 0.0) -> float1093```1094 1095Read a float value from the message buffer.1096 1097**Arguments**:1098 1099- `default_value`: Default value to use if the end of the message is reached.1100 1101**Returns**:1102 1103The value read from the message, or the default value if the end was reached.1104 1105<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.read_float32_list"></a>1106#### read\_float32\_list1107 1108```python1109 | read_float32_list(default_value: List[float] = None) -> List[float]1110```1111 1112Read a list of float values from the message buffer.1113 1114**Arguments**:1115 1116- `default_value`: Default value to use if the end of the message is reached.1117 1118**Returns**:1119 1120The value read from the message, or the default value if the end was reached.1121 1122<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.read_string"></a>1123#### read\_string1124 1125```python1126 | read_string(default_value: str = "") -> str1127```1128 1129Read a string value from the message buffer.1130 1131**Arguments**:1132 1133- `default_value`: Default value to use if the end of the message is reached.1134 1135**Returns**:1136 1137The value read from the message, or the default value if the end was reached.1138 1139<a name="mlagents_envs.side_channel.incoming_message.IncomingMessage.get_raw_bytes"></a>1140#### get\_raw\_bytes1141 1142```python1143 | get_raw_bytes() -> bytes1144```1145 1146Get a copy of the internal bytes used by the message.1147 1148<a name="mlagents_envs.side_channel.float_properties_channel"></a>1149# mlagents\_envs.side\_channel.float\_properties\_channel1150 1151<a name="mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel"></a>1152## FloatPropertiesChannel Objects1153 1154```python1155class FloatPropertiesChannel(SideChannel)1156```1157 1158This is the SideChannel for float properties shared with Unity.1159You can modify the float properties of an environment with the commands1160set_property, get_property and list_properties.1161 1162<a name="mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.on_message_received"></a>1163#### on\_message\_received1164 1165```python1166 | on_message_received(msg: IncomingMessage) -> None1167```1168 1169Is called by the environment to the side channel. Can be called1170multiple times per step if multiple messages are meant for that1171SideChannel.1172 1173<a name="mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.set_property"></a>1174#### set\_property1175 1176```python1177 | set_property(key: str, value: float) -> None1178```1179 1180Sets a property in the Unity Environment.1181 1182**Arguments**:1183 1184- `key`: The string identifier of the property.1185- `value`: The float value of the property.1186 1187<a name="mlagents_envs.side_channel.float_properties_channel.FloatPropertiesChannel.get_property"></a>1188#### get\_property1189 1190```python1191 | get_property(key: str) -> Optional[float]1192```1193 1194Gets a property in the Unity Environment. If the property was not1195found, will return None.1196 1197**Arguments**:1198 1199- `key`: The string identifier of the property.1200 

Showing the first 1,200 of 1353 lines. Download the file for the rest.