Team Ai
Modelpublic

AnnaMats/ppo-Pyramids-Training

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes110downloads
Learning-Environment-Examples.md501 linesDownload Raw Back to docs
1# Example Learning Environments2 3<img src="../images/example-envs.png" align="middle" width="3000"/>4 5The Unity ML-Agents Toolkit includes an expanding set of example environments6that highlight the various features of the toolkit. These environments can also7serve as templates for new environments or as ways to test new ML algorithms.8Environments are located in `Project/Assets/ML-Agents/Examples` and summarized9below.10 11For the environments that highlight specific features of the toolkit, we provide12the pre-trained model files and the training config file that enables you to13train the scene yourself. The environments that are designed to serve as14challenges for researchers do not have accompanying pre-trained model files or15training configs and are marked as _Optional_ below.16 17This page only overviews the example environments we provide. To learn more on18how to design and build your own environments see our19[Making a New Learning Environment](Learning-Environment-Create-New.md) page. If20you would like to contribute environments, please see our21[contribution guidelines](CONTRIBUTING.md) page.22 23## Basic24 25![Basic](images/basic.png)26 27- Set-up: A linear movement task where the agent must move left or right to28  rewarding states.29- Goal: Move to the most reward state.30- Agents: The environment contains one agent.31- Agent Reward Function:32  - -0.01 at each step33  - +0.1 for arriving at suboptimal state.34  - +1.0 for arriving at optimal state.35- Behavior Parameters:36  - Vector Observation space: One variable corresponding to current state.37  - Actions: 1 discrete action branch with 3 actions (Move left, do nothing, move38    right).39  - Visual Observations: None40- Float Properties: None41- Benchmark Mean Reward: 0.9342 43## 3DBall: 3D Balance Ball44 45![3D Balance Ball](images/balance.png)46 47- Set-up: A balance-ball task, where the agent balances the ball on it's head.48- Goal: The agent must balance the ball on it's head for as long as possible.49- Agents: The environment contains 12 agents of the same kind, all using the50  same Behavior Parameters.51- Agent Reward Function:52  - +0.1 for every step the ball remains on it's head.53  - -1.0 if the ball falls off.54- Behavior Parameters:55  - Vector Observation space: 8 variables corresponding to rotation of the agent56    cube, and position and velocity of ball.57  - Vector Observation space (Hard Version): 5 variables corresponding to58    rotation of the agent cube and position of ball.59  - Actions: 2 continuous actions, with one value corresponding to60    X-rotation, and the other to Z-rotation.61  - Visual Observations: Third-person view from the upper-front of the agent. Use62    `Visual3DBall` scene.63- Float Properties: Three64  - scale: Specifies the scale of the ball in the 3 dimensions (equal across the65    three dimensions)66    - Default: 167    - Recommended Minimum: 0.268    - Recommended Maximum: 569  - gravity: Magnitude of gravity70    - Default: 9.8171    - Recommended Minimum: 472    - Recommended Maximum: 10573  - mass: Specifies mass of the ball74    - Default: 175    - Recommended Minimum: 0.176    - Recommended Maximum: 2077- Benchmark Mean Reward: 10078 79## GridWorld80 81![GridWorld](images/gridworld.png)82 83- Set-up: A multi-goal version of the grid-world task. Scene contains agent, goal,84  and obstacles.85- Goal: The agent must navigate the grid to the appropriate goal while86  avoiding the obstacles.87- Agents: The environment contains nine agents with the same Behavior88  Parameters.89- Agent Reward Function:90  - -0.01 for every step.91  - +1.0 if the agent navigates to the correct goal (episode ends).92  - -1.0 if the agent navigates to an incorrect goal (episode ends).93- Behavior Parameters:94  - Vector Observation space: None95  - Actions: 1 discrete action branch with 5 actions, corresponding to movement in96    cardinal directions or not moving. Note that for this environment,97    [action masking](Learning-Environment-Design-Agents.md#masking-discrete-actions)98    is turned on by default (this option can be toggled using the `Mask Actions`99    checkbox within the `trueAgent` GameObject). The trained model file provided100    was generated with action masking turned on.101  - Visual Observations: One corresponding to top-down view of GridWorld.102  - Goal Signal : A one hot vector corresponding to which color is the correct goal103  for the Agent104- Float Properties: Three, corresponding to grid size, number of green goals, and105  number of red goals.106- Benchmark Mean Reward: 0.8107 108## Push Block109 110![Push](images/push.png)111 112- Set-up: A platforming environment where the agent can push a block around.113- Goal: The agent must push the block to the goal.114- Agents: The environment contains one agent.115- Agent Reward Function:116  - -0.0025 for every step.117  - +1.0 if the block touches the goal.118- Behavior Parameters:119  - Vector Observation space: (Continuous) 70 variables corresponding to 14120    ray-casts each detecting one of three possible objects (wall, goal, or121    block).122  - Actions: 1 discrete action branch with 7 actions, corresponding to turn clockwise123    and counterclockwise, move along four different face directions, or do nothing.124- Float Properties: Four125  - block_scale: Scale of the block along the x and z dimensions126    - Default: 2127    - Recommended Minimum: 0.5128    - Recommended Maximum: 4129  - dynamic_friction: Coefficient of friction for the ground material acting on130    moving objects131    - Default: 0132    - Recommended Minimum: 0133    - Recommended Maximum: 1134  - static_friction: Coefficient of friction for the ground material acting on135    stationary objects136    - Default: 0137    - Recommended Minimum: 0138    - Recommended Maximum: 1139  - block_drag: Effect of air resistance on block140    - Default: 0.5141    - Recommended Minimum: 0142    - Recommended Maximum: 2000143- Benchmark Mean Reward: 4.5144 145## Wall Jump146 147![Wall](images/wall.png)148 149- Set-up: A platforming environment where the agent can jump over a wall.150- Goal: The agent must use the block to scale the wall and reach the goal.151- Agents: The environment contains one agent linked to two different Models. The152  Policy the agent is linked to changes depending on the height of the wall. The153  change of Policy is done in the WallJumpAgent class.154- Agent Reward Function:155  - -0.0005 for every step.156  - +1.0 if the agent touches the goal.157  - -1.0 if the agent falls off the platform.158- Behavior Parameters:159  - Vector Observation space: Size of 74, corresponding to 14 ray casts each160    detecting 4 possible objects. plus the global position of the agent and161    whether or not the agent is grounded.162  - Actions: 4 discrete action branches:163    - Forward Motion (3 possible actions: Forward, Backwards, No Action)164    - Rotation (3 possible actions: Rotate Left, Rotate Right, No Action)165    - Side Motion (3 possible actions: Left, Right, No Action)166    - Jump (2 possible actions: Jump, No Action)167  - Visual Observations: None168- Float Properties: Four169- Benchmark Mean Reward (Big & Small Wall): 0.8170 171## Crawler172 173![Crawler](images/crawler.png)174 175- Set-up: A creature with 4 arms and 4 forearms.176- Goal: The agents must move its body toward the goal direction without falling.177- Agents: The environment contains 10 agents with same Behavior Parameters.178- Agent Reward Function (independent):179  The reward function is now geometric meaning the reward each step is a product180  of all the rewards instead of a sum, this helps the agent try to maximize all181  rewards instead of the easiest rewards.182  - Body velocity matches goal velocity. (normalized between (0,1))183  - Head direction alignment with goal direction. (normalized between (0,1))184- Behavior Parameters:185  - Vector Observation space: 172 variables corresponding to position, rotation,186    velocity, and angular velocities of each limb plus the acceleration and187    angular acceleration of the body.188  - Actions: 20 continuous actions, corresponding to target189    rotations for joints.190  - Visual Observations: None191- Float Properties: None192- Benchmark Mean Reward: 3000193 194## Worm195 196![Worm](images/worm.png)197 198- Set-up: A worm with a head and 3 body segments.199- Goal: The agents must move its body toward the goal direction.200- Agents: The environment contains 10 agents with same Behavior Parameters.201- Agent Reward Function (independent):202  The reward function is now geometric meaning the reward each step is a product203  of all the rewards instead of a sum, this helps the agent try to maximize all204  rewards instead of the easiest rewards.205  - Body velocity matches goal velocity. (normalized between (0,1))206  - Body direction alignment with goal direction. (normalized between (0,1))207- Behavior Parameters:208  - Vector Observation space: 64 variables corresponding to position, rotation,209    velocity, and angular velocities of each limb plus the acceleration and210    angular acceleration of the body.211  - Actions: 9 continuous actions, corresponding to target212    rotations for joints.213  - Visual Observations: None214- Float Properties: None215- Benchmark Mean Reward: 800216 217## Food Collector218 219![Collector](images/foodCollector.png)220 221- Set-up: A multi-agent environment where agents compete to collect food.222- Goal: The agents must learn to collect as many green food spheres as possible223  while avoiding red spheres.224- Agents: The environment contains 5 agents with same Behavior Parameters.225- Agent Reward Function (independent):226  - +1 for interaction with green spheres227  - -1 for interaction with red spheres228- Behavior Parameters:229  - Vector Observation space: 53 corresponding to velocity of agent (2), whether230    agent is frozen and/or shot its laser (2), plus grid based perception of231    objects around agent's forward direction (40 by 40 with 6 different categories).232  - Actions:233    - 3 continuous actions correspond to Forward Motion, Side Motion and Rotation234    - 1 discrete acion branch for Laser with 2 possible actions corresponding to235      Shoot Laser or No Action236  - Visual Observations (Optional): First-person camera per-agent, plus one vector237    flag representing the frozen state of the agent. This scene uses a combination238    of vector and visual observations and the training will not succeed without239    the frozen vector flag. Use `VisualFoodCollector` scene.240- Float Properties: Two241  - laser_length: Length of the laser used by the agent242    - Default: 1243    - Recommended Minimum: 0.2244    - Recommended Maximum: 7245  - agent_scale: Specifies the scale of the agent in the 3 dimensions (equal246    across the three dimensions)247    - Default: 1248    - Recommended Minimum: 0.5249    - Recommended Maximum: 5250- Benchmark Mean Reward: 10251 252## Hallway253 254![Hallway](images/hallway.png)255 256- Set-up: Environment where the agent needs to find information in a room,257  remember it, and use it to move to the correct goal.258- Goal: Move to the goal which corresponds to the color of the block in the259  room.260- Agents: The environment contains one agent.261- Agent Reward Function (independent):262  - +1 For moving to correct goal.263  - -0.1 For moving to incorrect goal.264  - -0.0003 Existential penalty.265- Behavior Parameters:266  - Vector Observation space: 30 corresponding to local ray-casts detecting267    objects, goals, and walls.268  - Actions: 1 discrete action Branch, with 4 actions corresponding to agent269    rotation and forward/backward movement.270- Float Properties: None271- Benchmark Mean Reward: 0.7272  - To train this environment, you can enable curiosity by adding the `curiosity` reward signal273    in `config/ppo/Hallway.yaml`274 275## Soccer Twos276 277![SoccerTwos](images/soccer.png)278 279- Set-up: Environment where four agents compete in a 2 vs 2 toy soccer game.280- Goal:281  - Get the ball into the opponent's goal while preventing the ball from282    entering own goal.283- Agents: The environment contains two different Multi Agent Groups with two agents in each.284  Parameters : SoccerTwos.285- Agent Reward Function (dependent):286  - (1 - `accumulated time penalty`) When ball enters opponent's goal287    `accumulated time penalty` is incremented by (1 / `MaxStep`) every fixed288    update and is reset to 0 at the beginning of an episode.289  - -1 When ball enters team's goal.290- Behavior Parameters:291  - Vector Observation space: 336 corresponding to 11 ray-casts forward292    distributed over 120 degrees and 3 ray-casts backward distributed over 90293    degrees each detecting 6 possible object types, along with the object's294    distance. The forward ray-casts contribute 264 state dimensions and backward295    72 state dimensions over three observation stacks.296  - Actions: 3 discrete branched actions corresponding to297    forward, backward, sideways movement, as well as rotation.298  - Visual Observations: None299- Float Properties: Two300  - ball_scale: Specifies the scale of the ball in the 3 dimensions (equal301    across the three dimensions)302    - Default: 7.5303    - Recommended minimum: 4304    - Recommended maximum: 10305  - gravity: Magnitude of the gravity306    - Default: 9.81307    - Recommended minimum: 6308    - Recommended maximum: 20309 310## Strikers Vs. Goalie311 312![StrikersVsGoalie](images/strikersvsgoalie.png)313 314- Set-up: Environment where two agents compete in a 2 vs 1 soccer variant.315- Goal:316  - Striker: Get the ball into the opponent's goal.317  - Goalie: Keep the ball out of the goal.318- Agents: The environment contains two different Multi Agent Groups. One with two Strikers and the other one Goalie.319  Behavior Parameters : Striker, Goalie.320- Striker Agent Reward Function (dependent):321  - +1 When ball enters opponent's goal.322  - -0.001 Existential penalty.323- Goalie Agent Reward Function (dependent):324  - -1 When ball enters goal.325  - 0.001 Existential bonus.326- Behavior Parameters:327  - Striker Vector Observation space: 294 corresponding to 11 ray-casts forward328    distributed over 120 degrees and 3 ray-casts backward distributed over 90329    degrees each detecting 5 possible object types, along with the object's330    distance. The forward ray-casts contribute 231 state dimensions and backward331    63 state dimensions over three observation stacks.332  - Striker Actions: 3 discrete branched actions corresponding333    to forward, backward, sideways movement, as well as rotation.334  - Goalie Vector Observation space: 738 corresponding to 41 ray-casts335    distributed over 360 degrees each detecting 4 possible object types, along336    with the object's distance and 3 observation stacks.337  - Goalie Actions: 3 discrete branched actions corresponding338    to forward, backward, sideways movement, as well as rotation.339  - Visual Observations: None340- Float Properties: Two341  - ball_scale: Specifies the scale of the ball in the 3 dimensions (equal342    across the three dimensions)343    - Default: 7.5344    - Recommended minimum: 4345    - Recommended maximum: 10346  - gravity: Magnitude of the gravity347    - Default: 9.81348    - Recommended minimum: 6349    - Recommended maximum: 20350 351## Walker352 353![Walker](images/walker.png)354 355- Set-up: Physics-based Humanoid agents with 26 degrees of freedom. These DOFs356  correspond to articulation of the following body-parts: hips, chest, spine,357  head, thighs, shins, feet, arms, forearms and hands.358- Goal: The agents must move its body toward the goal direction without falling.359- Agents: The environment contains 10 independent agents with same Behavior360  Parameters.361- Agent Reward Function (independent):362  The reward function is now geometric meaning the reward each step is a product363  of all the rewards instead of a sum, this helps the agent try to maximize all364  rewards instead of the easiest rewards.365  - Body velocity matches goal velocity. (normalized between (0,1))366  - Head direction alignment with goal direction. (normalized between (0,1))367- Behavior Parameters:368  - Vector Observation space: 243 variables corresponding to position, rotation,369    velocity, and angular velocities of each limb, along with goal direction.370  - Actions: 39 continuous actions, corresponding to target371    rotations and strength applicable to the joints.372  - Visual Observations: None373- Float Properties: Four374  - gravity: Magnitude of gravity375    - Default: 9.81376    - Recommended Minimum:377    - Recommended Maximum:378  - hip_mass: Mass of the hip component of the walker379    - Default: 8380    - Recommended Minimum: 7381    - Recommended Maximum: 28382  - chest_mass: Mass of the chest component of the walker383    - Default: 8384    - Recommended Minimum: 3385    - Recommended Maximum: 20386  - spine_mass: Mass of the spine component of the walker387    - Default: 8388    - Recommended Minimum: 3389    - Recommended Maximum: 20390- Benchmark Mean Reward : 2500391 392 393## Pyramids394 395![Pyramids](images/pyramids.png)396 397- Set-up: Environment where the agent needs to press a button to spawn a398  pyramid, then navigate to the pyramid, knock it over, and move to the gold399  brick at the top.400- Goal: Move to the golden brick on top of the spawned pyramid.401- Agents: The environment contains one agent.402- Agent Reward Function (independent):403  - +2 For moving to golden brick (minus 0.001 per step).404- Behavior Parameters:405  - Vector Observation space: 148 corresponding to local ray-casts detecting406    switch, bricks, golden brick, and walls, plus variable indicating switch407    state.408  - Actions: 1 discrete action branch, with 4 actions corresponding to agent rotation and409    forward/backward movement.410- Float Properties: None411- Benchmark Mean Reward: 1.75412 413## Match 3414![Match 3](images/match3.png)415 416- Set-up: Simple match-3 game. Matched pieces are removed, and remaining pieces417drop down. New pieces are spawned randomly at the top, with a chance of being418"special".419- Goal: Maximize score from matching pieces.420- Agents: The environment contains several independent Agents.421- Agent Reward Function (independent):422  - .01 for each normal piece cleared. Special pieces are worth 2x or 3x.423- Behavior Parameters:424  - None425  - Observations and actions are defined with a sensor and actuator respectively.426- Float Properties: None427- Benchmark Mean Reward:428  - 39.5 for visual observations429  - 38.5 for vector observations430  - 34.2 for simple heuristic (pick a random valid move)431  - 37.0 for greedy heuristic (pick the highest-scoring valid move)432 433## Sorter434![Sorter](images/sorter.png)435 436 - Set-up: The Agent is in a circular room with numbered tiles. The values of the437 tiles are random between 1 and 20. The tiles present in the room are randomized438 at each episode. When the Agent visits a tile, it turns green.439 - Goal: Visit all the tiles in ascending order.440 - Agents: The environment contains a single Agent441 - Agent Reward Function:442  - -.0002 Existential penalty.443  - +1 For visiting the right tile444  - -1 For visiting the wrong tile445 - BehaviorParameters:446  - Vector Observations : 4 : 2 floats for Position and 2 floats for orientation447  - Variable Length Observations : Between 1 and 20 entities (one for each tile)448  each with 22 observations, the first 20 are one hot encoding of the value of the tile,449  the 21st and 22nd represent the position of the tile relative to the Agent and the 23rd450  is `1` if the tile was visited and `0` otherwise.451  - Actions: 3 discrete branched actions corresponding to forward, backward,452  sideways movement, as well as rotation.453  - Float Properties: One454    - num_tiles: The maximum number of tiles to sample.455      - Default: 2456      - Recommended Minimum: 1457      - Recommended Maximum: 20458  - Benchmark Mean Reward: Depends on the number of tiles.459 460## Cooperative Push Block461![CoopPushBlock](images/cooperative_pushblock.png)462 463- Set-up: Similar to Push Block, the agents are in an area with blocks that need464to be pushed into a goal. Small blocks can be pushed by one agents and are worth465+1 value, medium blocks require two agents to push in and are worth +2, and large466blocks require all 3 agents to push and are worth +3.467- Goal: Push all blocks into the goal.468- Agents: The environment contains three Agents in a Multi Agent Group.469- Agent Reward Function:470  - -0.0001 Existential penalty, as a group reward.471  - +1, +2, or +3 for pushing in a block, added as a group reward.472- Behavior Parameters:473  - Observation space: A single Grid Sensor with separate tags for each block size,474    the goal, the walls, and other agents.475  - Actions: 1 discrete action branch with 7 actions, corresponding to turn clockwise476    and counterclockwise, move along four different face directions, or do nothing.477- Float Properties: None478- Benchmark Mean Reward: 11 (Group Reward)479 480## Dungeon Escape481![DungeonEscape](images/dungeon_escape.png)482 483- Set-up: Agents are trapped in a dungeon with a dragon, and must work together to escape.484  To retrieve the key, one of the agents must find and slay the dragon, sacrificing itself485  to do so. The dragon will drop a key for the others to use. The other agents can then pick486  up this key and unlock the dungeon door. If the agents take too long, the dragon will escape487  through a portal and the environment resets.488- Goal: Unlock the dungeon door and leave.489- Agents: The environment contains three Agents in a Multi Agent Group and one Dragon, which490  moves in a predetermined pattern.491- Agent Reward Function:492  - +1 group reward if any agent successfully unlocks the door and leaves the dungeon.493- Behavior Parameters:494  - Observation space: A Ray Perception Sensor with separate tags for the walls, other agents,495    the door, key, the dragon, and the dragon's portal. A single Vector Observation which indicates496    whether the agent is holding a key.497  - Actions: 1 discrete action branch with 7 actions, corresponding to turn clockwise498    and counterclockwise, move along four different face directions, or do nothing.499- Float Properties: None500- Benchmark Mean Reward: 1.0 (Group Reward)501