Team Ai
Modelpublic

AnnaMats/ppo-Pyramids-Training

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes110downloads
Background-Machine-Learning.md196 linesDownload Raw Back to docs
1# Background: Machine Learning2 3Given that a number of users of the ML-Agents Toolkit might not have a formal4machine learning background, this page provides an overview to facilitate the5understanding of the ML-Agents Toolkit. However, we will not attempt to provide6a thorough treatment of machine learning as there are fantastic resources7online.8 9Machine learning, a branch of artificial intelligence, focuses on learning10patterns from data. The three main classes of machine learning algorithms11include: unsupervised learning, supervised learning and reinforcement learning.12Each class of algorithm learns from a different type of data. The following13paragraphs provide an overview for each of these classes of machine learning, as14well as introductory examples.15 16## Unsupervised Learning17 18The goal of19[unsupervised learning](https://en.wikipedia.org/wiki/Unsupervised_learning) is20to group or cluster similar items in a data set. For example, consider the21players of a game. We may want to group the players depending on how engaged22they are with the game. This would enable us to target different groups (e.g.23for highly-engaged players we might invite them to be beta testers for new24features, while for unengaged players we might email them helpful tutorials).25Say that we wish to split our players into two groups. We would first define26basic attributes of the players, such as the number of hours played, total money27spent on in-app purchases and number of levels completed. We can then feed this28data set (three attributes for every player) to an unsupervised learning29algorithm where we specify the number of groups to be two. The algorithm would30then split the data set of players into two groups where the players within each31group would be similar to each other. Given the attributes we used to describe32each player, in this case, the output would be a split of all the players into33two groups, where one group would semantically represent the engaged players and34the second group would semantically represent the unengaged players.35 36With unsupervised learning, we did not provide specific examples of which37players are considered engaged and which are considered unengaged. We just38defined the appropriate attributes and relied on the algorithm to uncover the39two groups on its own. This type of data set is typically called an unlabeled40data set as it is lacking these direct labels. Consequently, unsupervised41learning can be helpful in situations where these labels can be expensive or42hard to produce. In the next paragraph, we overview supervised learning43algorithms which accept input labels in addition to attributes.44 45## Supervised Learning46 47In [supervised learning](https://en.wikipedia.org/wiki/Supervised_learning), we48do not want to just group similar items but directly learn a mapping from each49item to the group (or class) that it belongs to. Returning to our earlier50example of clustering players, let's say we now wish to predict which of our51players are about to churn (that is stop playing the game for the next 30 days).52We can look into our historical records and create a data set that contains53attributes of our players in addition to a label indicating whether they have54churned or not. Note that the player attributes we use for this churn prediction55task may be different from the ones we used for our earlier clustering task. We56can then feed this data set (attributes **and** label for each player) into a57supervised learning algorithm which would learn a mapping from the player58attributes to a label indicating whether that player will churn or not. The59intuition is that the supervised learning algorithm will learn which values of60these attributes typically correspond to players who have churned and not61churned (for example, it may learn that players who spend very little and play62for very short periods will most likely churn). Now given this learned model, we63can provide it the attributes of a new player (one that recently started playing64the game) and it would output a _predicted_ label for that player. This65prediction is the algorithms expectation of whether the player will churn or66not. We can now use these predictions to target the players who are expected to67churn and entice them to continue playing the game.68 69As you may have noticed, for both supervised and unsupervised learning, there70are two tasks that need to be performed: attribute selection and model71selection. Attribute selection (also called feature selection) pertains to72selecting how we wish to represent the entity of interest, in this case, the73player. Model selection, on the other hand, pertains to selecting the algorithm74(and its parameters) that perform the task well. Both of these tasks are active75areas of machine learning research and, in practice, require several iterations76to achieve good performance.77 78We now switch to reinforcement learning, the third class of machine learning79algorithms, and arguably the one most relevant for the ML-Agents Toolkit.80 81## Reinforcement Learning82 83[Reinforcement learning](https://en.wikipedia.org/wiki/Reinforcement_learning)84can be viewed as a form of learning for sequential decision making that is85commonly associated with controlling robots (but is, in fact, much more86general). Consider an autonomous firefighting robot that is tasked with87navigating into an area, finding the fire and neutralizing it. At any given88moment, the robot perceives the environment through its sensors (e.g. camera,89heat, touch), processes this information and produces an action (e.g. move to90the left, rotate the water hose, turn on the water). In other words, it is91continuously making decisions about how to interact in this environment given92its view of the world (i.e. sensors input) and objective (i.e. neutralizing the93fire). Teaching a robot to be a successful firefighting machine is precisely94what reinforcement learning is designed to do.95 96More specifically, the goal of reinforcement learning is to learn a **policy**,97which is essentially a mapping from **observations** to **actions**. An98observation is what the robot can measure from its **environment** (in this99case, all its sensory inputs) and an action, in its most raw form, is a change100to the configuration of the robot (e.g. position of its base, position of its101water hose and whether the hose is on or off).102 103The last remaining piece of the reinforcement learning task is the **reward104signal**. The robot is trained to learn a policy that maximizes its overall rewards. When training a robot to be a mean firefighting machine, we provide it105with rewards (positive and negative) indicating how well it is doing on106completing the task. Note that the robot does not _know_ how to put out fires107before it is trained. It learns the objective because it receives a large108positive reward when it puts out the fire and a small negative reward for every109passing second. The fact that rewards are sparse (i.e. may not be provided at110every step, but only when a robot arrives at a success or failure situation), is111a defining characteristic of reinforcement learning and precisely why learning112good policies can be difficult (and/or time-consuming) for complex environments.113 114<div style="text-align: center"><img src="../images/rl_cycle.png" alt="The reinforcement learning lifecycle."></div>115 116[Learning a policy](https://blogs.unity3d.com/2017/08/22/unity-ai-reinforcement-learning-with-q-learning/)117usually requires many trials and iterative policy updates. More specifically,118the robot is placed in several fire situations and over time learns an optimal119policy which allows it to put out fires more effectively. Obviously, we cannot120expect to train a robot repeatedly in the real world, particularly when fires121are involved. This is precisely why the use of122[Unity as a simulator](https://blogs.unity3d.com/2018/01/23/designing-safer-cities-through-simulations/)123serves as the perfect training grounds for learning such behaviors. While our124discussion of reinforcement learning has centered around robots, there are125strong parallels between robots and characters in a game. In fact, in many ways,126one can view a non-playable character (NPC) as a virtual robot, with its own127observations about the environment, its own set of actions and a specific128objective. Thus it is natural to explore how we can train behaviors within Unity129using reinforcement learning. This is precisely what the ML-Agents Toolkit130offers. The video linked below includes a reinforcement learning demo showcasing131training character behaviors using the ML-Agents Toolkit.132 133<p align="center">134  <a href="http://www.youtube.com/watch?feature=player_embedded&v=fiQsmdwEGT8" target="_blank">135    <img src="http://img.youtube.com/vi/fiQsmdwEGT8/0.jpg" alt="RL Demo" width="400" border="10" />136  </a>137</p>138 139Similar to both unsupervised and supervised learning, reinforcement learning140also involves two tasks: attribute selection and model selection. Attribute141selection is defining the set of observations for the robot that best help it142complete its objective, while model selection is defining the form of the policy143(mapping from observations to actions) and its parameters. In practice, training144behaviors is an iterative process that may require changing the attribute and145model choices.146 147## Training and Inference148 149One common aspect of all three branches of machine learning is that they all150involve a **training phase** and an **inference phase**. While the details of151the training and inference phases are different for each of the three, at a152high-level, the training phase involves building a model using the provided153data, while the inference phase involves applying this model to new, previously154unseen, data. More specifically:155 156- For our unsupervised learning example, the training phase learns the optimal157  two clusters based on the data describing existing players, while the158  inference phase assigns a new player to one of these two clusters.159- For our supervised learning example, the training phase learns the mapping160  from player attributes to player label (whether they churned or not), and the161  inference phase predicts whether a new player will churn or not based on that162  learned mapping.163- For our reinforcement learning example, the training phase learns the optimal164  policy through guided trials, and in the inference phase, the agent observes165  and takes actions in the wild using its learned policy.166 167To briefly summarize: all three classes of algorithms involve training and168inference phases in addition to attribute and model selections. What ultimately169separates them is the type of data available to learn from. In unsupervised170learning our data set was a collection of attributes, in supervised learning our171data set was a collection of attribute-label pairs, and, lastly, in172reinforcement learning our data set was a collection of173observation-action-reward tuples.174 175## Deep Learning176 177[Deep learning](https://en.wikipedia.org/wiki/Deep_learning) is a family of178algorithms that can be used to address any of the problems introduced above.179More specifically, they can be used to solve both attribute and model selection180tasks. Deep learning has gained popularity in recent years due to its181outstanding performance on several challenging machine learning tasks. One182example is [AlphaGo](https://en.wikipedia.org/wiki/AlphaGo), a183[computer Go](https://en.wikipedia.org/wiki/Computer_Go) program, that leverages184deep learning, that was able to beat Lee Sedol (a Go world champion).185 186A key characteristic of deep learning algorithms is their ability to learn very187complex functions from large amounts of training data. This makes them a natural188choice for reinforcement learning tasks when a large amount of data can be189generated, say through the use of a simulator or engine such as Unity. By190generating hundreds of thousands of simulations of the environment within Unity,191we can learn policies for very complex environments (a complex environment is192one where the number of observations an agent perceives and the number of193actions they can take are large). Many of the algorithms we provide in ML-Agents194use some form of deep learning, built on top of the open-source library,195[PyTorch](Background-PyTorch.md).196