awacke1/Transcript-AI-Learner-From-Youtube
2
1https://www.youtube.com/watch?v=9EN_HoEk3KY&t=172s2 3 41:425program the does very very well on your data then you will achieve the best61:487generalization possible with a little bit of modification you can turn it into a precise theorem81:549and on a very intuitive level it's easy to see what it should be the case if you102:0111have some data and you're able to find a shorter program which generates this122:0613data then you've essentially extracted all the all conceivable regularity from142:1115this data into your program and then you can use these objects to make the best predictions possible like if if you have162:1917data which is so complex but there is no way to express it as a shorter program182:2519then it means that your data is totally random there is no way to extract any regularity from it whatsoever now there202:3221is little known mathematical theory behind this and the proofs of these statements actually not even that hard222:3823but the one minor slight disappointment is that it's actually not possible at242:4425least given today's tools and understanding to find the best short program that 26 27 28 29https://youtu.be/9EN_HoEk3KY?t=44230531to talk a little bit about reinforcement learning so reinforcement learning is a framework it's a framework of evaluating326:5333agents in their ability to achieve goals and complicated stochastic environments346:5835you've got an agent which is plugged into an environment as shown in the figure right here and for any given367:0637agent you can simply run it many times and compute its average reward now the387:1339thing that's interesting about the reinforcement learning framework is that there exist interesting useful407:2041reinforcement learning algorithms the framework existed for a long time it427:2543became interesting once we realized that good algorithms exist now these are there are perfect algorithms but they447:3145are good enough to do interesting things and all you want the mathematical467:3747problem is one where you need to maximize the expected reward now one487:4449important way in which the reinforcement learning framework is not quite complete is that it assumes that the reward is507:5051given by the environment you see this picture the agent sends an action while527:5653the reward sends it an observation in a both the observation and the reward backwards that's what the environment548:0155communicates back the way in which this is not the case in the real world is that we figure out568:1157what the reward is from the observation we reward ourselves we are not told588:1659environment doesn't say hey here's some negative reward it's our interpretation over census that lets us determine what608:2361the reward is and there is only one real true reward in life and this is628:2863existence or nonexistence and everything else is a corollary of that so well what648:3565should our agent be you already know the answer should be a neural network because whenever you want to do668:4167something dense it's going to be a neural network and you want the agent to map observations to actions so you let688:4769it be parametrized with a neural net and you apply learning algorithm so I want to explain to you how reinforcement708:5371learning works this is model free reinforcement learning the reinforcement learning has actually been used in practice everywhere but it's