Prithvi-Raj-Dixit/Large_Language_Model_LLMs
0
11. How AI learns?2 3AI models are tained on massive data(internet,books,wikipedia etc.).4Based on this data machine recognize some patterns. Machine recognize this pattern using a lot of math5that happens behind the scene. 6Based on this pattern a model is created- ML model7 8Ex- We fed machine thousand of smiling photos. Based on the photos machine identify the patterns(smile face etc.)9Based on that patterns machine created a model. This ML model we can think of as a mathematical function(y=f(x))10which does future predictions for us.11So after that if we give any picture to model on which it is not trained, it will also able to predict it.12 132. What is AI?14AI is a system which performs tasks that require human intelligence(Image recognition, Human translation etc.)15 163. What is ML?17ML is a subset of AI which works on data. ML models are trained on lot of data.18 194. What is DL?20DL is a subset of ML which deals with neural networks.21In DL we try to mimic human brain using neural networks.22 235. Neural Networks24Neural network consists of- neurons(circular circle), connections(lines connecting neurons)25First layer is called Input layer, Last layer is called Output layer and in between those there are many hidden layers.26Our data propogate from left to right side.27 28Initially we get some data as input(parameters of data). We feed the data to neural network and at the end29neural network do some predictions.30 31When we feed data to neural network and it goes layer after layer and at end we get predictions- Forward Propagation32If predction is wrong we calculate how much is difference- Loss value33We feed this loss to our neural network in opposite direction to adjust weights- Backward Propagation34 35Each neuron has a bias value(b) also along with weights(w) and input data(x)36Neuron value- x1w1 + x2w2 + x3w3 + b37Neuron output- f(x1w1 + x2w2 + x3w3 + b) (We apply Activation function on summation of neuron values)38 396. What are LLMs?40Large Language model are those DL models or DL neural network which are trained on large amount of textual data.41Large- almost entire internet, wikipedia, books etc.42 43LLMs predict the next token based on input prompt from user based on lot of math and pattern.44 457. Tokens/Tokenization-46Machine does not understand human lamguage. It understand numbers.47Its very hard to convert entire text in numbers,so LLMs break the text into small parts called tokens(3/4th of English word).48Each token has an ID also49 50Ex- I like Generative AI 51tokens=5, characters=2052token ID= [40, 1299, 4140, 1799, 20837]53 54I like Machine Learning55tokens=4, characters=2356token ID= [40, 1299, 19121, 25392]57 588. Context Window59User input prompt + LLMs generated response + All of previous message of the conversation + uploading pdf, images etc60These things fill up context window61gpt 6= ~1M token62fable= ~1M token63 64Context window is not permanent. It is only for current active conversation.65Larger context window does not mean LLM will perform good.66 679. Vector/Vector Embeddings68From token and token ID LLMs does not understand the meaning of the word.69It understands meaning of the word by using vectors(1D array of different numbers)70Cat vector=[0.02 , 0.3 , 0.5]71Vectors of related word are near to each other72Ex- Vector of dog and pet are more near than lion and pet73 7410. Attention75I code in python76A python is going77 78Both has python word so if we convert these words to vector embeddings and feed to LLM that is not enough79to understand the meaning of the word because both words carry different meaning.80So LLMs use a special mechanism to do that. It checks in data for a token how much other tokens in that sentence81in related to that token.Ex- I code in first sentence and is going in second82 83To tackle this LLMs create a Query vector of the word and key vector for other tokens in a sentence.84I love python 85python- Query vector86I love- Key vector87 88It then perform vector operations using query and key(like dot product etc.) and calculate Attention score.89Ex- I code in python90code and python has high attention score as compared to other tokens in this sentence91This helps to understand meaning of python more accurately.92Attention score is calculated of every token with every other token and itself also.93 94This attention mechanism is introduced in research paper- Attention is all you need (2017)95 9611. Transformers97Transformers architecture is based on attention mechanism98 9912. Self supervised learning100LLMs on their training data do self supervised learning.101They hide some parts of data from token and try to predict it and then it will see how close its prediction is to 102the target value.103 10413. RAG(Retrieval Augmented Generation)105Allow LLM to get information of company specific documents, information.106So a company created a RAG pipeline to allow LLM to generate better output based on company specific documents.107The data access we give to LLM for RAG can be given in the form of- Graph db, SQL db, Vector db(manjority RAG pipeline we use)108 109Adding documents in context video is different than RAG pipleline110 11114. Vector database112We do not directly store company specific documents in vector database.113First documents are divided into chunks, then these chunks are converted into vector embeddings114 11515. MCP(Model Context Protocol)116What is LLM needs to access another website like Jira, github repository to fetch details from that website.117So LLM has a MCP client will connect to MCP server of github from where it will fetch relevant information and 118this will pass along with out query.119 12016. AI Agent121LLM does not perform any operation like- pushin code on github etc.122AI Agents help to execute tasks for us123 124 125 