cj-dev-code/semantic_search
0
1- created conda env for project2- installed PyMuPDF for book conversation to .txt3- implemented script for conversion to .txt4- converted book to .txt5- explored semantic chunking strategies. 6 - 500 max: chunks usually if not exclusively seem 500 chars long, disrespecting sentence boundaries.7 - (100, 1000): usually about 100 chars long. again, disrespecting idea and sentence boundaries.8 - 200 max: continues to violate sentence boundaries.9 - (10, 500): statements seem too small and still violate sentence boundaries.10 - (100, 500): better, but sentence boundaries unpreserved.11 conclusion: conventional sentence boundary preserving text splitting module may work better than semantic-text-splitter12 options: langchain recursive text splitter or nltk.tokenize13 - LC recursive text splitter: helpful for paragraph level retrieval. lots of context. struggles for quote level extraction.14 - nltk: helpful for quote level retrieval. could struggle to push mvp out.15 - LC SemanticChunker: could also work. for MVP, let's do paragraphs and scope this for later.16 decision: langchain. conventional. paragraph level. yields the quote inside anyway. enduser can select from output.17- explored recursive text splitter18 - separators=["\n\n", "\n", ".", " "], chunk_size=1000, chunk_overlap=100: chunks sometimes contain too much info. two paras... eh.. not a fan. trying smaller chunk size.19 - separators=["\n\n", "\n", ".", " "], chunk_size=500, chunk_overlap=100: better. same issue. trying to make strict on paragraphs next20 - separators=["\n\n"], chunk_size=500, chunk_overlap=100: this did sections. trying paragraphs21 - separators=["\n"], chunk_size=500, chunk_overlap=100: this did paragraphs. improved cogency. probably higher max chunk size for big paragraphs.22 - separators=["\n"], chunk_size=1000, chunk_overlap=100: this did paragraphs. too big though. 23 conclusion: probably worth exploring the langchain semantic approach 24- explored semantic splitter (langchain, different embeddings):25 - with openai embeddings: it's okay. some chunks are really long. but the're coherent. it may be important to do a search inside one of these chunks before outputting to enduser. 26 - with openai embeddings stddev = 1.5: still long chunks27 - with openai embeddings stddev = 6.0: nonexistent chunks28 - with openai embeddings stddev = .5: reasonably good chunks29 - with bge-base-en-v1.5 (huggingface) stddev = .5: comparable to openai, but 30 - with bge-base-en-v1.5 (huggingface) stddev = 1.5: unusably long31 - with bge-base-en-v1.5 (huggingface) stddev = .25: unusably long32 33 - with anthropic embeddings: 34 