Team Ai
Apppublic

cj-dev-code/semantic_search

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes
journey.md34 linesDownload Raw Back to docs
1- created conda env for project2- installed PyMuPDF for book conversation to .txt3- implemented script for conversion to .txt4- converted book to .txt5- explored semantic chunking strategies. 6  - 500 max: chunks usually if not exclusively seem 500 chars long, disrespecting sentence boundaries.7  - (100, 1000): usually about 100 chars long. again, disrespecting idea and sentence boundaries.8  - 200 max: continues to violate sentence boundaries.9  - (10, 500): statements seem too small and still violate sentence boundaries.10  - (100, 500): better, but sentence boundaries unpreserved.11  conclusion: conventional sentence boundary preserving text splitting module may work better than semantic-text-splitter12  options: langchain recursive text splitter or nltk.tokenize13  - LC recursive text splitter: helpful for paragraph level retrieval. lots of context. struggles for quote level extraction.14  - nltk: helpful for quote level retrieval. could struggle to push mvp out.15  - LC SemanticChunker: could also work. for MVP, let's do paragraphs and scope this for later.16  decision: langchain. conventional. paragraph level. yields the quote inside anyway. enduser can select from output.17- explored recursive text splitter18  - separators=["\n\n", "\n", ".", " "], chunk_size=1000, chunk_overlap=100: chunks sometimes contain too much info. two paras... eh.. not a fan. trying smaller chunk size.19  - separators=["\n\n", "\n", ".", " "], chunk_size=500, chunk_overlap=100: better. same issue. trying to make strict on paragraphs next20  - separators=["\n\n"], chunk_size=500, chunk_overlap=100: this did sections. trying paragraphs21  - separators=["\n"], chunk_size=500, chunk_overlap=100: this did paragraphs. improved cogency. probably higher max chunk size for big paragraphs.22  - separators=["\n"], chunk_size=1000, chunk_overlap=100: this did paragraphs. too big though. 23  conclusion: probably worth exploring the langchain semantic approach 24- explored semantic splitter (langchain, different embeddings):25  - with openai embeddings: it's okay. some chunks are really long. but the're coherent. it may be important to do a search inside one of these chunks before outputting to enduser. 26  - with openai embeddings stddev = 1.5: still long chunks27  - with openai embeddings stddev = 6.0: nonexistent chunks28  - with openai embeddings stddev = .5: reasonably good chunks29  - with bge-base-en-v1.5 (huggingface) stddev = .5: comparable to openai, but 30  - with bge-base-en-v1.5 (huggingface) stddev = 1.5: unusably long31  - with bge-base-en-v1.5 (huggingface) stddev = .25: unusably long32 33  - with anthropic embeddings: 34