Team Ai
Datasetpublic

Angshul/SparseGeometricRAG

SparseGeometricRAG CPU-first sparse geometric retrieval for practical top-10 RAG No transformer inference at retrieval time. No retrieval GPU requirement. No dense document-vector dot products. No external API. SparseGeometricRAG is a retrieval system built around one systems objective: make the retrieval layer cheap enough to run on ordinary multicore CPU hardware without turning the corpus into a dense embedding database. It uses sparse TF-IDF geometry, fuzzy… See the full description on the dataset page: https://huggingface.co/datasets/Angshul/SparseGeometricRAG.

sourceHugging Facefair-noncommercial-research-licenseupdated 2mo agoView on Hugging Face
0likes331downloads
README.md434 linesDownload Raw Back to root
1---2license: fair-noncommercial-research-license3language:4- en5pretty_name: SparseGeometricRAG6tags:7- information-retrieval8- retrieval-augmented-generation9- rag10- sparse-retrieval11- cpu12- low-latency13- low-resource14- beir15- msmarco16---17 18# SparseGeometricRAG19 20## CPU-first sparse geometric retrieval for practical top-10 RAG21 22**No transformer inference at retrieval time. No retrieval GPU requirement. No dense document-vector dot products. No external API.**23 24SparseGeometricRAG is a retrieval system built around one systems objective: **make the retrieval layer cheap enough to run on ordinary multicore CPU hardware without turning the corpus into a dense embedding database.** It uses sparse TF-IDF geometry, fuzzy branch localization, and a tiny signed local residual code. The richer chunk-level evidence is delayed until after routing and shortlist reduction.25 26The project is not positioned as an accuracy-at-any-cost replacement for the strongest neural retrievers. Its selling point is the **quality / latency / hardware tradeoff**: useful top-10 retrieval with small structured state, bounded local computation, no retrieval-time transformer stack, and no requirement for a GPU or hosted inference service.27 28### At a glance29 30| Property | Frozen design |31|---|---|32| Query representation | sparse TF-IDF |33| Fuzzy memberships per chunk | `F = 4` |34| Sparse branch-center support | `B = 64` coordinates |35| Signed residual support | `S = 16` coordinates per membership |36| Weak routing expansion | bounded sparse neighborhood |37| Large-route shortlist | `P = 100` for the frozen six-dataset row |38| Final RAG output | top 10 chunks |39| Retrieval-time transformer | **none** |40| Retrieval-time GPU | **not required** |41| Dense vector per document | **not required** |42 43---44 45## 1. Why this design exists46 47Most modern retrieval systems optimize a learned representation and then optimize the search engine around that representation. SparseGeometricRAG changes the question: **can the representation itself be made sufficiently small and local that the retrieval engine no longer needs heavyweight dense-vector machinery?**48 49<p align="center">50  <img src="figures/fig01_positioning.png" width="900" alt="SparseGeometricRAG positioning against dense and learned sparse retrieval">51</p>52<p align="center"><em>Figure 1. SparseGeometricRAG changes the cost structure of retrieval. Dense and learned-sparse stacks retain a neural representation stage; the proposed stack remains sparse and CPU-native at retrieval time.</em></p>53 54The key design choice is to store a **coarse sparse location plus a tiny local directional code**, rather than a dense vector for every chunk. The query remains sparse and real-valued, so it supplies fine amplitude information at runtime while the database stores only coarse branch position and signed local deviations.55 56This gives three practical consequences:57 581. the stored geometric state per chunk is controlled by small fixed handles;592. the decisive local comparison is bounded by only **16 residual coordinates per routed membership**; and603. detailed lexical, semantic-support, and diversity calculations are postponed until the candidate set has already collapsed.61 62---63 64## 2. Architecture65 66### 2.1 Offline indexing67 68<p align="center">69  <img src="figures/fig02_offline_indexing.png" width="900" alt="SparseGeometricRAG offline indexing architecture">70</p>71<p align="center"><em>Figure 2. Offline indexing converts each chunk into sparse lexical support, four fuzzy branch memberships, and a 16-sign local residual code. Branch centers and sparse term graphs are shared structures.</em></p>72 73The index is constructed from sparse normalized TF-IDF. A chunk is assigned to its strongest fuzzy branches, each branch is represented by a sparse center, and the chunk's deviation from that center is compressed to a small signed residual code. A bounded sparse term graph provides weak second-order routing support. The result is a compact index consisting of branch postings, membership weights, local sign codes, binary term support, and shared sparse structures.74 75The frozen structural handles are:76 77| Handle | Frozen value | Role |78|---|---:|---|79| `F` | 4 | fuzzy branch memberships per chunk |80| `B` | 64 | sparse coordinates retained in each branch center |81| `S` | 16 | signed residual coordinates per chunk-branch membership |82| `L` | 12 | sparse chunk terms retained in the frozen geometry path |83| `P` | 100 | large-route shortlist for the final six-dataset row |84 85### 2.2 Query-time retrieval86 87<p align="center">88  <img src="figures/fig03_querytime_retrieval.png" width="900" alt="SparseGeometricRAG query-time retrieval architecture">89</p>90<p align="center"><em>Figure 3. Query-time computation is staged. Sparse routing finds candidate branch memberships; the local geometric comparison touches only 16 coordinates; detailed chunk evidence is evaluated only after shortlist reduction.</em></p>91 92A query is converted to sparse TF-IDF amplitudes. Weak second-order expansion is used only for routing; it is not a dense semantic representation. Routed branch postings produce candidate memberships. Each candidate is compared locally using the 16 retained signed residual coordinates, then aggregated at the document level. Cheap whole-chunk support provides an early lexical rescue/pre-score. Only a small shortlist proceeds to the richer final evidence calculation.93 94The final relevance score combines the geometric tail with whole-chunk lexical evidence, sparse semantic support, rare-term coverage, and coordination. Branch quality is estimated from the strongest branch-specific evidences; only high-quality branches are eligible for the small diversity bonus used for ranks 2-10. Rank 1 remains pure relevance.95 96### 2.3 What one chunk actually stores97 98<p align="center">99  <img src="figures/fig04_representation_anatomy.png" width="900" alt="SparseGeometricRAG per-chunk representation anatomy">100</p>101<p align="center"><em>Figure 4. The chunk-local geometric state is deliberately tiny: four fuzzy memberships and sixteen signed residual positions per membership. Branch centers and term-neighbor graphs are shared across chunks.</em></p>102 103The asymmetry between document and query representations is intentional. Document residual amplitudes are discarded after their signs and reliability structure have been retained; query amplitudes remain real-valued. The database therefore carries **direction**, while the query supplies **magnitude** at runtime.104 105This differs from dense retrieval, where every chunk generally contributes a full dense vector to the search object. Here, the local geometry is explicitly bounded by `F` and `S`, with sparse lexical support retained separately for the rescue and final evidence stages.106 107---108 109## 3. Complexity and why the method is fast110 111### 3.1 Query-time computation112 113Let:114 115- `Q` be the number of nonzero query terms;116- `K_r` be the retained routing neighbors per query term;117- `C` be the number of routed branch-membership hits;118- `U` be the number of unique routed documents;119- `S = 16` be the residual support;120- `P` be the final shortlist size; and121- `L_d` be the average binary-support length of a shortlisted chunk.122 123The main query-time stages are:124 125| Stage | Work |126|---|---:|127| Sparse query construction | `O(query tokens)` |128| Weak routing expansion | `O(Q K_r)` |129| Posting traversal | `O(C)` |130| Local geometric scoring | **`O(C S) = O(16 C)`** |131| Candidate aggregation | `O(C log C)` in the frozen reference path; `O(C)` in the preserved stamp-aggregation optimization |132| Cheap lexical rescue | proportional to routed/gated binary support |133| Shortlist selection | approximately linear partial selection in `U` |134| Final evidence extraction | performed only on the shortlist `P` |135| Top-10 construction | small overhead after shortlist features are available |136 137<p align="center">138  <img src="figures/fig05_computation_funnel.png" width="900" alt="SparseGeometricRAG computation funnel">139</p>140<p align="center"><em>Figure 5. Corpus scale does not imply corpus-wide expensive scoring. Each stage reduces the active set before the next, richer computation is permitted to run.</em></p>141 142The decisive bounded term is the local geometric score: **only 16 coordinates are consulted for each routed membership**. Rich chunk-level evidence is deliberately positioned after routing and shortlist reduction. This is the main reason the system can stay CPU-native without replacing one expensive dense-search primitive by another.143 144### 3.2 Storage complexity145 146Ignoring implementation dtypes and small metadata, the structured state scales conceptually as147 148```text149O(N F S)150+ O(M B)151+ O(M (K_assoc + K_route))152+ O(total binary term support)153```154 155where `N` is the number of chunks and `M` is the vocabulary size. The first term is the chunk-local geometric state; the second and third are shared sparse structures; the final term is the whole-chunk binary lexical support.156 157With the frozen `F = 4` and `S = 16`, the local residual layer retains only **64 residual positions across the four memberships of a chunk**. This is the structural reason the method does not require an `N x d` dense document matrix.158 159### 3.3 Preserved post-benchmark optimizations160 161The repository also preserves a later optimization branch that was developed after the frozen six-dataset benchmark. It introduces:162 163- stamp-based `O(C)` candidate aggregation instead of sort-based deduplication; and164- a gate before whole-chunk lexical scanning.165 166These optimizations are kept separate from the canonical benchmark implementation so that the reported frozen results are not silently changed after the fact.167 168---169 170## 4. Hardware and deployment requirements171 172<p align="center">173  <img src="figures/fig06_low_cost_deployment.png" width="900" alt="SparseGeometricRAG low-cost deployment architecture">174</p>175<p align="center"><em>Figure 6. Retrieval needs only CPU, system RAM, and local corpus/index storage. A GPU may still be used by the generator, but it is not a dependency of the retriever.</em></p>176 177The low-cost hardware story is central, not incidental. SparseGeometricRAG is designed for environments where a dedicated retrieval GPU is undesirable or unavailable: inexpensive servers, lab workstations, teaching machines, air-gapped systems, and deployments where accelerator memory is reserved for generation.178 179| Deployment | Retrieval requirement | Typical reason to use it |180|---|---|---|181| Laptop / teaching machine | ordinary CPU + modest RAM | development, instruction, small corpora |182| Commodity workstation/server | multicore CPU + more RAM | larger corpora and batch evaluation |183| Air-gapped / cost-constrained | CPU + local storage | no hosted model/API dependency |184| GPU-equipped RAG system | GPU optional for generator | retrieval does not compete for accelerator memory |185 186The claim is **not** that GPUs are undesirable. The claim is that **the retriever is designed so they are optional rather than mandatory**.187 188---189 190## 5. Why the objective is top-10 RAG191 192A practical generator usually consumes only a small number of retrieved chunks. For that reason, this repository treats shortlist size as a RAG operating parameter rather than assuming that the setting that maximizes deep recall must also be best for top-10 context selection.193 194<p align="center">195  <img src="figures/fig07_shortlist_sweep.png" width="900" alt="SparseGeometricRAG shortlist sweep on TREC-COVID and SciFact">196</p>197<p align="center"><em>Figure 7. TREC-COVID and SciFact expose opposite regimes. On the large TREC-COVID route, `P = 100` is a useful denoising operating point. On the tiny SciFact route, quality continues improving as aggressive pruning is relaxed.</em></p>198 199This is why the final benchmark emphasizes `nDCG@10`, `MRR@10`, `P@10`, `R@10`, `Hit@10`, and query latency. Deep recall remains useful as a diagnostic, but it is not allowed to determine the final RAG shortlist by itself.200 201---202 203## 6. Frozen six-dataset CPU results204 205<p align="center">206  <img src="figures/fig08_frozen_results.png" width="900" alt="SparseGeometricRAG frozen six-dataset CPU benchmark heatmap">207</p>208<p align="center"><em>Figure 8. Frozen `100 -> 10` effectiveness across six datasets, with representative median latencies. The result should be read as a quality/cost tradeoff rather than an accuracy-at-any-cost claim.</em></p>209 210### 6.1 OURS: concise benchmark row211 212| Dataset | nDCG@10 | MRR@10 | P@10 | R@10 | Hit@10 | Median latency (ms) | p95 latency (ms) |213|---|---:|---:|---:|---:|---:|---:|---:|214| SciFact | 0.5685 | 0.5452 | 0.0737 | 0.6663 | 0.6833 | 0.947 | 1.031 |215| TREC-COVID | 0.5990 | 0.8252 | 0.6520 | 0.0163 | 1.0000 | 1.051 | 1.193 |216| Quora | 0.7366 | 0.7287 | 0.1124 | 0.8407 | 0.8904 | 120.301 | 175.494 |217| MS MARCO / DL19 | 0.3400 | 0.5189 | 0.2674 | 0.0916 | 0.7209 | 65.195 | 142.647 |218| HotpotQA | 0.4670 | 0.6255 | 0.0964 | 0.4820 | 0.7507 | 28.944 | 42.568 |219| NQ | 0.2579 | 0.2262 | 0.0462 | 0.3983 | 0.4342 | 33.686 | 44.161 |220 221The strongest neural systems remain ahead in pure effectiveness on many datasets. SparseGeometricRAG instead targets the low-cost corner of the design space: **CPU-first retrieval with bounded sparse computation and no retrieval-time neural inference**.222 223---224 225## 7. Full benchmark suite226 227The full suite is intentionally retained. Missing entries are shown as `NR`; baselines are not removed merely because a compatible value is unavailable.228 229### 7.2.1 nDCG@10230 231| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |232|---|---:|---:|---:|---:|---:|---:|233| Exact TF-IDF | 0.5780 | 0.3738 | NR | NR | NR | NR |234| BM25 | 0.6650 | 0.6560 | 0.7890 | 0.2280 | 0.6030 | 0.3290 |235| MiniLM + FAISS Flat | 0.6451 | 0.4725 | 0.8756 | 0.3654 | 0.4651 | 0.4387 |236| BGE-base + FAISS Flat | 0.7404 | 0.7807 | 0.8890 | 0.4135 | 0.7260 | 0.5415 |237| BGE-base + FAISS HNSW | 0.7404 | 0.7807† | 0.8890† | 0.4135† | 0.7260† | 0.5415† |238| BGE-base + FAISS IVF-Flat | 0.7255 | 0.7807† | 0.8890† | 0.4135† | 0.7260† | 0.5415† |239| BGE-base + FAISS IVF-PQ | 0.6979 | 0.7807† | 0.8890† | 0.4135† | 0.7260† | 0.5415† |240| BGE-base + hnswlib HNSW | 0.7404 | 0.7807† | 0.8890† | 0.4135† | 0.7260† | 0.5415† |241| BGE-base + ScaNN | 0.6783 | 0.7807† | 0.8890† | 0.4135† | 0.7260† | 0.5415† |242| Contriever-MS MARCO + FAISS | 0.6770 | 0.5960 | 0.8650 | 0.4070 | 0.6380 | 0.4980 |243| SPLADE++ | 0.7040 | 0.7270 | 0.8340 | 0.4330 | 0.6870 | 0.5370 |244| Modern ColBERT | 0.7645 | 0.8341 | 0.8754 | 0.4499 | 0.7667 | 0.6169 |245| **OURS — CPU, 100→10** | **0.5685** | **0.5990** | **0.7366** | **0.3400** | **0.4670** | **0.2579** |246 247### 7.2.2 MRR@10248 249| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |250|---|---:|---:|---:|---:|---:|---:|251| Exact TF-IDF | 0.5437 | 0.5915 | NR | NR | NR | NR |252| BM25 | 0.6460 | 0.8530 | 0.7790 | 0.1800 | 0.8030 | 0.2630 |253| MiniLM + FAISS Flat | 0.6110 | 0.7244 | NR | NR | 0.4446 | NR |254| BGE-base + FAISS Flat | 0.7034 | 0.9180 | 0.8823 | 0.3502 | 0.8611 | 0.4924 |255| BGE-base + FAISS HNSW | 0.7034 | 0.9180† | 0.8823† | 0.3502† | 0.8611† | 0.4924† |256| BGE-base + FAISS IVF-Flat | 0.6879 | 0.9180† | 0.8823† | 0.3502† | 0.8611† | 0.4924† |257| BGE-base + FAISS IVF-PQ | 0.6615 | 0.9180† | 0.8823† | 0.3502† | 0.8611† | 0.4924† |258| BGE-base + hnswlib HNSW | 0.7034 | 0.9180† | 0.8823† | 0.3502† | 0.8611† | 0.4924† |259| BGE-base + ScaNN | 0.6494 | 0.9180† | 0.8823† | 0.3502† | 0.8611† | 0.4924† |260| Contriever-MS MARCO + FAISS | 0.6207 | NR | NR | NR | NR | NR |261| SPLADE++ | 0.6699 | NR | NR | 0.3830 | NR | NR |262| Modern ColBERT | 0.7390 | 0.9533 | 0.8671 | 0.3849 | 0.9188 | 0.5655 |263| **OURS — CPU, 100→10** | **0.5452** | **0.8252** | **0.7287** | **0.5189** | **0.6255** | **0.2262** |264 265### 7.2.3 Precision@10266 267| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |268|---|---:|---:|---:|---:|---:|---:|269| Exact TF-IDF | 0.0797 | 0.4020 | NR | NR | NR | NR |270| BM25 | 0.0863 | 0.6360 | ~0.1200 | NR | NR | NR |271| MiniLM + FAISS Flat | 0.0883 | 0.5040 | 0.1337 | 0.0591 | 0.0974 | 0.0770 |272| BGE-base + FAISS Flat | 0.0987 | 0.8300 | 0.1346 | 0.0656 | 0.1515 | 0.0884 |273| BGE-base + FAISS HNSW | 0.0987 | 0.8300† | 0.1346† | 0.0656† | 0.1515† | 0.0884† |274| BGE-base + FAISS IVF-Flat | 0.0973 | 0.8300† | 0.1346† | 0.0656† | 0.1515† | 0.0884† |275| BGE-base + FAISS IVF-PQ | 0.0950 | 0.8300† | 0.1346† | 0.0656† | 0.1515† | 0.0884† |276| BGE-base + hnswlib HNSW | 0.0987 | 0.8300† | 0.1346† | 0.0656† | 0.1515† | 0.0884† |277| BGE-base + ScaNN | 0.0887 | 0.8300† | 0.1346† | 0.0656† | 0.1515† | 0.0884† |278| Contriever-MS MARCO + FAISS | 0.0883 | NR | NR | NR | NR | NR |279| SPLADE++ | 0.0937 | NR | NR | NR | NR | NR |280| Modern ColBERT | 0.0977 | 0.8820 | 0.1327 | 0.0701 | 0.1548 | 0.0979 |281| **OURS — CPU, 100→10** | **0.0737** | **0.6520** | **0.1124** | **0.2674** | **0.0964** | **0.0462** |282 283### 7.2.4 Recall@10284 285| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |286|---|---:|---:|---:|---:|---:|---:|287| Exact TF-IDF | 0.7135 | 0.0105 | NR | NR | NR | NR |288| BM25 | 0.7809 | 0.0158 | 0.8854 | NR | 0.6531 | NR |289| MiniLM + FAISS Flat | 0.7833 | 0.0128 | 0.9503 | 0.5676 | 0.4870 | 0.6471 |290| BGE-base + FAISS Flat | 0.8742 | 0.0221 | 0.9574 | 0.6277 | 0.7574 | 0.7469 |291| BGE-base + FAISS HNSW | 0.8742 | 0.0221† | 0.9574† | 0.6277† | 0.7574† | 0.7469† |292| BGE-base + FAISS IVF-Flat | 0.8609 | 0.0221† | 0.9574† | 0.6277† | 0.7574† | 0.7469† |293| BGE-base + FAISS IVF-PQ | 0.8441 | 0.0221† | 0.9574† | 0.6277† | 0.7574† | 0.7469† |294| BGE-base + hnswlib HNSW | 0.8742 | 0.0221† | 0.9574† | 0.6277† | 0.7574† | 0.7469† |295| BGE-base + ScaNN | 0.7814 | 0.0221† | 0.9574† | 0.6277† | 0.7574† | 0.7469† |296| Contriever-MS MARCO + FAISS | 0.7868 | NR | NR | NR | NR | NR |297| SPLADE++ | 0.8230 | NR | NR | NR | NR | NR |298| Modern ColBERT | 0.8647 | 0.0230 | 0.9516 | 0.6710 | 0.7739 | 0.8239 |299| **OURS — CPU, 100→10** | **0.6663** | **0.0163** | **0.8407** | **0.0916** | **0.4820** | **0.3983** |300 301### 7.2.5 Hit@10302 303| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |304|---|---:|---:|---:|---:|---:|---:|305| Exact TF-IDF | 0.7333 | 0.8600 | NR | NR | NR | NR |306| BM25 | 0.8033 | 1.0000 | 0.9286 | NR | NR | NR |307| MiniLM + FAISS Flat | NR | NR | NR | NR | NR | NR |308| BGE-base + FAISS Flat | 0.8833 | NR | NR | NR | NR | NR |309| BGE-base + FAISS HNSW | 0.8833 | NR | NR | NR | NR | NR |310| BGE-base + FAISS IVF-Flat | 0.8700 | NR | NR | NR | NR | NR |311| BGE-base + FAISS IVF-PQ | 0.8567 | NR | NR | NR | NR | NR |312| BGE-base + hnswlib HNSW | 0.8833 | NR | NR | NR | NR | NR |313| BGE-base + ScaNN | 0.7900 | NR | NR | NR | NR | NR |314| Contriever-MS MARCO + FAISS | 0.7967 | NR | NR | NR | NR | NR |315| SPLADE++ | 0.8333 | NR | NR | NR | NR | NR |316| Modern ColBERT | 0.8100 | NR | NR | NR | NR | NR |317| **OURS — CPU, 100→10** | **0.6833** | **1.0000** | **0.8904** | **0.7209** | **0.7507** | **0.4342** |318 319### 7.2.6 Median query latency (ms)320 321| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |322|---|---:|---:|---:|---:|---:|---:|323| Exact TF-IDF | 6.197 | 50.185 | NR | NR | NR | NR |324| BM25 | 0.561 | 4.912 | NR | NR | NR | NR |325| MiniLM + FAISS Flat | NR | NR | NR | NR | NR | NR |326| BGE-base + FAISS Flat | 10.013 | NR | NR | NR | NR | NR |327| BGE-base + FAISS HNSW | 10.369 | NR | NR | NR | NR | NR |328| BGE-base + FAISS IVF-Flat | 10.050 | NR | NR | NR | NR | NR |329| BGE-base + FAISS IVF-PQ | 13.910 | NR | NR | NR | NR | NR |330| BGE-base + hnswlib HNSW | 10.594 | NR | NR | NR | NR | NR |331| BGE-base + ScaNN | 9.947 | NR | NR | NR | NR | NR |332| Contriever-MS MARCO + FAISS | 7.472 | NR | NR | NR | NR | NR |333| SPLADE++ | 13.493 | NR | NR | NR | NR | NR |334| Modern ColBERT | 62.671 | NR | NR | NR | NR | NR |335| **OURS — CPU, 100→10** | **0.947** | **1.051** | **120.301** | **65.195** | **28.944** | **33.686** |336 337### 7.2.7 p95 query latency (ms)338 339| Method | SciFact | TREC-COVID | Quora | MS MARCO / DL19 | HotpotQA | NQ |340|---|---:|---:|---:|---:|---:|---:|341| Exact TF-IDF | 14.605 | 56.716 | NR | NR | NR | NR |342| BM25 | 5.645 | 7.336 | NR | NR | NR | NR |343| MiniLM + FAISS Flat | NR | NR | NR | NR | NR | NR |344| BGE-base + FAISS Flat | 15.398 | NR | NR | NR | NR | NR |345| BGE-base + FAISS HNSW | 12.925 | NR | NR | NR | NR | NR |346| BGE-base + FAISS IVF-Flat | 16.303 | NR | NR | NR | NR | NR |347| BGE-base + FAISS IVF-PQ | 18.496 | NR | NR | NR | NR | NR |348| BGE-base + hnswlib HNSW | 13.293 | NR | NR | NR | NR | NR |349| BGE-base + ScaNN | 11.858 | NR | NR | NR | NR | NR |350| Contriever-MS MARCO + FAISS | 9.628 | NR | NR | NR | NR | NR |351| SPLADE++ | 19.156 | NR | NR | NR | NR | NR |352| Modern ColBERT | 70.994 | NR | NR | NR | NR | NR |353| **OURS — CPU, 100→10** | **1.031** | **1.193** | **175.494** | **142.647** | **42.568** | **44.161** |354 355**Notes.** `NR` means “not reported under a compatible metric / protocol in the current ledger.” `†` indicates that, outside SciFact, the BGE ANN-backend rows mirror the BGE-base representation-level effectiveness reference rather than a separately rerun backend-specific effectiveness experiment.356 357---358 359## 8. Interpreting the latency numbers360 361Speed comparisons in retrieval are easy to misstate. ANN papers frequently report **search-only** latency after a dense query embedding already exists, whereas a deployed RAG request pays for query representation, retrieval, shortlist scoring, and final selection.362 363This repository therefore follows two rules:364 3651. **do not silently compare ANN-only latency with end-to-end retrieval latency**; and3662. **do not fill missing latency cells using measurements from incompatible hardware or protocols**.367 368The full provenance policy is documented in `docs/BASELINE_SUITE.md` and `docs/LITERATURE_AND_SPEED.md`. The frozen OURS timings include the retrieval path used by the reported experiment. For the MS MARCO column, OURS is the 43-query TREC-DL19 run on the full 8.84M-passage corpus; published model-reference values in that column may use MS MARCO dev where applicable.369 370---371 372## 9. Repository organization373 374The repository is deliberately split between reusable retrieval code, clean benchmark runners, the full-scale experimental campaign, and frozen outputs.375 376| Path | Purpose |377|---|---|378| `geomretrieval/` | reusable sparse geometric retriever |379| `experiments/beir/` | BEIR runners and pool-sweep experiments |380| `experiments/msmarco_scale/` | full MS MARCO scale campaign |381| `experiments/postbenchmark_optimizations/` | preserved later optimization branch |382| `results/` | frozen JSON outputs and benchmark artifacts |383| `configs/` | reproduction handles |384| `baselines/` | baseline utilities |385| `scripts/` | runnable helpers |386| `docs/` | method, RAG protocol, baseline, speed, and reproducibility notes |387| `tests/` | smoke tests |388| `manifests/` | corpus / run manifests |389 390The final six-dataset artifacts are under `results/final_100_to_10/`. The later stamp-aggregation and lexical-gating optimization results are preserved separately and are not used to rewrite the frozen benchmark row.391 392---393 394## 10. Reproducibility395 396The repository contains the exact frozen result JSONs, benchmark scripts, configuration handles, and tests used to reconstruct the final evaluation. The reference package was checked with the project smoke tests before release.397 398For a clean reproduction path, start with:399 4001. `docs/METHOD.md` - algorithmic description;4012. `docs/RAG_PROTOCOL.md` - top-10 evaluation protocol;4023. `docs/REPRODUCIBILITY.md` - environment and run guidance;4034. `docs/BASELINE_SUITE.md` - baseline/provenance policy; and4045. `results/final_100_to_10/` - frozen final outputs.405 406The code is intentionally CPU-first. Baseline packages that depend on neural encoders or ANN libraries are listed separately from the core requirements.407 408---409 410## 11. Scope of the claim411 412### What this repository claims413 414- a **CPU-first** retrieval architecture with no retrieval-time transformer inference;415- no requirement for a dense vector per document or a GPU-based retrieval service;416- fixed small structural handles (`F = 4`, `B = 64`, `S = 16`) controlling local geometry;417- local geometric scoring bounded by `O(CS)` with `S = 16` fixed;418- deliberate postponement of richer chunk-level computation until after routing and shortlist reduction;419- a practical top-10 RAG operating point validated across six datasets;420- complete benchmark tables rather than selective reporting of only OURS.421 422### What it does not claim423 424- dominance over the strongest neural retrievers in pure effectiveness;425- that every latency cell in the literature is directly comparable across hardware and protocol;426- that one shortlist size is mathematically optimal for every dataset or route size;427- that deep recall is irrelevant. It is retained as a diagnostic, but it is not the sole deployment objective.428 429---430 431## 12. License432 433This repository is released under the **Fair Noncommercial Research License** selected on the Hugging Face repository. Check the repository license metadata and license text before redistribution or commercial use.434