Retrieval Systems
Series · 3 posts
-
AskChem: Making Provenance-Carrying Claims the Retrieval Unit
Advanced Retrieval, memory, and production RAGA critical reading of AskChem's atomic claims, source locators, faceted taxonomy, evidence graph, and AskChem-Bench results, with a clear boundary between citation traceability and scientific correctness.
Understand it in 90 seconds
- Problem
- paper/chunk retrieval leaves a reader or agent to find the supporting sentence, establish claim location, and synthesize across papers.
- Core insight
- AskChem makes typed atomic claims with DOI and quote/evidence locators the retrieval unit, then connects them through taxonomy, an evidence graph, and shared REST/SDK/MCP interfaces.
- Strongest evidence
- the 2.4M-claim, 147K-paper system reports 100% DOI resolvability for AskChem-grounded answers versus 88.3% for LLM-only on 30 chemistry synthesis questions (Section 7; Table 1).
- Main boundary
- DOI resolvability and citation density are provenance proxies, not proof of claim truth, complete literature coverage, or usable chemical conclusions.
-
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
Intermediate Retrieval, memory, and production RAGA deep reading of Wang et al.'s arXiv v3 study: across 28 nested enterprise-shaped corpus tiers with fixed questions, evidence, and adversarial documents, why BM25 crosses over at roughly 10 million corpus tokens and why agents should begin after global candidate discovery.
Understand it in 90 seconds
- Problem
- RAG paradigms are often compared at one corpus size, hiding joint accuracy, construction-cost, query-cost, and latency scaling.
- Core insight
- 28 nested corpus tiers (1,144 to 511,959 documents) hold reader/judge and adversarial bedrock fixed while comparing lexical, dense, graph, and file-system agency; a retrieval-swap control isolates access substrate.
- Strongest evidence
- at large shared tiers, BM25 reportedly overtakes raw file-system agency around 10M corpus tokens; a matched 150-question resweep gives Agent+BM25 69.4 versus raw-file agency 36.9 (Section 5.1; Figure 4; Table 4).
- Main boundary
- EnterpriseRAG-Bench is fictional and enterprise-shaped, with 500 questions and one main reader/judge; no public executable data/benchmark artifact is confirmed, so this is not “BM25 always wins.”
-
RubricRanker Deep Read: RAG Needs the Right Document Set, Not Just the Most Relevant Documents
Advanced Retrieval, memory, and production RAGA close reading of how RubricRanker uses query-specific search rubrics, SFT, and GRPO to train a document reranker, and what its deep-research and RAG benchmark results actually establish.
Understand it in 90 seconds
- Problem
- Traditional rerankers score documents independently, so the top k need not be complete, concise, consistent, or authoritative as a set.
- Core insight
- Change the output target from a document ranking to an evidence set that jointly supports the answer, using query-specific rubrics for labels and rewards.
- Strongest evidence
- Tables 1–3 show downstream gains, while the ablation points to rubric labels and cold-start SFT rather than RL alone.
- Main boundary
- Final answers are still produced by agents and scored by LLM judges; a better evidence set does not guarantee correct citation, reasoning, or facts.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact