← Back to Paper Reading

  • AskChem: Making Provenance-Carrying Claims the Retrieval Unit

    Advanced Retrieval, memory, and production RAG
    Retrieval Systems · Part 1 · Note · Aug 7, 2026 · Paper · 2026 · NLP

    A critical reading of AskChem's atomic claims, source locators, faceted taxonomy, evidence graph, and AskChem-Bench results, with a clear boundary between citation traceability and scientific correctness.

    Understand it in 90 seconds
    Problem
    paper/chunk retrieval leaves a reader or agent to find the supporting sentence, establish claim location, and synthesize across papers.
    Core insight
    AskChem makes typed atomic claims with DOI and quote/evidence locators the retrieval unit, then connects them through taxonomy, an evidence graph, and shared REST/SDK/MCP interfaces.
    Strongest evidence
    the 2.4M-claim, 147K-paper system reports 100% DOI resolvability for AskChem-grounded answers versus 88.3% for LLM-only on 30 chemistry synthesis questions (Section 7; Table 1).
    Main boundary
    DOI resolvability and citation density are provenance proxies, not proof of claim truth, complete literature coverage, or usable chemical conclusions.
    Read the full deep dive
  • BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

    Intermediate Retrieval, memory, and production RAG
    Retrieval Systems Deep Dive · Part 2 · Note · Aug 7, 2026 · Paper · 2026 · NLP

    A deep reading of Wang et al.'s arXiv v3 study: across 28 nested enterprise-shaped corpus tiers with fixed questions, evidence, and adversarial documents, why BM25 crosses over at roughly 10 million corpus tokens and why agents should begin after global candidate discovery.

    Understand it in 90 seconds
    Problem
    RAG paradigms are often compared at one corpus size, hiding joint accuracy, construction-cost, query-cost, and latency scaling.
    Core insight
    28 nested corpus tiers (1,144 to 511,959 documents) hold reader/judge and adversarial bedrock fixed while comparing lexical, dense, graph, and file-system agency; a retrieval-swap control isolates access substrate.
    Strongest evidence
    at large shared tiers, BM25 reportedly overtakes raw file-system agency around 10M corpus tokens; a matched 150-question resweep gives Agent+BM25 69.4 versus raw-file agency 36.9 (Section 5.1; Figure 4; Table 4).
    Main boundary
    EnterpriseRAG-Bench is fictional and enterprise-shaped, with 500 questions and one main reader/judge; no public executable data/benchmark artifact is confirmed, so this is not “BM25 always wins.”
    Read the full deep dive
  • RubricRanker Deep Read: RAG Needs the Right Document Set, Not Just the Most Relevant Documents

    Advanced Retrieval, memory, and production RAG
    Retrieval Systems · Part 3 · Note · Aug 7, 2026 · Paper · 2026 · Retrieval Systems

    A close reading of how RubricRanker uses query-specific search rubrics, SFT, and GRPO to train a document reranker, and what its deep-research and RAG benchmark results actually establish.

    Understand it in 90 seconds
    Problem
    Traditional rerankers score documents independently, so the top k need not be complete, concise, consistent, or authoritative as a set.
    Core insight
    Change the output target from a document ranking to an evidence set that jointly supports the answer, using query-specific rubrics for labels and rewards.
    Strongest evidence
    Tables 1–3 show downstream gains, while the ablation points to rubric labels and cold-start SFT rather than RL alone.
    Main boundary
    Final answers are still produced by agents and scored by LLM judges; a better evidence set does not guarantee correct citation, reasoning, or facts.
    Read the full deep dive

For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.

Speaking & contact