PAPERS / REPRODUCTION / ENGINEERING NOTES
Paper reading for working engineers
Structured readings that extract architecture decisions, reproducible methods, evaluation limits, and what changes when research meets a real system.
- reading notes
- 81
- research topics
- 7
- learning paths
- 3
READING LIBRARY
Paper archive — page 3
Showing notes 49–72 of 81.
-
Reflexion: Write Verbal Reflections into Memory — Do Not Mistake Retries for Weight Learning
Intermediate Agent runtime, safety, and evaluationA deep read of Shinn et al., NeurIPS 2023: frozen weights plus linguistic feedback stored in episodic memory for across-trial verbal credit assignment. HumanEval pass@1 91.0 vs GPT-4 80.1 is a programming setup with self-tests and retries; WebShop and MBPP mark the boundary.
Understand it in 90 seconds
-
MemGPT: Treat Context as Paged Memory — Do Not Mistake the OS Metaphor for Enterprise Memory
Intermediate Agent runtime, safety, and evaluationA deep read of Packer et al., arXiv:2310.08560 v2: finite context as RAM, OS-like tiers, and function-mediated paging. DMR moves GPT-4 from 32.1% to 92.5%; nested KV shows multi-hop lookup—not ACL memory governance.
Understand it in 90 seconds
-
CoT: Make the Model Write the Reasoning, but Do Not Treat It as an Agent That Moves
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Wei et al., NeurIPS 2022: few-shot exemplars with intermediate steps elicit multi-step reasoning in large frozen models. PaLM 540B on GSM8K moves from 17.9 to 56.9; this is still a prompt, not tools, an environment, or memory paging.
Understand it in 90 seconds
-
WebGPT: Let the Model Browse for Answers, but Do Not Treat It as a Reasoning Agent Loop
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Nakano et al., arXiv:2112.09332 v3: GPT-3 is given a text browser and trained with human demonstrations and preference / reward modeling to search, quote, and answer. The 175B best-of-64 model is preferred 56% versus demonstrators and 69% versus Reddit; this is browsing QA, not ReAct’s thought–action–observation contract.
Understand it in 90 seconds
-
RAG: Attach Retrieval to Generation, but Do Not Treat 2020 RAG as a Production RAG Platform
Intermediate Retrieval, memory, and production RAGA source-grounded reading of Lewis et al., NeurIPS 2020: BART is paired with a dense retriever over Wikipedia, and RAG-Sequence / RAG-Token condition generation on retrieved passages. RAG-Seq reaches 44.5 Exact Match on NQ; this is a 2020 method paper, not a 2025 production RAG platform and not an agent loop.
Understand it in 90 seconds
-
DPR: Turn Open-Domain QA into Dense Passage Retrieval, but Do Not Treat the Dual Encoder as Production RAG
Intermediate Retrieval, memory, and production RAGA source-grounded reading of Karpukhin et al., EMNLP 2020: a BERT dual encoder trained with question–passage pairs and in-batch negatives replaces BM25 over Wikipedia passages via MIPS. On NQ, top-20 retrieval is 78.4% vs BM25 59.1%; end-to-end Exact Match is 41.5. This is the retriever RAG uses, not a generation platform.
Understand it in 90 seconds
-
Self-RAG: Let the Model Decide When to Retrieve, but Do Not Treat Reflection Tokens as a Production RAG Gate
Intermediate Retrieval, memory, and production RAGA source-grounded reading of Asai et al., ICLR 2024: an LM is trained with reflection tokens (Retrieve / Relevant / Supported / Useful) for on-demand retrieval and self-critique. Self-RAG 7B / 13B reach 54.9 / 55.8 on PopQA; this is a when-to-retrieve method paper, not a production RAG platform and not an agent tool loop.
Understand it in 90 seconds
-
REALM: Wire Retrieval into LM Pre-Training, but Do Not Treat Joint Training as a Ready-Made RAG Stack
Intermediate Retrieval, memory, and production RAGA source-grounded reading of Guu et al., ICML 2020: a differentiable knowledge retriever is pre-trained with an MLM signal, an asynchronously refreshed MIPS index, and Open-QA fine-tuning. With CC-News / Wikipedia, NQ Exact Match is 40.4, above ORQA and T5-11B. This is costly retrieval-augmented pre-training—not Lewis RAG generation and not DPR’s cheaper dual-encoder recipe.
Understand it in 90 seconds
-
Gorilla: Turn a Large API Catalog into Retrievable Tools, but APIBench Does Not Establish MCP Product Behavior
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Patil et al., NeurIPS 2024: retriever-aware finetuning of LLaMA-7B on APIBench (TorchHub / TensorHub / HuggingFace) so catalog-scale API calls can be retrieved and checked. Zero-shot overall accuracy and hallucination beat prompted GPT-4 on that table—this is not a ReAct loop, MidTool mid-training, or RAG-MCP product routing.
Understand it in 90 seconds
-
MidTool: Does Teaching Tool Use During Mid-Training Make Agents More Reliable?
Advanced Agent runtime, safety, and evaluationA deep reading of MidTool, which moves schema grounding, workflow composition, and recovery under incomplete information into a 20.3B-token mid-training mixture—while web search remains at 0%.
Understand it in 90 seconds
-
SWE-Bench ProMax: Can Large-Scale Multilingual Refactoring Measure Long-Horizon Coding Agents?
Advanced Agent runtime, safety, and evaluationA deep reading of SWE-Bench ProMax, which uses 170 cross-file, multilingual, behavior-preserving refactoring tasks to test whether coding agents can complete large changes rather than merely fix a nearby test.
Understand it in 90 seconds
-
ADIAS: Turning Agent Self-Improvement into Traceable Issue Repair
Advanced Agent runtime, safety, and evaluationA deep reading of ADIAS: persistent issue state organizes failure evidence across optimization rounds so a full-code agent designer can remember what was tried, what regressed, and when a repair is actually confirmed.
Understand it in 90 seconds
-
DocMemo: Letting Long-Document RAG Recover from a Bad First Retrieval
Advanced Retrieval, memory, and production RAGA deep reading of DocMemo: document schema, page belief, and question episodic memory preserve retrieval state across rounds, while Bayesian updates, Thompson sampling, and adaptive granularity recover missed evidence.
Understand it in 90 seconds
-
Agentic Configuration Management: Treating Agent Systems as Governed Configuration, Not Just One Execution
Advanced Agent runtime, safety, and evaluationA deep reading of how ACM uses a framework-independent Configuration Graph, immutable revisions, dependency-aware impact propagation, and runtime provenance to govern heterogeneous agent configurations across LangGraph, CrewAI, and the OpenAI Agents SDK.
Understand it in 90 seconds
-
FinRank: Hard-Negative Retrieval Evaluation for Financial-Document RAG
Intermediate Retrieval, memory, and production RAGA deep reading of FinRank: how company, year, and disclosure boundaries create deceptively plausible evidence, and why pooled retrieval, hard negatives, and metadata filters must be evaluated together.
Understand it in 90 seconds
-
A²E: A Traceable, Re-Evaluable Engine for Agent Auditing
Intermediate Agent runtime, safety, and evaluationA deep reading of A²E: ATP aligns benchmarks with agent harnesses, span-based traces preserve execution causality, and lifecycle-aligned metrics analyze correctness, tools, cost, and safety.
Understand it in 90 seconds
-
ContextWeave Deep Read: Does Memory Actually Make Agents Better at Work?
Advanced Agent runtime, safety, and evaluationA close reading of how ContextWeave reconstructs multi-month workflows into an executable benchmark and measures memory's effect on workspace outcomes, preference adherence, continuity, and misleading recall.
Understand it in 90 seconds
-
Argus Deep Read: Long-Running Agents Need a Runtime, Not a Longer Prompt
Advanced Agent runtime, safety, and evaluationA critical reading of Argus's Manager–Planner–Engineer–Reviewer runtime, durable state, verification-gated evolution, and rollback, separating benchmark results from author-operated case studies and unproven self-learning claims.
Understand it in 90 seconds
-
AskChem: Making Provenance-Carrying Claims the Retrieval Unit
Advanced Retrieval, memory, and production RAGA critical reading of AskChem's atomic claims, source locators, faceted taxonomy, evidence graph, and AskChem-Bench results, with a clear boundary between citation traceability and scientific correctness.
Understand it in 90 seconds
-
AgentS4D Deep Read: The Task Finished—Is the Runtime Safe?
Advanced Agent runtime, safety, and evaluationA critical reading of how AgentS4D places workspace-agent risk entry, induction strategy, target harm, and lifecycle evidence in one sandbox benchmark, and why completion rate cannot stand in for safety.
Understand it in 90 seconds
-
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
Intermediate Retrieval, memory, and production RAGA deep reading of Wang et al.'s arXiv v3 study: across 28 nested enterprise-shaped corpus tiers with fixed questions, evidence, and adversarial documents, why BM25 crosses over at roughly 10 million corpus tokens and why agents should begin after global candidate discovery.
Understand it in 90 seconds
-
Real-Time Detection and Repair of LLM Agent Failures: A Deep Read of AgentTrajectorySentinel
Advanced Agent runtime, safety, and evaluationA critical reading of AgentTrajectorySentinel's low-cost healthy-only temporal monitor, deterministic verification, and rollback-and-retry loop, separating measured detection and repair gains from calibration dependence, content blind spots, and reproducibility limits.
Understand it in 90 seconds
-
Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG
Advanced Retrieval, memory, and production RAGA deep read of how Before Reasoning Can Fail turns answer-before-reading into an observable trajectory failure, and tests whether Read-Gate actually improves multi-hop QA.
Understand it in 90 seconds
-
PAST-Bench: What Did a Persistent Agent Actually Learn from the Past?
Advanced Agent runtime, safety, and evaluationA deep read of how PAST-Bench uses fresh-session task families, matched persistence controls, and trace-level mechanism evidence to separate genuine retained-experience gains from higher scores with unrelated causes.
Understand it in 90 seconds
RESEARCH EXCHANGE