PAPERS / REPRODUCTION / ENGINEERING NOTES
Paper reading for working engineers
Structured readings that extract architecture decisions, reproducible methods, evaluation limits, and what changes when research meets a real system.
- reading notes
- 81
- research topics
- 7
- learning paths
- 3
READING LIBRARY
Paper archive — page 2
Showing notes 25–48 of 81.
-
SilentProbe: When HTTP 200 Did Not Answer the Question
Advanced Agent runtime, safety, and evaluationA deep reading of SilentProbe (arXiv:2609.00035 v1): from OpenAPI constraint gaps and live differential probes to agent false negatives, separating how disclosure and machine-readability determine whether a tool fails honestly.
Understand it in 90 seconds
-
RAGSieve: Detecting RAG Knowledge-Poisoning Promotion with Self-Referenced Local Contrast
Advanced Retrieval, memory, and production RAGA deep read of RAGSieve: query-local and corpus-local references expose suspicious retrieval promotion without a trusted clean corpus, while the paper's boundary remains essential—promotion detection is not truth verification.
Understand it in 90 seconds
-
BTS-AgentBench: Compiling Read-Only Telemetry into Replayable Agent Episodes
Advanced Agent runtime, safety, and evaluationA deep read of Jeong-Yoon Kim's BTS-AgentBench (arXiv:2608.27334 v1): a deterministic path from building telemetry to read-only tools, executable tasks, bounded interaction contracts, and evidence-grounded evaluation—strong on replay consistency, bounded beyond production safety or arbitrary-domain portability.
Understand it in 90 seconds
-
When Failure Propagates, Can We Still Find the Start? Causal Failure Attribution in Agentic RAG
Advanced Retrieval, memory, and production RAGA deep reading of When Failures Propagate: an interventional benchmark, three-hop MuSiQue, and certified content corruption that separate failure detection, causal attribution, propagation, and recovery in agentic RAG.
Understand it in 90 seconds
-
ACE: Let a Canvas Agent Understand Structure Before It Corrects Itself
Advanced Agent runtime, safety, and evaluationA deep reading of ACE (arXiv:2608.24103 v1): hierarchical scene graphs, CARE routing, and an instruction-following judge turn multi-slide editing into a scoped, diffable, rollback-aware loop, with explicit limits around benchmarks, human raters, mock mode, and live reproduction.
Understand it in 90 seconds
-
Tool Call Succeeds, Workflow Fails: External-Effect Anomalies at the Agent–Tool Boundary
Advanced Agent runtime, safety, and evaluationA deep reading of the effect-history model behind Agent–Tool Boundary: why a successful tool response cannot guarantee a coherent external world state, and what MCP annotations and transactional contracts still leave unspecified.
Understand it in 90 seconds
-
EvoOntology: Turning a Static Data-Agent Semantic Layer into a Verifiable, Self-Evolving Interface
Advanced Retrieval, memory, and production RAGA deep reading of EvoOntology: an MCP ontology layer built by evidence-grounded probing, then refined through attribution-guided typed edits and a backbone-conditional paired gate.
Understand it in 90 seconds
-
Corrupt Plans, Clean Traces: How Plan Injection Evades Chain-of-Thought Monitoring
Advanced Agent runtime, safety, and evaluationA deep reading of Plan Injection: when an adversarial plan enters context and an actor rewrites it into plausible Chain-of-Thought, why the monitor’s clean trace can disconnect from behavioral causality, and where APPS, Bio-Math, and budget-sweep evidence stops.
Understand it in 90 seconds
-
K-Bench: Why Agentic Unlearning Cannot Be Certified from the Final Answer Alone
Advanced Agent runtime, safety, and evaluationA deep read of Yu et al.'s K-Bench (arXiv:2609.12808 v1): an end-to-end agent deployment benchmark that evaluates six observable channels across four memory substrates, using OR-of-channels leakage, collapse-aware K-Scores, and pre-registered statistical tests to separate forgetting from channel migration and agent collapse.
Understand it in 90 seconds
-
REVA: Moving RAG compression into reusable evidence views instead of paying per request
Advanced Retrieval, memory, and production RAGA critical reading of Nguyen et al.'s REVA (arXiv 2609.11209 v1): historical generator attention becomes a document-keyed score store, separating offline scoring from online rendering while exposing unseen-document fallback, local/global budgets, quality, and latency boundaries.
Understand it in 90 seconds
-
VikingRAG: Fewer Retrieval Rounds, Less Context Waste for Structured-Document RAG
Advanced Retrieval, memory, and production RAGA deep reading of VikingRAG: hierarchy-preserving URI-addressable storage, Search/List/Grep/Read tools, reusable experience edges, and adaptive escalation for reducing repeated retrieval tokens and latency.
Understand it in 90 seconds
-
DRACO: Sending Long-Horizon Agent Credit Back to the Steps
Advanced Agent runtime, safety, and evaluationA source-grounded reading of DRACO (arXiv:2609.04094): dynamic per-trajectory rubrics create an outcome-blind reward, then a closed-form rule redistributes GRPO advantage to the steps cited by the judge.
Understand it in 90 seconds
-
CONTINUITY: Keeping provenance, authorization, and tool effects continuous across Agent composition
Advanced Agent runtime, safety, and evaluationA critical reading of Zheng and Yang's CONTINUITY (arXiv:2609.05269 v1): security-context contracts, field-level provenance, transformation witnesses, and effect-bound permits for preserving an LLM Agent's instruction-to-effect boundary.
Understand it in 90 seconds
-
Parsing the Stream: Long-Horizon Agents Need Auditable Live State, Not Just Memory
Advanced Agent runtime, safety, and evaluationA critical reading of Pakhomov and Nijkamp's Parsing the Stream (arXiv:2609.01466): an append-only trace is folded into typed RunState and compiled into observer and worker views. The paper reports gains on specific accumulation tasks and monitoring costs, but does not show that fixed aggregates replace every form of trace memory.
Understand it in 90 seconds
-
Generative Agents: Observe–Reflect–Plan in a Multi-Agent Sandbox — Do Not Mistake Sandbox Memory for MemGPT OS Paging
Intermediate Agent runtime, safety, and evaluationA deep read of Park et al., UIST 2023 / arXiv:2304.03442 v2: 25 agents in Smallville use a memory stream, periodic reflection, and retrieval-based planning. Interview ablations hit TrueSkill μ 29.89 vs 21.21 fully ablated; two-day sandbox diffusion and party coordination are qualitative evidence, not production runtime.
Understand it in 90 seconds
-
ResNet: Residuals Make Depth Trainable, but ImageNet 2015 Is Not a Ready-Made Detection or ViT Contract
Intermediate Build the foundations firstA source-grounded reading of He et al., CVPR 2016 / arXiv:1512.03385: identity shortcuts let stacked layers learn residual F(x)+x and fix plain-net degradation. ResNet-152 reaches 4.49% top-5 validation error on ImageNet; this is 2015 classification evidence, not a YOLO, ViT, or modern ConvNet leaderboard contract.
Understand it in 90 seconds
-
YOLO: Detect From One Full-Image Pass, but VOC 2016 Does Not Represent the Later YOLO Family
Intermediate Build the foundations firstA source-grounded reading of Redmon et al., CVPR 2016 / arXiv:1506.02640: object detection as a single forward-pass regression—an S×S grid, B boxes, and C class probabilities in one shot. On VOC 2007, YOLO reaches 63.4% mAP at 45 FPS; this is 2016 unified-detection evidence, not YOLOv3 COCO or Ultralytics product numbers.
Understand it in 90 seconds
-
Transformer: Drop Recurrence for Self-Attention, but WMT 2017 BLEU Does Not Represent Later LLMs
Intermediate Build the foundations firstA source-grounded reading of Vaswani et al., NeurIPS 2017 / arXiv:1706.03762: stacked encoder-decoder with multi-head self-attention and positional encodings replaces RNNs and convolutions for machine translation. On WMT 2014, the big model reaches 28.4 BLEU EN-DE and 41.8 BLEU EN-FR; this is 2017 sequence-transduction evidence, not a BERT, GPT-3, or ViT contract.
Understand it in 90 seconds
-
InstructGPT: Align Instructions with Human Feedback, but 2022 Preference Win Rates Do Not Represent Later ChatGPT
Intermediate Build the foundations firstA source-grounded reading of Ouyang et al., NeurIPS 2022 / arXiv:2203.02155: on a frozen GPT-3 architecture, a three-stage pipeline—SFT demonstrations, a reward model, and PPO (default PPO-ptx)—aligns a pretrained LM to human preferences. 175B InstructGPT is preferred over 175B GPT-3 85±3% of the time; this is 2022 API labeler evidence, not a ChatGPT product SLA, GPT-4 eval, or DPO contract.
Understand it in 90 seconds
-
Speculative Decoding: Draft with a Small Model, Verify in Parallel, but T5 Speedups Do Not Represent Every Serving Stack
Intermediate Build the foundations firstA source-grounded reading of Leviathan et al., ICML 2023 / arXiv:2211.17192: a cheap draft model M_q proposes tokens, the target model M_p verifies a chunk in parallel, and rejection sampling keeps the output distribution identical to target-only decoding. T5-XXL 11B reaches 2.3X-3.4X wall-clock speedup on T5X; this is 2023 lossless inference-algorithm evidence, not a GPTQ, FlashAttention, vLLM, Medusa, or EAGLE contract.
Understand it in 90 seconds
-
Indirect Prompt Injection: Web Pages and Tool Returns Become Instruction Channels, but 2023 Cases Do Not Represent Later Guard Products
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Greshake et al., arXiv:2302.12173 v2: when LLM-integrated apps retrieve web pages, email, or tool output, untrusted data enters the prompt as if it were instructions. The authors demonstrate indirect prompt injection on Bing Chat, GitHub Copilot, and synthetic GPT-4 apps and give a computer-security threat taxonomy. This is 2023 control-plane evidence, not a Llama-Guard, Constitutional AI, OWASP Top-10, or jailbreak-benchmark product SLA.
Understand it in 90 seconds
-
ReAct: Interleave Thought and Action, but Do Not Treat a Few-Shot Loop as an Agent Runtime
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Yao et al., ICLR 2023: language thoughts join the action space, while HotpotQA, FEVER, ALFWorld, and WebShop keep hallucination, search failure, and the abstract's +34% / +10% in separate buckets.
Understand it in 90 seconds
-
Toolformer: Self-Supervised API Calls Are Not an Agent Loop
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Schick et al., NeurIPS 2023: future-token loss filters QA, Wikipedia, calculator, calendar, and translation calls on CCNet for GPT-J. LAMA and math jump; this is still not a chainable agent runtime.
Understand it in 90 seconds
-
SWE-bench: Real GitHub Issues as Evaluation, but 1.96% Is Not a Model Ceiling
Intermediate Agent runtime, safety, and evaluationA source-grounded reading of Jimenez et al., ICLR 2024 Oral: the evaluation unit is a real GitHub issue, a full Python repository, and tests. Claude 2 resolves 1.96% under BM25; that number is a protocol, not a model ranking.
Understand it in 90 seconds
RESEARCH EXCHANGE