The paper in 90 seconds
- Problem: Existing autonomous agent systems frequently borrow standard RAG pipelines for long-term memory. However, interaction logs form a bounded, coherent, highly correlated, and near-duplicate conversational stream. Relying on flat Top- vector retrieval leads to redundant collapse into dense semantic clusters, while post-hoc token pruning easily shatters fragile, temporally linked evidence chains.
- Core insight: xMemory proposes decoupling before aggregation: raw interaction streams are decomposed into atomic units and organized into a four-tier hierarchy (Message, Episode, Semantic, Theme). A dual Sparsity–Semantics objective dynamically balances cluster split and merge operations, while an adaptive top-down retrieval process expands to fine-grained episodes or raw messages only when it meaningfully reduces predictive uncertainty.
- Strongest evidence: On the long-dialogue benchmark LoCoMo (averaging ~9,000 tokens across ~300 turns), xMemory achieves superior average F1 and BLEU scores across three distinct backbones (Qwen3-8B, Llama-3.1-8B-Instruct, GPT-5 nano), outperforming Naive RAG, A-Mem, MemoryOS, LightMem, and Nemori. On temporal reasoning questions, F1 improves by 3.72 to 11.23 points, while reducing inference context token consumption by approximately 39% to 48% (Table 1).
- Main boundary: Hierarchy quality heavily depends on initial dialogue segmentation and semantic embedding representations. Benchmarks focus on academic multi-turn QA proxies rather than live production environments with concurrent writes, right-to-be-forgotten deletion mandates, semantic drift, or tool-grounded workspace state mutations. High benchmark accuracy does not establish solved memory governance or transaction consistency.
This reading is based on the arXiv:2602.02007 preprint (Hu et al., King’s College London and The Alan Turing Institute).
What to know first
To understand xMemory’s contribution, it is essential to identify why traditional RAG assumptions fail when applied to autonomous agent memory, and where previous methods encounter systemic bottlenecks:
- RAG corpus assumptions vs. Agent memory reality: Standard RAG is founded on large-scale, heterogeneous, diverse, and relatively independent knowledge bases. In contrast, agent memory originates from a bounded, continuous, single-entity or task-focused conversational stream.
- Redundant collapse in dense vector space: In long interactions, successive turns often revisit the same core topic with minor updates. Flat Top- similarity queries inevitably retrieve multiple near-duplicate chunks from the same dense semantic neighborhood. This exhausts the context budget without supplying complementary information.
- Temporal entanglement and broken prerequisites: Vital facts in conversation depend heavily on temporal sequences, coreference, ellipsis, and chronological overrides. For example, if a user revokes a permission granted two weeks earlier, flat retrieval may surface the original authorization while missing the subsequent revocation, leading to critical failure.
- Brittleness of post-hoc pruning: Compression approaches (such as LightMem paired with LLMLingua-2) attempt to eliminate redundant tokens after retrieval. However, these compressors assume independent passages; deleting tokens within conversational streams frequently severs causal and pronoun-referent relationships.
| Dimension | Standard RAG Architecture | Agent Long-Term Memory Setting |
|---|---|---|
| Corpus Nature | Large, heterogeneous, multi-domain | Bounded, coherent, single conversational stream |
| Candidate Spans | Diverse, largely independent | Highly correlated, near-duplicate, iterative |
| Primary Failure Mode | Irrelevance (fetching unrelated docs) | Redundant collapse (fetching duplicate views) |
| Evidence Structure | Unordered or parallel passages | Temporally entangled prerequisite chains |
| Retrieval Unit | Arbitrary fixed-length raw chunks | Multi-scale latent semantic components |
Consequently, the central challenge in agent memory is not simply tuning a better reranker, but fundamentally redesigning the organizational granularity during writing and the search scale during retrieval.

Figure 1, Section 1 of the paper (motivation and paradigm comparison): agent memory forms a coherent temporal sequence where traditional similarity top-k leads to redundant collapse, whereas structured retrieval navigates latent multi-scale components. See the original Figure 1 anchor and arXiv HTML figure endpoint. arXiv source identifies a perpetual non-exclusive license; reproduced under arXiv reuse terms with attribution.
Core intuition
The core intuition of xMemory is that memory retrieval should not be surface-level text span matching, but topological navigation across multi-scale latent components.
Under previous decision rules, memory systems rank all raw chunks by global similarity, allowing redundant variations of the same event to dominate the context window. xMemory replaces this with a two-phase decision paradigm: decoupling before aggregation.
The system decomposes raw conversational turns into minimal editable units, then aggregates them bottom-up into a four-tier hierarchy governed by semantic coherence and structural sparsity:
- Message: Indivisible single-turn raw utterances;
- Episode: Temporal summaries of continuous message blocks preserving sequential context;
- Semantic: Reusable long-term atomic facts extracted from episodes, acting as the primary retrieval indices;
- Theme: Broad topic clusters grouping semantically related fact nodes.
During inference, retrieval proceeds top-down rather than flatly scanning all text: the query first isolates relevant Themes, explores representative Semantic facts, and finally expands down to intact Episodes or raw Messages only when doing so reduces downstream model uncertainty. Every downward step incurs an explicit token cost, systematically preventing redundant collapse.

Figure 2, the paper’s Section 2 methodology overview: the figure places the four-tier hierarchy, sparsity–semantics objective, and top-down retrieval in one method context. See the original Figure 2 anchor and arXiv HTML figure endpoint. The arXiv source states a perpetual non-exclusive license; this article preserves attribution and follows the arXiv reuse terms.
Walk one example through the method
To trace the end-to-end mechanics of xMemory, consider a long-horizon engineering operations scenario:
- Input:
- Historical Memory : Two months of accumulated dialogue across 35 multi-turn sessions (tens of thousands of tokens), discussing database migrations, service architectures, and security compliance changes.
- User Query : “Is last week’s production database deployment exception still active?”
- Intermediate representation:
- The memory store is structured into four tiers: at the scale of LoCoMo, this comprises roughly 650 Theme nodes, 2,900 Semantic nodes, and 750 Episode blocks.
- Theme and Semantic nodes maintain bidirectional NN graph edges.
- A Theme node “Production Database Policy” encompasses several Semantic nodes, including “Initial Exception Approval”, “Connection Timeout Tuning”, and “Security Review Revocation Notice”, each linked to their underlying Episode blocks.
- Decision or transformation:
- Stage I: Query-Aware Representative Selection: Query navigates the Theme graph. Using Eq. (4) to balance coverage and relevance, the algorithm selects “Production Database Policy” while pruning unrelated topics like “Frontend Styling”. Within the induced Semantic subgraph, the algorithm greedily selects representative facts. Because the objective rewards cluster coverage over pure lexical overlap, it selects not only the “Initial Exception Approval” node but also the subsequent, lexically distinct “Security Review Revocation Notice” node.
- Stage II: Uncertainty-Adaptive Evidence Inclusion: The system retrieves the intact Episode summaries backing the selected Semantic nodes. Rather than indiscriminately concatenating all text, it measures whether including each Episode reduces the predictive entropy of the reader model. Incorporating the “Revocation Notice” Episode drastically reduces uncertainty regarding the exception’s status. The algorithm detects that additional raw historical logs offer negligible information gain, triggering early stopping and preventing context bloat.
- Output:
- A highly concise, temporally sound context is constructed (consuming a few hundred tokens rather than several thousand).
- The downstream LLM generates an accurate, chronologically faithful response: “The deployment exception was revoked during last Friday’s security review meeting; standard approval requirements are now back in effect.”
- Likely failure point:
- If the Sparsity–Semantics objective misclusters “Security Audit” and “Routine Maintenance” into a single bloated Theme, Stage I might fail to select the critical revocation node due to candidate dispersion;
- If the entropy proxy model misjudges uncertainty reduction, early stopping could terminate prematurely after loading the approval episode, omitting the subsequent revocation.
Technical mechanism
xMemory’s architecture centers on two core mechanisms: hierarchical memory construction with dynamic plasticity, and two-stage adaptive retrieval.
1. Problem Formalization
Let the agent’s interaction history be a chronological sequence of messages . Given a query and a token context budget , the objective is to construct an optimal context satisfying that maximizes response quality while preserving critical relational and temporal evidence structure.
2. Four-Tier Hierarchy and Scale
Memory units are strictly partitioned across four levels:
Original Messages (Message) → Episode Blocks → Semantic Facts → Themes
| Tier Name | Definition | Mapping and Structural Invariants |
|---|---|---|
| Message | Raw dialogue turns | Sequential turns form contiguous blocks mapped to 1 Episode |
| Episode | Temporal window abstractions | Captures chronological summaries; 1 Episode maps to multiple Semantics |
| Semantic | Atomic reusable facts | Each Semantic belongs strictly to exactly 1 Theme, ensuring orthogonality |
| Theme | High-level topic clusters | 1 Theme aggregates multiple semantically aligned Semantic nodes |
In the LoCoMo benchmark setting (Figure 2 caption), a representative dialogue history is organized into approximately 650 Themes, 2,900 Semantics, and 750 Episodes.
3. Sparsity–Semantics Objective and Structural Plasticity
To prevent unconstrained cluster expansion (which causes candidate explosion) or over-fragmentation into isolated clusters, xMemory defines an optimization objective over partition (Section 3.2, Eq. 1–3):
- SparsityScore (Eq. 2): Penalizes disproportionately large or highly skewed theme sizes, encouraging balanced cluster cardinalities and preventing retrieval from collapsing into oversized subgraphs.
- SemScore (Eq. 3): Evaluates distances between theme centroids, penalizing centroids that are too close (semantic redundancy) or excessively distant (semantic islands lacking cross-topic connectivity).
- Split & Merge: When new interactions generate new Semantic nodes, any Theme exceeding a capacity threshold (e.g., a cap of 12 Semantics per Theme) triggers sub-cluster splitting to maximize . Conversely, undersized or redundant themes are merged into neighbors.
- Graph Maintenance: Theme and Semantic nodes continuously update Top- cosine similarity edges, providing the topology for graph-based traversal during retrieval.
4. Two-Stage Adaptive Retrieval
Stage I: Query-Aware Representative Selection
Over the topology, the system applies a greedy selection algorithm to choose representative node set , optimizing trade-offs between graph coverage and query relevance (Section 3.3, Eq. 4):
Where:
- is the set of candidate nodes in the current tier, and is the set of already selected representatives;
- denotes the set of newly covered neighbors outside when node is added;
- represents the edge weight between node and neighbor , and is a normalizer;
- is the semantic similarity between query and candidate node ;
- is a tunable hyperparameter balancing coverage and relevance.
The system first solves for representative Themes, then restricts Stage I to the induced Semantic subgraph. This mechanism naturally facilitates multi-hop reasoning and set-level evidence gathering, pulling together disparate facts scattered across thematic branches.
Stage II: Uncertainty-Adaptive Evidence Inclusion
Following Semantic node selection, the system retrieves backing Episode blocks. Crucial design invariant: Episodes are included as intact units, with no internal word-level pruning.
- The system evaluates marginal information gain by measuring predictive entropy reduction in the reader model conditioned on the candidate context;
- An Episode is incorporated into context only if it produces significant entropy reduction;
- If necessary, the system can selectively expand down to raw Messages;
- When additional nodes yield negligible uncertainty reduction, Early Stopping terminates context expansion.
Unlike RAG pruning methods (e.g., LLMLingua-2) that arbitrarily drop words under independence assumptions, xMemory preserves narrative integrity and chronological prerequisites.
How to read the evidence
The paper evaluates xMemory across multi-turn dialogue QA, personal lifelong memory, backbone generalization, and extensive component ablations.
1. Experimental Setup and Evaluation Dimensions
- Datasets:
- LoCoMo: 50 long-horizon dialogues averaging ~9,000 tokens and ~300 turns across up to 35 sessions, evaluated on Single-hop, Multi-hop, Temporal, and Open-domain questions.
- PerLTQA: Personal lifelong memory benchmark assessing profiles, interpersonal relationships, and chronological events with sentence-level answers.
- Baselines:
- Naive RAG: Top-20 raw dialogue chunks retrieved via cosine similarity;
- Structured Memory Systems: A-Mem, MemoryOS, Nemori;
- Pruning Baselines: LightMem (integrating LLMLingua-2 compression).
- Backbone LLMs:
- Open models: Qwen3-8B, Llama-3.1-8B-Instruct;
- Proprietary models: GPT-5 nano.
- Compute, Embeddings, and Metrics:
- Embeddings generated via
text-embedding-3-small; - Greedy decoding () for response generation on A100 80GB GPUs;
- Evaluated using BLEU-1, Token F1, ROUGE-L, and Token/query (context efficiency).
- Embeddings generated via
2. Main Results Analysis
Table 1: LoCoMo Long-Dialogue Benchmark
Under the Qwen3-8B backbone, xMemory demonstrates pronounced advantages (excerpted from Table 1):
| Method | Avg F1 | Avg BLEU | Temporal F1 | Multi-hop F1 | Token/query |
|---|---|---|---|---|---|
| Naive RAG | 40.45 | 28.51 | 32.14 | 17.01 | 7754.66 |
| Nemori | 40.45 | 28.51 | 33.74 | 18.25 | — |
| LightMem | 30.28 | 23.77 | 26.50 | 12.30 | 5545.35 |
| A-Mem | 21.78 | 19.49 | 19.20 | 11.15 | 9103.46 |
| MemoryOS | 33.76 | 29.20 | 28.40 | 14.80 | 7234.66 |
| xMemory (Paper) | 43.98 | 34.48 | 37.46 | 20.69 | 4711.29 |
Key findings:
- Temporal Reasoning Gains: On questions requiring temporal prerequisite resolution, xMemory reaches an F1 of 37.46, exceeding Nemori (33.74) by 3.72 points and Naive RAG by 5.32 points. BLEU-1 jumps from Nemori’s 23.60 to 29.58.
- Multi-Hop Association: In questions demanding cross-session synthesis, xMemory scores 20.69 F1, substantially outperforming Naive RAG (17.01) and A-Mem (11.15).
- Substantial Context Efficiency: xMemory consumes only 4711.29 tokens per query—a 48% reduction compared to A-Mem (9103.46) and a 39% reduction compared to Naive RAG (7754.66). Accuracy improvements stem from precise structural filtering rather than context stuffing.
- Cross-Model Consistency:
- On GPT-5 nano: xMemory scores Avg F1 50.00 (vs. Nemori’s 48.17) while reducing tokens from 9155 to 6581;
- On Llama-3.1-8B-Instruct: xMemory achieves Avg F1 34.77, BLEU 24.73, and tokens 5539.97, sustaining top average performance across all tested backbones.
Table 2: PerLTQA Personal Lifelong Memory
On the PerLTQA dataset under Qwen3-8B (Table 2):
| Method | BLEU-1 | Token F1 | ROUGE-L | Token/query |
|---|---|---|---|---|
| Naive RAG | 32.08 | 41.37 | 35.95 | 6274.38 |
| MemoryOS | 35.14 | 42.35 | 38.48 | 6499.47 |
| xMemory (Paper) | 36.24 | 47.08 | 42.50 | 5087.18 |
Under Llama-3.1-8B, xMemory reaches an F1 of 52.37. In stark contrast, LightMem collapses to a BLEU of 23.47 and F1 of 35.93, as word-level pruning severs biographical facts and event linkages. This confirms xMemory’s design principles transfer effectively to lifelong personal memory domains.
3. Ablation and Diagnostic Analysis (Figure 3–5, Table 3)

Figure 3, Section 5 of the paper (component ablation analysis): left axis tracks BLEU and F1 scores while right axis shows token usage per query; comparing Naive RAG, Memory-only, w/o Stage II, w/o Split & Merge, and full xMemory to isolate the contribution of each architectural layer. See the original Figure 3 anchor and arXiv HTML figure endpoint. arXiv source identifies a perpetual non-exclusive license; reproduced under arXiv reuse terms with attribution.
- Figure 3 Architecture Ablations (LoCoMo, Qwen3-8B):
- Naive RAG (raw chunks): Low accuracy and inflated token costs;
- Memory-only (static hierarchy without adaptive retrieval): Shows structural hierarchy alone cannot cure retrieval collapse;
- w/o Stage II (no uncertainty-based inclusion): Marked F1 degradation, demonstrating the importance of entropy-based episode gating;
- w/o Split & Merge (frozen clusters): Downstream performance consistently falls;
- Full xMemory: Optimal performance across all metrics.
- Figure 4 Evidence Hit Distribution: xMemory produces a higher proportion of Multi-hit blocks on multi-fact questions, whereas pruning baselines concentrate heavily on 1-hit blocks, indicating evidence fragmentation.
- Figure 5 Structural Plasticity: Disabling retroactive restructuring during memory ingestion leads to compounding degradation over successive sessions.
- Table 3 Coverage Efficiency: xMemory achieves superior golden evidence coverage while requiring significantly fewer tokens.
4. Method Comparison with Related Paradigms
| Paradigm | Representatives | Architectural Trait | Critical Flaw in Agent Memory |
|---|---|---|---|
| Flat Context | MemGPT, MemoryOS | Paging or FIFO dialogue queues | Operates on raw text; cannot prevent redundant collapse |
| Structured Memory | MemoryBank, Zep, A-Mem | Entity graphs or memory cards | Triggers unguided cross-layer expansion during queries |
| RAG Pruning | LightMem + LLMLingua-2 | Token importance pruning | Assumes passage independence; fragments conversational prerequisites |
| xMemory (Paper) | xMemory | Decouple-then-aggregate four-tier topology | Dynamic write-time restructuring, adaptive two-stage expansion |
Evidence map
To establish rigorous boundaries between empirical findings and interpretations, the conclusions are organized into four explicit tiers:
Direct paper evidence
- Benchmark Scores and Context Efficiency: Across LoCoMo and PerLTQA, xMemory achieves superior F1, BLEU, and ROUGE metrics across Qwen3-8B, Llama-3.1-8B, and GPT-5 nano compared to Naive RAG, A-Mem, MemoryOS, LightMem, and Nemori;
- Temporal and Multi-Hop Strengths: On LoCoMo’s temporal split, xMemory records 37.46 F1 (vs. Nemori’s 33.74); on multi-hop questions, it achieves 20.69 F1 (vs. Naive RAG’s 17.01);
- Inference Token Compression: xMemory lowers per-query context tokens by ~39% relative to Naive RAG and ~48% relative to A-Mem while improving QA accuracy;
- Component Ablations: Figures 3, 4, and 5 and Table 3 establish that removing Split & Merge, disabling Stage II uncertainty filtering, or freezing the graph hierarchy causes measurable performance drops.
Author causal claim
- Decoupling and Aggregation Eliminates Collapse: The authors attribute the mitigation of redundant collapse directly to elevating the retrieval unit from text chunks to latent semantic components;
- Intact Episodes Preserve Temporal Continuity: The authors assert that refraining from word-level pruning within episodes is the primary causal driver of temporal reasoning gains over LightMem;
- Entropy Reduction is an Optimal Gating Criterion: The authors treat predictive uncertainty reduction as the ideal stopping mechanism for balancing token budgets against retrieval precision.
Unsupported claims
- Feasibility in Real-Time Production Streams: The paper does not provide benchmarks for live concurrent write latency, graph recalculation overhead, or end-to-end service SLAs;
- Structural Stability Under Deletion Mandates: The paper does not analyze compliance with privacy regulations (e.g., GDPR right-to-be-forgotten). The impact of deleting an arbitrary Episode on existing centroids and graph connectivity remains unmeasured;
- Robustness Against Adversarial Poisoning: The authors explicitly exclude LoCoMo’s adversarial subset (Section 4.1). Behavior under deliberate misinformation injection or corrupted interaction logs remains unverified;
- Generalization to Tool-Using Workspace Agents: Experiments are confined to text QA proxies; effectiveness in executing terminal commands or multi-step workspace file modifications is not established;
- Total Cost of Ownership: Table 1 accounts only for inference tokens; computational costs for hierarchy building, summarization LLM calls, embedding generation, and graph maintenance are omitted.
Bloss0m engineering synthesis
- Paradigm Shift in Memory Architecture: xMemory demonstrates that write-time topological structuring is far more impactful than expanding vector retrieval dimensions;
- Multi-Layer Defensive Guardrails: Production systems must wrap clustering layers in data provenance tracking, physical timestamps, and explicit TTL policies;
- Separation of Immutable Logs and Derived Projections: Raw dialogue messages should be stored as append-only event logs, while Episodes, Semantics, and Themes operate as rebuildable derived indices to accommodate schema migrations and deletion requests.
Artifacts and reproducibility
This reading follows the arXiv:2602.02007 preprint (Hu et al.). The authors maintain an official xMemory GitHub repository and academic project page.
Public endpoint status as of 2026-08-09:
- Code Accessibility: The official repository is publicly available under the MIT license, including an
environment.ymlspecification, dataset download references, and execution scripts for LoCoMo construction and evaluation on Llama-3.1-8B using a single A100 80GB GPU. A turn-key, multi-model evaluation harness is not provided. - Model Checkpoints and Releases (Announced but Unavailable): The GitHub Releases section was empty as of this date. While the README mentions pre-built LoCoMo Llama memory snapshots in releases, no files are hosted.
- Benchmark Reproduction Scope:
- Experimental figures cited herein reflect author-reported results;
- Dataset links route to external upstream sources rather than pre-packaged evaluation artifacts;
- Hyperparameters, prompts, entropy decision logic, random seeds, and raw logs for GPT-5 nano and Qwen3-8B remain unpublished. Teams can verify single Llama pipeline mechanics locally, but reproducing every row in Tables 1 and 2 is not directly supported out of the box.
Recommended Minimal Local Verification Path: Teams considering xMemory should avoid attempting a full multi-model benchmark rerun. Instead, select 10 representative multi-session dialogue traces with clear temporal dependencies, run Naive RAG alongside xMemory’s hierarchical retrieval script on an open model (e.g., Llama-3.1-8B), and compare token consumption, answer precision, and redundancy rates directly.
Bloss0m engineering judgment and when not to use it
Drawing on Bloss0m’s engineering experience with autonomous agents and retrieval architectures, the following adoption criteria apply:
When Adoption is Justified
| Operational Context | Recommended Strategy | Engineering Rationale |
|---|---|---|
| Multi-session personal assistant dialogues | Run shadow replay on frozen interaction logs | Closely matches Figure 1 and Table 1 conditions; measure baseline redundancy and temporal breakdown first. |
| Strict context budget constraints | Adopt two-stage top-down retrieval | Table 1 documents ~40% token savings, offering meaningful cost relief under high API pricing. |
| Traceable, versioned memory stores | Implement dynamic hierarchy with rollback support | Figure 5 confirms restructuring is vital for longevity, but requires rebuildable indices. |
When Not to Use It
- Single-turn QA, short documentation, or static corporate knowledge bases: Do not deploy xMemory’s four-tier hierarchy for standard enterprise document retrieval or static FAQs. The overhead of summarization, clustering, and graph maintenance will outweigh retrieval gains; standard RAG with a high-performance reranker is far more economical and reliable.
- Strict regulatory deletion requirements (GDPR / Right to be Forgotten): The paper provides no mechanism for dynamic node eviction. In systems subject to frequent user deletion requests, cascading cluster recalculations present significant stability risks. Do not deploy without custom tombstone mechanisms.
- High-concurrency, unbuffered real-time write streams: Split & Merge operations rely on batch clustering and cannot guarantee millisecond-level write consistency. High-throughput real-time ingest will cause database lock contention and latency spikes.
- Adversarial or prompt-injection-prone environments: Without robust semantic sanitization, adversarial inputs can manipulate clustering algorithms and propagate malicious instructions into broad Theme nodes, corrupting entire memory partitions.
Four Architectural Guardrails for Production
Teams implementing xMemory principles should adopt four core engineering guardrails:
- Immutable Log Principle: Raw messages must be retained in append-only storage, treating Themes and Semantics as ephemeral, rebuildable projections;
- Hard Physical Timestamps: Inject explicit timestamps into Semantic nodes and combine graph traversal with temporal filters to prevent stale memories from overriding current state;
- Entropy Threshold Fallback: Enforce a minimum expansion floor in Stage II to prevent smaller reader models from terminating prematurely due to overconfidence;
- Total Cost of Ownership Monitoring: Track ingestion LLM summarization costs, graph indexing latency, and inference token savings holistically to confirm net operational benefit.
Three things to remember
- Technical idea: Agent memory bottlenecks stem from temporal entanglement and redundant collapse; xMemory decouples raw streams and aggregates them into a four-tier topology, elevating retrieval from flat text matching to multi-scale topological navigation.
- Core evidence: Across long-horizon dialogues on LoCoMo, xMemory achieves top performance across three diverse backbones, notably boosting temporal reasoning F1 while cutting per-query context tokens by 39% to 48%.
- Engineering boundary: Hierarchical memory introduces substantial construction overhead and lacks validation in concurrent real-time write, GDPR deletion, or adversarial contexts; never deploy without data provenance and defensive guardrails.
Primary sources
- arXiv Preprint Record (arXiv:2602.02007): Version history, author roster, abstract, and project links.
- Full Paper HTML (arXiv:2602.02007v1): Figures 1–5, Sections 2–3 methodology, Section 4 experimental data, and Tables 1–3.
- Official xMemory GitHub Repository (HU-xiaobai/xMemory): MIT license, environment specifications, LoCoMo evaluation scripts, and dataset links.
- Official xMemory Project Page: Architectural overview and academic presentation materials.
- arXiv Non-Exclusive License Terms: Open scholarly license terms governing figure reproduction (Figure 2).