← Blog

Engineering note

Pydantic AI v2.45: Durable Agent Reliability Is a Session and Trace Contract

Pydantic AI v2.45: Durable Agent Reliability Is a Session and Trace Contract

Pydantic AI v2.45.0 shipped on September 17, 2026. The release note appears to contain several unrelated items: a new TypeSafeModel, an Amazon Bedrock xhigh effort fix, changes to DynamicToolset and MCP sessions, preserved MCP sampling history, and corrected agent-run usage. The more useful way to read them is through one question: when a durable execution engine splits an agent into retryable units, what counts as the same run?

That question is closer to production reliability than “did one call succeed?” If every unit re-resolves tools, creates a new MCP session, loses the previous tool result, or charges a child agent’s tokens to its parent span, the workflow may still finish. It will be much harder to replay, explain cost, or tell whether recovery continued the original state.

This article separates three layers: behavior explicitly promised by the release and its PRs, engineering meaning inferred from those changes, and limitations that still need verification in your own engine/provider combination. It is not a complete Pydantic AI tutorial, and it does not turn maintainer tests into a cross-environment SLA.

What the release actually changed

The official v2.45.0 release note groups naturally into a new interface and a set of durable-execution fixes. The distinction matters: TypeSafeModel expands the model surface, while the later changes repair how the agent runtime carries state and telemetry across units.

New interface: TypeSafeModel is a typed decision model, not a chat LLM

TypeSafeModel brings TypeSafe’s Jev into Pydantic AI. PR #8450 is explicit: Jev receives text and typed questions, then answers each field with probabilities; it does not generate long-form text, read files, or fill arbitrary tool arguments. The fields in output_type are therefore closer to calibrated decision questions than to a free-form generation schema.

There are two practical uses in an agent runtime: put low-latency triage, routing, or rubric decisions in a decision layer; then use a general language model or FallbackModel when the task needs open-ended answers or parameterized tool calls. A correct output type does not grant tool permission or replace a policy engine. Jev confidence is also a model signal, not proof of authorization.

Durable execution: correcting the unit boundary to a run boundary

The four durability-related fixes share one misalignment: the durable engine records, retries, and replays units, while some runtime resources should live for the whole run.

Changev2.45.0 behaviorEngineering meaning
DynamicToolsetWith per_run_step=False, resolve once per durable run; enter on the first unit that needs it, then close at run endFactories, toolset caches, and connections are not rebuilt for every step
MCP sessionApplicable DBOS and Prefect paths hold one session per durable runinitialize, tools/list, and server-side caches do not cold-start at every unit
MCP tool historyMCPSamplingModel keeps native tool_use/tool_result blocks, IDs, arguments, results, and retry feedbackA continuation or server/client round trip can see the previous tool exchange
Agent-run usage spanEach run span is attributed with requests made by that run, rather than shared or nested usageParent and delegate costs can be added without duplicate attribution

“Once per run” is not an unconditional global guarantee. per_run_step=True still asks DynamicToolset to resolve per unit. When the run context must cross a process boundary, such as Temporal worker/activity boundaries, the engine may not see the resolved toolset and can fall back to per-unit behavior. MCP engine lifecycles also differ: PR #8463 deliberately keeps Temporal at per-activity sessions because an activity can execute on another worker.

Why session lifetime changes reliability

1. Toolset resolution determines what a replay is replaying

Under the old mismatch, a DynamicToolset factory could run again inside every durable unit. If it returned an MCP toolset, the connection and cache_tools were rebuilt too. The second model request in one run could see a copy of the first unit’s environment rather than a continuing tool session.

PR #8455 separates object resolution from session entry, which performs I/O: resolve the factory once in container code; enter on the first durable unit that actually needs the tool; hold it for the run; close it at the end. The durable engine keeps its deterministic unit sequence, while connection failure returns to a retryable unit.

That design adds a factory contract worth putting into migration review: a factory executed in container code must deterministically build a toolset from run dependencies. It cannot quietly connect, read transient unit-only state, or rely on an unreplayable side effect. per_run_step=True remains an explicit choice when parallel tool-call units should not share mutable toolset state.

2. Session lifetime changes MCP network and cache cost

An MCP client obtains tool definitions before calling a tool; the server may also retain initialization state and cache_tools. PR #8455 measured three model requests and two tool calls against an in-process server: a dynamic toolset took five initialize and five tools/list calls before the fix, then one of each after it. PR #8463 applies the run-held lifecycle to the static MCPToolset path for DBOS and Prefect; discovery remains a recorded unit rather than an opaque process-local cache.

This is not only a performance tweak. Rebuilding a session adds connection, authentication, schema-discovery, and server warm-up failure points. A longer-lived session also means credentials, tenant scope, server-side state, and concurrency must remain correct within the run. If tool access can be revoked mid-run, the team cannot extend a session indefinitely just to save round trips.

3. Tool history determines whether a continuation is actually a continuation

The MCPSamplingModel fix does not flatten history into plain text. PR #8466 preserves native MCP tool_use and tool_result blocks, including original tool IDs, arguments, results, and retry feedback. Parallel results remain in a separate user message. After a durable restart or a server/client round trip, the next model can identify which tool was called, what it returned, and whether a retry was requested.

The format boundary is explicit: tool history requires the MCP 2025-11-25 sampling format, while new tool execution and multimodal tool results remain unsupported. Migration therefore needs a version matrix across the Python package, client, server, sampling adapter, and persisted message history—not only a package upgrade.

Usage spans: observability must respect the run boundary too

Session lifetime fixes the problem of execution state being cut apart; usage attribution fixes the problem of after-the-fact telemetry charging the wrong run. PR #8456 addresses two cases: a later run carrying a previous RunUsage object, and a parent agent handing the same usage object to concurrent delegates. The old behavior could expose a conversation total on a later run or count sibling tokens inside each delegate span.

v2.45.0 attributes usage to the run that actually made the request while its agent-run span is open, separating nested and concurrent tasks through context attribution. result.usage, UsageLimits, and per-request chat spans are unchanged; durable activities still return usage_delta to the run span. This is an observability-contract correction, not a new definition of provider billing.

Operationally, aggregation needs a rule: when the backend sums parent and child span attributes, sum the outermost agent-run spans rather than every nested span. Otherwise the same tokens are counted twice. That is why “each run reports its own usage” must ship with trace-topology guidance, not only a renamed field.

Bedrock and TypeSafeModel: adjacent fixes, separate promises

v2.45.0 also contains two provider changes that are easy to blend into the same release summary.

First, Bedrock Converse no longer hardcodes thinking='xhigh' to effort='max'. Per PR #8392, BedrockConverseModel reads the merged model profile: when the model supports xhigh, the original value is forwarded; otherwise it still maps to Bedrock’s accepted max. The PR names Opus 4.7, Opus 4.8, and Sonnet 5 as examples that accept xhigh, while older models keep the fallback.

This improves request semantics and provider compatibility; it does not prove that xhigh produces better reasoning than max. The PR’s live check confirms that Bedrock accepts the value, not how the two settings differ in inference quality. Production traces should still record the resolved model profile, thinking level, provider error, and fallback path.

Second, TypeSafeModel lets an agent use Jev as a typed decision provider, but it does not automatically share a lifecycle with durable sessions. If Jev is placed inside a durable workflow, the application still has to define whether the decision is replayable, version confidence thresholds, record fallback-to-LLM behavior, and include provider details in the trace. A new provider does not complete the runtime contract for you.

Migration checklist: verify these five things first

1. Inventory toolset factory I/O and scope

Find every DynamicToolset, record per_run_step, and verify that per_run_step=False factories perform deterministic object construction only. Put user, tenant, and credential scope into run dependencies; do not rely on state that merely happens to exist inside one unit.

2. Add engine-specific wire-count and lifecycle tests

Measure initialize, tools/list, tools/call, enter, close, and retry counts for one run. DBOS and Prefect should confirm the applicable path reaches one session per run; Temporal should confirm that per-activity sessions are an intentional cross-worker design rather than a regression.

3. Build a message-history version matrix

Use real tool calls, results, parallel results, and retry feedback to test MCPSamplingModel across MCP 2025-11-25 and older SDK combinations. Treat unsupported multimodal tool results as an explicit failure mode rather than silently downgrading them to text that looks complete.

4. Reconcile usage aggregation and budget dashboards

Check whether the logging backend aggregates chat spans, agent-run spans, or both. Add fixtures for three-level nesting, parallel delegates, usage reuse across conversations, and durable usage_delta; verify that result.usage, UsageLimits, and span totals agree.

5. Make provider profiles and fallbacks trace fields

Bedrock xhigh, Jev confidence, fallback model, MCP protocol version, and engine name all affect the meaning of a run. Without those fields, “completed” or “failed” is not enough to explain changes in cost, latency, or quality.

Failure modes v2.45.0 does not solve automatically

Failure modeWhat v2.45.0 helps withStill owned by the application/platform
Unit retry creates a cold MCP sessionApplicable engines hold the toolset/session for the runCredential rotation, tenant isolation, server timeouts, concurrency limits
Factory has an external side effect during replayDocs and PRs state the deterministic-factory contractCode review, replay tests, and banning container-side I/O
Tool history loses a retry or resultNative history blocks and IDs are preservedProtocol version, schema migration, and unsupported-content policy
Parent/delegate costs are counted twiceRun spans report their own usageChoosing outermost spans and configuring backend aggregation
Bedrock rejects the requested effortModel profile selects xhigh or falls back to maxModel availability, quality comparison, cost, and latency SLOs
Typed decision has the right shape but the wrong meaningTypeSafeModel provides a typed decision interfaceGolden sets, calibration, human escalation, permissions, and side-effect gates

If the goal is a replayable, inspectable agent runtime, start with the state, tools, and observability layers in the AI Agent guide, then use the Agentic AI platform contract to turn trace, policy, and evaluation into release criteria. For MCP identity and connection attribution, compare the multi-user Forge MCP auth runtime; for decision gates inside an agent loop, read the Jev confidence-gated runtime.

Closing: treat v2.45 as a lifecycle-contract reminder

The value of Pydantic AI v2.45.0 is not that every changelog item can independently claim “more reliable.” It is that several details commonly handled in isolation now line up: toolset resolution fixes when the tool environment becomes stable, MCP session lifetime fixes how long the external connection lives, tool history fixes whether a restart can understand the previous exchange, and usage spans fix whether operations can charge the right run.

The limitations are part of the story. Temporal’s worker boundary, the MCP sampling version, Bedrock profile capabilities, and Jev’s typed-decision scope still require an application-specific compatibility matrix. After upgrading, the most valuable test is not another happy path. Intentionally interrupt a run, retry a tool, and fan out two delegates, then answer three questions: Which session did it use? Which history did it see? Whose usage did it record?

Sources and further reading

For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.

Speaking & contact