Transport-free public Memory facade.
Memory is the surface platform users interact with. It delegates the six high-level workflow operations to an injected MemoryRuntime and the CRUD conveniences to an injected MemoryCrudClient. Identity is resolved per call via resolve_identity; tenancy never appears in any public signature.
Memory
class Memory()
Public, transport-free memory facade.
__init__
def __init__(*, runtime: MemoryRuntime | None = None, client: MemoryCrudClient | None = None, _bound_ctx: MemoryRequestContext | None = None, service_account_token: str | None = None, api_key: str | None = None, base_url: str | None = None, project_id: str | None = None) -> None
Construct a Memory facade.
Pass the inputs directly. Each one also falls back to an AGENTIC_MEMORY_* environment variable when the argument is omitted, so an explicit argument always wins over the environment.
service_account_token(AGENTIC_MEMORY_SERVICE_ACCOUNT_TOKEN) is auth — a service-account access token minted from a service account’s client ID and secret atPOST /api/v1/oauth/token(the README’s hosted guide walks through it). It becomes a bearer token when set, and there is no auth header when it is absent.base_url(AGENTIC_MEMORY_BASE_URL) is the backend address.project_id(AGENTIC_MEMORY_PROJECT_ID) selects the route shape. Set it to use project-scoped routes (/api/v1/projects/{id}/memory/*). Leave it empty to use flat routes (/api/v1/memory/*). Auth does not select the route shape, so an authenticated caller with noproject_idstill uses flat routes.
api_key (AGENTIC_MEMORY_API_KEY) is a deprecated alias for service_account_token: still accepted, emits a DeprecationWarning. Setting both the new and the legacy input — in any mix of argument and environment variable — raises ValueError. A blank environment variable counts as unset (matching base_url and project_id), so an empty-exported placeholder never conflicts with an explicit argument. A blank argument still raises ValueError.
base_url resolves from the argument, then AGENTIC_MEMORY_BASE_URL, then the hosted default when an auth token is set. With neither a resolvable URL nor an auth token, construction raises ValueError.
Type-specific CRUD is available on hosted (Gateway/project_id set) and direct (OE/flat) HTTP modes. Per-backend gaps (including app-bound) are documented in docs/capability-matrix.md and raise MemoryNotSupportedError when hit. Application code does not construct app-bound mode; the platform supplies runtime / client injection instead.
bind
def bind(ctx: MemoryRequestContext) -> Memory
Return a new Memory scoped to ctx without mutating this handle.
The bound handle shares this handle’s transport; closing either one closes the shared transport. Ownership propagates, so a bound handle of a self-constructed (api-key) Memory still releases the transport on close(), while a bound handle of an injected runtime leaves it to its owner.
close
def close() -> None
Release the underlying transport(s) when this handle owns them.
A runtime this Memory constructed — and, in direct mode, the CRUD adapter alongside it — is closed; an injected or bound runtime belongs to its provider and is left untouched.
record_turn
def record_turn(*, role: str, content: str | None = None, session_id: str | None = None, tool_calls: list[dict[str, Any]] | None = None, tool_call_id: str | None = None, tool_name: str | None = None, is_error: bool = False, model_name: str | None = None, metadata: dict[str, Any] | None = None, idempotency_key: str | None = None) -> WriteTurnResult
Record a single conversation turn (write-accepted semantics).
Arguments:
metadata- Arbitrary key-value metadata persisted on the turn and returned on retrieval in its ownchunk.metadata["metadata"]slot, kept apart from the platform’s turn fields, so any key is safe to use. Matchable by the retrievalmetadata_filter, where the key is the qualified index path (metadata.<name>) because the filter addresses the stored document. Omitting it preserves prior behavior.idempotency_key- Deduplication key for the write. Auto-generated per call when omitted, which protects transport-level retries beneath this call only; callers that retryrecord_turnitself must supply their own key for cross-call dedupe.
Notes:
In app-bound mode the returned WriteTurnResult.id is "" and turn_seq is 0 — the platform runtime fires the write asynchronously and returns no per-turn identifier or sequence number.
save_semantic
def save_semantic(*, text: str, label: str, user_id: str | None = None, source: str = "agent", visibility: str = "private", metadata: dict[str, Any] | None = None, upsert: bool = True, agent_id: str | None = None) -> CreateSemanticResult
Save a semantic memory (fact, customer profile).
Arguments:
visibility- Access scope of the memory (“private” or “org”).metadata- Optional caller-supplied metadata dictionary stored with the memory.upsert- When True (the default), replace an existing memory with the same label instead of creating a duplicate.
Raises:
MemoryIdentityError- Ifuser_idis not supplied and cannot be resolved from the bound context.MemoryNotSupportedError- If this handle has no CRUD client configured.
search_semantic
def search_semantic(query: str, user_id: str | None = None, visibility: str | None = None, top_k: int = 50) -> list[MemoryChunk]
Search semantic memories.
Arguments:
visibility- When supplied, the ambient (runtime-context) principal is not auto-substituted foruser_id; the caller must supply one explicitly if a user-scoped filter is needed.
get_semantic
def get_semantic(label: str, user_id: str | None = None, visibility: str | None = None) -> Any | None
Get a semantic memory by label.
Raises:
MemoryNotSupportedError- If this handle has no CRUD client configured.
save_episode
def save_episode(*, title: str, content: str, user_id: str | None = None, summary: str | None = None, session_id: str | None = None, participants: list[str] | None = None, tags: list[str] | None = None, visibility: str = "private", metadata: dict[str, Any] | None = None, agent_id: str | None = None) -> CreateEpisodicResult
Save an episodic memory (conversation summary).
Arguments:
visibility- Access scope of the memory (“private” or “org”).metadata- Optional caller-supplied metadata dictionary stored with the episode.
Raises:
MemoryIdentityError- Ifuser_idorsession_idis not supplied and cannot be resolved from the bound context.MemoryNotSupportedError- If this handle has no CRUD client configured.
search_episodes
def search_episodes(query: str, user_id: str | None = None, visibility: str | None = None, session_id: str | None = None, top_k: int = 50) -> list[MemoryChunk]
Search past episodes.
Arguments:
visibility- When supplied, the ambient (runtime-context) principal is not auto-substituted foruser_id; the caller must supply one explicitly if a user-scoped filter is needed.session_id- Episodic memory is stored without a session, so a bound or ambientsession_idis not inherited here. Pass one explicitly to filter to a single conversation.
list_episodes
def list_episodes(user_id: str | None = None, visibility: str | None = None, session_id: str | None = None, limit: int = 20) -> list[Any]
List episodic memories with optional filters.
Raises:
MemoryNotSupportedError- If this handle has no CRUD client configured.
save_taxonomic
def save_taxonomic( *, domain: str, term: str, definition: str, related_terms: list[str] | None = None, visibility: str = "org", user_id: str | None = None, metadata: dict[str, Any] | None = None) -> CreateTaxonomicResult
Save a taxonomic term (knowledge-base entry).
Arguments:
visibility- Access scope of the entry (“org” by default so the term is shared knowledge; “private” restricts it to the user).metadata- Arbitrary key-value metadata persisted on the entry and returned on retrieval; matchable by the retrievalmetadata_filter. Omitting it preserves prior behavior.
Raises:
MemoryIdentityError- Ifuser_idis not supplied and cannot be resolved from the bound context.MemoryNotSupportedError- If this handle has no CRUD client configured.
search_taxonomic
def search_taxonomic(query: str, user_id: str | None = None, domain: str | None = None, visibility: str | None = None, top_k: int = 50) -> list[MemoryChunk]
Search the knowledge base.
Arguments:
visibility- When supplied, the ambient (runtime-context) principal is not auto-substituted foruser_id; the caller must supply one explicitly if a user-scoped filter is needed.
get_taxonomic_term
def get_taxonomic_term(domain: str, term: str, user_id: str | None = None, visibility: str | None = None) -> Any | None
Get a specific knowledge-base term.
Raises:
MemoryNotSupportedError- If this handle has no CRUD client configured.
list_domains
def list_domains(visibility: str | None = None) -> list[str]
List all knowledge-base domains.
Raises:
MemoryNotSupportedError- If this handle has no CRUD client configured.
discover_procedures
def discover_procedures( query: str, user_id: str | None = None, *, visibility: str | None = None, tags: list[str] | None = None, top_k: int = 10, similarity_threshold: float = 0.0, metadata_filter: dict[str, Any] | None = None) -> list[dict[str, Any]]
Discover procedures matching a query without loading full content.
Arguments:
metadata_filter- Field-level filter on procedural metadata. Keys are fully qualified index paths, so an attribute declared astierin the project’s metadata partition configuration is filtered as{"metadata.tier": "gold"}; nothing is prefixed for you. Keys here are not validated against the declared set – an undeclared or misspelled path is not rejected and silently yields no or degraded matches.
get_procedure
def get_procedure(procedure_name: str, *, user_id: str | None = None, visibility: str | None = None, include_deleted: bool = False) -> Any | None
Get a full procedural memory by procedure name.
Raises:
MemoryNotSupportedError- If this handle has no CRUD client configured.
save_procedure
def save_procedure(*, procedure: str, description: str, content: str, user_id: str | None = None, steps: list[dict[str, Any]] | None = None, resources: list[dict[str, Any]] | None = None, allowed_tools: list[str] | None = None, compatibility: str | None = None, license: str | None = None, trigger_conditions: list[str] | None = None, tags: list[str] | None = None, visibility: str = "private", agent_id: str | None = None, extraction_source: str | None = None, source_format: str | None = None, source_path: str | None = None, metadata: dict[str, Any] | None = None, update_existing: bool = False) -> CreateProceduralResult
Create or update a procedural memory.
Arguments:
visibility- Access scope of the procedure (“private” or “org”).metadata- Arbitrary key-value metadata persisted on the procedure and returned on retrieval; matchable by the retrievalmetadata_filter. Applies to both create and (in app-bound mode)update_existing. Omitting it preserves prior behavior.update_existing- When True, update the procedure with the same name instead of creating a duplicate. Honored only in app-bound mode; the HTTP backends serve a create-only route, so Gateway and OE modes raiseMemoryNotSupportedError.
Raises:
MemoryIdentityError- Ifuser_idis not supplied and cannot be resolved from the bound context.MemoryNotSupportedError- If this handle has no CRUD client configured, or ifupdate_existingis True against an HTTP backend.
save
def save( memory_type: str, content: str, *, tags: Mapping[str, Any] | None = None, contextual_metadata: dict[str, Any] | None = None ) -> CustomMemorySaveResult
Save a memory of a declared custom type.
The platform stamps identity (org, project, user) and enforces the type’s declared tag schema; this method validates only tag syntax client-side so mistakes fail fast with server-matching messages.
Arguments:
memory_type- A custom type declared in the project’s memory configuration. Built-in types (semantic, episodic, taxonomic, procedural) are rejected — use their dedicated methods.content- The content to store and embed.tags- Declared tag keys with scalar values. Dotted keys (“profile.location”) and one-level nested mappings are equivalent.contextual_metadata- Arbitrary additional context; stored as-is, never filterable.
Raises:
MemoryClientError- If the type name is built-in/empty or a tag fails syntax checks (aValueErrorsubclass).MemoryNotSupportedError- If this handle has no CRUD client, or the platform does not support custom memory types.
retrieve
def retrieve(memory_type: str, query: str, *, tags: Mapping[str, Any] | None = None, top_k: int = 10) -> CustomMemoryRetrieveResult
Retrieve memories of a declared custom type by semantic query.
Filters are exact-match equality on declared tag keys; result ordering may improve between releases and is not contractual.
Arguments:
memory_type- A custom type declared in the project’s memory configuration; one type per call.query- Query text for semantic search.tags- Equality filters on declared tag keys.top_k- Maximum number of results (server enforces 1-100).
Raises:
MemoryClientError- If the type name is built-in/empty or a tag fails syntax checks (aValueErrorsubclass).MemoryNotSupportedError- If this handle has no CRUD client, or the platform does not support custom memory types.
build_context
def build_context(query: str, user_id: str | None = None, session_id: str | None = None, visibility: str | None = None, metadata_filter: dict[str, Any] | None = None, enabled_sources: set[str] | None = None, *, thread_id: str | None = None, max_tokens: int | None = None, format_style: FormatStyle | str | None = None, include_memories: bool = False) -> ContextResponse
Build a unified memory context across memory types.
Omitted enabled_sources defaults to episodic and semantic; pass {"stm", "episodic", "semantic"} to include recent record_turn content.
thread_id is a deprecated alias for session_id; when both are provided, a non-blank session_id wins.
Arguments:
enabled_sources- Memory sources to include.Noneuses the server default (episodic, semantic).metadata_filter- One field-level filter applied to every enabled source. Keys are fully qualified index paths, so an attribute declared astierin the project’s metadata partition configuration is filtered as{"metadata.tier": "gold"}; nothing is prefixed for you. Unlikebuild_context_from_sources, these keys are not validated against the declared set – an undeclared or misspelled path is not rejected and silently yields no or degraded matches. Usebuild_context_from_sourceswhen you want a bad key to fail loudly.visibility- When supplied, the ambient (runtime-context) principal is not auto-substituted foruser_id; the caller must supply one explicitly if a user-scoped filter is needed.max_tokens- Optional gross context-construction budget. Omitted (None) preserves prior behavior. Non-positive, non-integer, non-finite, and bool values raiseValueErrorbefore runtime delegation. After retrieval and ranking, the server subtracts a 500-token formatting reserve, then greedily selects whole memory chunks that fit in the remainder. Positive values at or below 500 leave no budget for memories. Values above 500 can still yield empty context when no chunk fits.metadata.token_countreports formatted output only and excludes the reserve.format_style- Output format for the assembled context:"openai"(a chat-message list that keeps STM turn roles),"claude"(an XML string), or"jinja2"(a markdown string). Unknown values raiseValueErrorbefore any request is made. Omitted (None), the server infers the format from its configured model. Caveat: the server currently re-infers from its configured model when the explicit value matches its default model type, so an explicit"openai"is honored verbatim only on servers with an OpenAI-family model configured;"claude"and"jinja2"are always honored.include_memories- WhenTrue, the response’sselected_memoriescarries the post-budgetMemoryChunklist alongsideformatted_context, so callers can inspect exactly which memories were selected.
build_context_from_sources
def build_context_from_sources( query: str, sources: Sequence[SourceSpec], *, user_id: str | None = None, session_id: str | None = None, visibility: str | None = None, rerank: bool = False, thread_id: str | None = None, max_tokens: int | None = None, format_style: FormatStyle | str | None = None, include_memories: bool = False) -> ContextResponse
Build a memory context from an explicit, per-source-configured source set.
Unlike :meth:build_context, each source in sources declares its own retrieval mode, metadata filter, and candidate count. Results are merged and de-duplicated across sources, optionally reordered by a relevance rerank, then formatted and budgeted like :meth:build_context. metadata.ranking_strategy and metadata.source_outcomes report how the final order was produced and how each source fared.
thread_id is a deprecated alias for session_id; when both are provided, a non-blank session_id wins. session_id is only needed when the stm source is requested.
Arguments:
sources- One :class:SourceSpecper memory source to search. Must be non-empty, and each source may appear at most once.rerank- WhenTrue, reorder merged results by a relevance rerank if the service is available.max_tokens- Optional gross context-construction budget; see :meth:build_context.format_style- Output format for the assembled context; see :meth:build_context.include_memories- WhenTrue, the response’sselected_memoriescarries the post-budgetMemoryChunklist; see :meth:build_context.
Raises:
ValueError- Ifsourcesis empty or lists a source more than once, or ifmax_tokensis non-positive.MemoryIdentityError- If astmsource is requested but no non-blanksession_idresolves from the argument or execution context.
search
def search(query: str, *, sources: SearchSource | str | Sequence[SearchSource | str] | None = None, top_k: int = 10, user_id: str | None = None, visibility: str | None = None, session_id: str | None = None, domain: str | None = None, tags: list[str] | None = None, similarity_threshold: float = 0.0, metadata_filter: dict[str, Any] | None = None) -> list[MemoryChunk]
Search one or more memory sources and return a single ranked list.
sources defaults to [semantic, episodic], the two ranked similarity sources. It accepts a single source ("semantic" or SearchSource.SEMANTIC) or a sequence of them. Requesting a source the current backend does not serve raises MemoryNotSupportedError; see docs/capability-matrix.md for current per-backend gaps. Per-source arguments are forwarded only to the sources that use them: session_id (episodic), domain (taxonomic), and tags / similarity_threshold / metadata_filter (procedural). Results are merged, sorted by similarity_score (descending, unscored last) and truncated to top_k.
metadata_filter keys are fully qualified index paths – an attribute declared as tier is filtered as {"metadata.tier": "gold"}, and nothing is prefixed for you. They are not validated against the declared set, so an undeclared or misspelled path is not rejected and silently yields no or degraded matches.
Raises:
MemoryNotSupportedError- If a requested source is unavailable in this mode. Seedocs/capability-matrix.mdfor current gaps.
Seam primitives: the immutable request context and the runtime Protocol.
MemoryRequestContext
class MemoryRequestContext(BaseModel)
Immutable identity context carried across a memory call.
MemoryRuntime
class MemoryRuntime(Protocol)
Transport seam every memory backend implements.
Identity is passed as resolved keyword arguments; tenancy is internal to each runtime.
AmbientIdentityRuntime
class AmbientIdentityRuntime(Protocol)
Optional capability: a runtime that can supply per-call ambient identity.
Not all runtimes carry ambient identity — only app-bound runtimes running inside the platform stack have access to per-request contextvars. This is a separate Protocol (not merged into MemoryRuntime) so the facade can detect the capability without requiring every runtime to implement it.
MemoryCrudClient
class MemoryCrudClient(Protocol)
CRUD seam the facade’s type-specific conveniences delegate to.
Like MemoryRuntime, tenancy is internal to each implementation: the method shapes mirror MemoryClient with org_id/project_id stripped, so the raw HTTP client never satisfies this Protocol — each transport wires in an adapter that supplies tenancy itself.
create_custom
def create_custom( *, memory_type: str, content: str, tags: Mapping[str, Any] | None = None, contextual_metadata: dict[str, Any] | None = None ) -> CustomMemorySaveResult
Save a custom-type memory. Identity is stamped by the platform.
retrieve_custom
def retrieve_custom(*, memory_type: str, query: str, tags: Mapping[str, Any] | None = None, top_k: int = 10) -> CustomMemoryRetrieveResult
Retrieve custom-type memories. Identity is stamped by the platform.
Exceptions raised by the memory seam.
MemoryIdentityError
class MemoryIdentityError(ValueError)
Raised when a required identity dimension cannot be resolved.
MemoryClientError
class MemoryClientError(ValueError)
Base for pre-HTTP client/usage errors (no status code or response body).
Subclasses ValueError so callers that guard client-side validation with a plain except ValueError keep working.
MemoryNotSupportedError
class MemoryNotSupportedError(MemoryClientError)
Raised when an operation is unavailable in the active transport mode.
This is a client-side capability check raised before any HTTP request is attempted, so it carries no status code or response body.
MemoryAPIError
class MemoryAPIError(Exception)
Base for typed transport errors from the HTTP memory transports.
MemoryAuthError
class MemoryAuthError(MemoryAPIError)
Raised for 401 / 403 responses.
MemoryNotProvisionedError
class MemoryNotProvisionedError(MemoryAPIError)
Raised when the project’s memory runtime is not yet reachable.
MemoryBadRequestError
class MemoryBadRequestError(MemoryAPIError)
Raised for 4xx responses other than auth / not-provisioned.
MemoryRouteNotFoundError
class MemoryRouteNotFoundError(MemoryBadRequestError)
Raised when a core-loop request 404s, hinting a route-shape mismatch.
The SDK selects the route shape from project_id presence (set => project-scoped Gateway routes; empty => flat OE routes). A 404 on a core-loop POST most often means that shape does not match the backend the base_url points at, so this carries a directional, actionable hint rather than the opaque 404. It subclasses MemoryBadRequestError so existing 4xx handling still catches it.
MemoryServerError
class MemoryServerError(MemoryAPIError)
Raised for 5xx responses, and for success responses whose body is unparseable or has an unexpected shape.
MemoryConnectionError
class MemoryConnectionError(MemoryAPIError)
Raised when the transport fails to reach the memory backend (connect/timeout).
Canonical Pydantic models for Memory API.
These models define the request/response types for the Memory Server HTTP API and are used by both MemoryClient and the server routes.
Single source of truth for Memory API types.
WriteTurnResult
class WriteTurnResult(BaseModel)
Result of writing a conversation turn.
CreateSemanticResult
class CreateSemanticResult(BaseModel)
Result of creating a semantic memory.
BulkCreateSemanticResult
class BulkCreateSemanticResult(BaseModel)
Result of bulk creating semantic memories.
CreateEpisodicResult
class CreateEpisodicResult(BaseModel)
Result of creating an episodic memory.
CreateTaxonomicResult
class CreateTaxonomicResult(BaseModel)
Result of creating a taxonomic memory.
CreateProceduralResult
class CreateProceduralResult(BaseModel)
Result of creating a procedural memory.
CreateUserContextResult
class CreateUserContextResult(BaseModel)
Result of creating a user context memory.
CreateSnapshotResult
class CreateSnapshotResult(BaseModel)
Result of creating a snapshot memory.
PromoteSnapshotResult
class PromoteSnapshotResult(BaseModel)
Result of promoting STM to snapshot.
DeleteResult
class DeleteResult(BaseModel)
Result of a delete operation.
InternalStateResult
class InternalStateResult(BaseModel)
Result of updating internal state.
Returns key information for verification: - session_id: Which session was updated - version: New version number (increments on each update) - token_count: Token count of the new content - content_hash: SHA256[:16] for integrity verification
ISGenerationResult
class ISGenerationResult(BaseModel)
Result of automatic IS generation via generate_is_from_turn().
Tracks whether the update happened, was skipped, or coalesced.
MemorySource
class MemorySource(str, Enum)
Source type for memory chunks.
SearchSource
class SearchSource(str, Enum)
A memory source that Memory.search can query.
Mirrors MemorySource minus STM (short-term turns are not a search target). Callers may pass either the enum or its string value.
RetrievalMode
class RetrievalMode(str, Enum)
How a single source is searched in per-source context building.
TEXT is lexical ($search), SEMANTIC is vector ($vectorSearch), and HYBRID fuses both. Callers may pass either the enum or its string value.
FormatStyle
class FormatStyle(str, Enum)
Output format for built context.
Mirrors the server’s format styles: OPENAI is a chat-message list that keeps STM turn roles, CLAUDE an XML string, and JINJA2 a markdown string. Callers may pass either the enum or its string value.
SourceSpec
class SourceSpec(BaseModel)
Per-source retrieval configuration for Memory.build_context_from_sources.
Each enabled source declares its own retrieval mode, metadata filter, and candidate count, unlike build_context where one filter and one mode apply to every source.
A metadata_filter key is the fully qualified index path, not the bare attribute name from the project config. An attribute declared as tier is indexed at metadata.tier, and that is what the filter must say – nothing is prefixed for you.
# project config declares metadata partition keys: tier, score SourceSpec( source="semantic", mode="hybrid", metadata_filter={"metadata.tier": "gold", "metadata.score": {"$gt": 0.8}}, )
Unlike build_context ’s single filter, these keys are validated: an undeclared or misspelled path raises MemoryBadRequestError listing the declared paths, so a typo fails at the call instead of silently narrowing the result set.
MemoryChunk
class MemoryChunk(BaseModel)
Unified representation of memory from any source (STM, episodic, semantic).
This model provides a common interface for memory chunks retrieved from different memory types, enabling uniform processing in the retrieval pipeline.
org_id
def org_id() -> str | None
Get org_id from metadata if available.
SourceOutcome
class SourceOutcome(TypedDict)
One entry of ContextMetadata.source_outcomes — what happened for a single source on the per-source context path.
The wire shape the server emits (per-source builder serializes its internal SourceOutcome dataclass to this dict). Enum-valued fields arrive as their string values. error is None unless the source failed; a requested_mode/effective_mode mismatch flags a source that ran in a degraded mode.
ContextMetadata
class ContextMetadata(BaseModel)
Metadata about the context building process.
Provides information about token usage, memory counts, and timing.
ContextResponse
class ContextResponse(BaseModel)
Final response model for context building.
Contains the formatted context and metadata about the retrieval process.
CustomMemorySaveResult
class CustomMemorySaveResult(BaseModel)
Result of saving a custom-type memory.
RetrievedCustomMemory
class RetrievedCustomMemory(BaseModel)
One retrieved custom-type memory.
CustomMemoryRetrieveResult
class CustomMemoryRetrieveResult(BaseModel)
Result of retrieving custom-type memories.