diff --git a/docs/guides/README.md b/docs/guides/README.md index ac6a56066e..222eab434e 100644 --- a/docs/guides/README.md +++ b/docs/guides/README.md @@ -20,6 +20,12 @@ The record of everything that happens during an invocation, and the side effects - [Event](events/event/index.md) - The `Event` and `EventActions` shapes, `isFinalResponse`, and the fields that diverge from adk-python. +### Memory + +Cross-session memory ingestion (`addSessionToMemory`) and retrieval (`searchMemory`) across in-memory keyword stores, Vertex AI RAG Engine corpora, and Vertex AI Agent Engine Memory Bank. + +- [Memory](memory/index.md) - `BaseMemoryService`, `InMemoryMemoryService`, `VertexAiRagMemoryService`, `VertexAiMemoryBankService`, and the `LOAD_MEMORY` / `PRELOAD_MEMORY` tools. + ### Planners Planning for an `LlmAgent` through its `planner` option: the model's built-in thinking, or a Plan-ReAct instruction for a model without it. diff --git a/docs/guides/memory/index.md b/docs/guides/memory/index.md new file mode 100644 index 0000000000..632c642c8b --- /dev/null +++ b/docs/guides/memory/index.md @@ -0,0 +1,168 @@ +# Memory (`BaseMemoryService`, `InMemoryMemoryService`, `VertexAiRagMemoryService`, and `VertexAiMemoryBankService`) + +Memory services in the Agent Development Kit store completed conversation sessions and retrieve relevant turns across sessions for the same application and user. + +Every backend implements `BaseMemoryService` so agents and tools can search historical sessions without coupling to a storage engine: + +- `InMemoryMemoryService` - Stores session events in process memory and matches queries by case-insensitive keyword overlap for local development and unit tests. +- `VertexAiRagMemoryService` - Uploads session transcripts into a Vertex AI RAG Engine corpus and retrieves matching chunks with tenant isolation across shared corpora. +- `VertexAiMemoryBankService` - Generates, stores, and retrieves structured user facts in Vertex AI Agent Engine Memory Bank. + +## Introduction + +A `Session` holds the turn-by-turn history of one conversation, whereas `BaseMemoryService` indexes completed sessions so an agent can recall facts from earlier conversations with the same user. Calling `addSessionToMemory(session)` ingests a `Session` into the configured backend, and calling `searchMemory({appName, userId, query})` returns a `SearchMemoryResponse` containing `MemoryEntry` items whose `content`, `author`, and ISO 8601 `timestamp` come from matching session events. + +During an agent run, `Runner` attaches the configured `BaseMemoryService` to `InvocationContext.memoryService` (`InMemoryRunner` defaults to an `InMemoryMemoryService` instance). Agents reach memory either on demand by registering `LOAD_MEMORY` (`LoadMemoryTool`), which exposes the `load_memory` function declaration to the model, or automatically on every turn by registering `PRELOAD_MEMORY` (`PreloadMemoryTool`), which injects `` into the outgoing `LlmRequest` instructions. Custom tools and callbacks call `toolContext.searchMemory(query)` on `Context`, which automatically forwards the active `appName` and `userId`. + +## Get started + +Create an `InMemoryMemoryService`, ingest a completed `Session` containing user preferences, and query those memories from a later turn for the same `appName` and `userId`. + +```ts +import { + createEvent, + InMemoryMemoryService, + InMemorySessionService, +} from '@google/adk'; + +const sessionService = new InMemorySessionService(); +const memoryService = new InMemoryMemoryService(); + +const pastSession = await sessionService.createSession({ + appName: 'travel_assistant', + userId: 'user-42', +}); + +await sessionService.appendEvent({ + session: pastSession, + event: createEvent({ + author: 'user', + timestamp: Date.parse('2025-02-10T09:15:00.000Z'), + content: { + role: 'user', + parts: [{text: 'I prefer aisle seats and vegetarian meals on flights.'}], + }, + }), +}); + +await memoryService.addSessionToMemory(pastSession); + +const searchResult = await memoryService.searchMemory({ + appName: 'travel_assistant', + userId: 'user-42', + query: 'aisle vegetarian flights', +}); + +for (const entry of searchResult.memories) { + const text = entry.content.parts?.map((part) => part.text ?? '').join(' '); + console.log(`[${entry.timestamp}] ${entry.author}: ${text}`); +} +``` + +## How it works + +1. **Session ingestion (`addSessionToMemory`)**: When a session completes or reaches a checkpoint, application code passes the `Session` object to `memoryService.addSessionToMemory(session)`. `InMemoryMemoryService` filters `session.events` to retain events where `(event.content?.parts?.length ?? 0) > 0` and stores them in a two-level null-prototype map (`Object.create(null)`) keyed by `${session.appName}/${session.userId}` and `session.id`. +2. **Tenant-scoped retrieval (`searchMemory`)**: `searchMemory` accepts a `SearchMemoryRequest` (`{appName, userId, query}`) and resolves to a `SearchMemoryResponse` (`{memories: MemoryEntry[]}`). Unlike `adk-python` v0.1.0, which groups events inside `MemoryResult` objects by `session_id`, `adk-js` returns a flat `MemoryEntry[]` array where each entry carries `content` (`@google/genai` `Content`), `author` (`string | undefined`), and `timestamp` (`string | undefined`, formatted via `new Date(event.timestamp).toISOString()`). +3. **Keyword matching in `InMemoryMemoryService`**: `InMemoryMemoryService` splits `req.query.toLowerCase()` on whitespace (`/\s+/`) and extracts lowercase alphabetic words (`/[A-Za-z]+/g`) from the joined `part.text` strings of each stored event. Any event whose word set contains at least one query word is returned as a `MemoryEntry`. Matching is whole-word, so `flight` does not match `flights`. This diverges from `adk-python` v0.1.0, which matched each keyword as a substring of the raw event text; `adk-js` follows the later whole-word behavior, so queries must use whole words that appear in the stored events. +4. **Transcript upload and chunk deduplication in `VertexAiRagMemoryService`**: `VertexAiRagMemoryService` serializes text-bearing session events into newline-delimited JSON lines (`{author, timestamp, text}` with `timestamp` in Unix epoch seconds `event.timestamp / 1000` so corpora stay interoperable with `adk-python`) and uploads the transcript as a RAG file named `adk-memory-v1...`. On `searchMemory`, it lists up to 10 pages of 100 files to narrow `ragFileIds` to the requesting tenant, calls `retrieveContexts`, filters every returned chunk through `parseSourceDisplayName`, and deduplicates overlapping chunks per session by event timestamp before sorting each session's events chronologically. +5. **Fact extraction and consolidation in `VertexAiMemoryBankService`**: `VertexAiMemoryBankService` sends session events to Vertex AI Agent Engine Memory Bank via `memories.generateInternal` and queries stored facts via `memories.retrieveInternal` scoped to `{app_name: request.appName, user_id: request.userId}`. It also provides `addEventsToMemory` for incremental event lists and `addMemory` for writing explicit `MemoryEntry` facts via `memories.createInternal` or batched consolidation when `customMetadata['enable_consolidation']` is `true`. + +## Configuration options + +### `SearchMemoryRequest` and `MemoryEntry` + +The `SearchMemoryRequest` interface defines the parameters passed to `BaseMemoryService.searchMemory`, and `MemoryEntry` defines each item in `SearchMemoryResponse.memories`. + +| Interface / Property | Type | Default | Description | +| :---------------------------- | :-------- | :---------- | :---------------------------------------------------------------------------------- | +| `SearchMemoryRequest.appName` | `string` | Required | Application name partitioning the memory store. | +| `SearchMemoryRequest.userId` | `string` | Required | User identifier whose past sessions are searched. | +| `SearchMemoryRequest.query` | `string` | Required | Natural-language or keyword query used to match stored memories. | +| `MemoryEntry.content` | `Content` | Required | `@google/genai` `Content` payload originally produced during a session event. | +| `MemoryEntry.author` | `string` | `undefined` | Producer of the event, such as `'user'`, `'model'`, or a sub-agent name. | +| `MemoryEntry.timestamp` | `string` | `undefined` | ISO 8601 timestamp string converted from the source `Event.timestamp` milliseconds. | + +### `VertexAiRagMemoryServiceOptions` + +The `VertexAiRagMemoryServiceOptions` interface configures `VertexAiRagMemoryService`. + +| Option | Type | Default | Description | +| :------------------------ | :------- | :-------------------------------------------------------------- | :-------------------------------------------------------------------------------------------------------- | +| `ragCorpus` | `string` | Required | Full resource name `projects/{project}/locations/{location}/ragCorpora/{id}` or bare `{id}`. | +| `similarityTopK` | `number` | `undefined` | Maximum number of RAG contexts requested in `query.ragRetrievalConfig.topK`. | +| `vectorDistanceThreshold` | `number` | `10` | Maximum vector distance (`query.ragRetrievalConfig.filter.vectorDistanceThreshold`) for retrieved chunks. | +| `projectId` | `string` | Segment 1 of `ragCorpus` or `process.env.GOOGLE_CLOUD_PROJECT` | Google Cloud project owning the RAG corpus. | +| `location` | `string` | Segment 3 of `ragCorpus` or `process.env.GOOGLE_CLOUD_LOCATION` | Google Cloud region hosting the RAG corpus endpoint. | + +`resolveRagCorpus` trims `ragCorpus` and throws `Error('ragCorpus is required for VertexAiRagMemoryService.')` when the string is empty. When `ragCorpus` is a bare corpus ID rather than a four-segment `projects/{project}/locations/{location}/...` resource path, both `projectId` and `location` must be supplied either on options or through `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION`. + +### `VertexAiMemoryBankServiceOptions` + +The `VertexAiMemoryBankServiceOptions` interface configures `VertexAiMemoryBankService`. + +| Option | Type | Default | Description | +| :------------------ | :------- | :-------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `agentEngineId` | `string` | Required | Reasoning Engine numeric or string ID backing the Memory Bank resource. | +| `projectId` | `string` | `undefined` | Google Cloud project passed to the `@google-cloud/vertexai` `Client`. | +| `location` | `string` | `undefined` | Google Cloud region passed to the `@google-cloud/vertexai` `Client`. | +| `expressModeApiKey` | `string` | `undefined` | Express Mode is not supported by the default Agent Engine client. `getExpressModeApiKey` throws when it is combined with `projectId` or `location`, and the constructor throws `EXPRESS_MODE_UNSUPPORTED_MESSAGE` when it is set without them. Use `projectId` and `location` with ADC, or inject a pre-configured `client`. | +| `client` | `Client` | `new Client({project, location})` | Optional pre-configured `@google-cloud/vertexai` `Client` instance. | + +## Advanced applications + +### Wiring `LOAD_MEMORY` and `PRELOAD_MEMORY` into an `LlmAgent` + +Give an `LlmAgent` either `LOAD_MEMORY` so the model calls `load_memory({query})` when it needs past context, or `PRELOAD_MEMORY` so ADK searches memory before every model call and appends `` to the system instructions. + +```ts +import { + InMemoryMemoryService, + LlmAgent, + LOAD_MEMORY, + PRELOAD_MEMORY, + Runner, + InMemorySessionService, +} from '@google/adk'; + +const memoryService = new InMemoryMemoryService(); +const sessionService = new InMemorySessionService(); + +export const memoryAwareAgent = new LlmAgent({ + name: 'travel_concierge', + model: 'gemini-flash-latest', + instruction: 'Answer travel questions using saved user preferences.', + tools: [LOAD_MEMORY, PRELOAD_MEMORY], +}); + +const runner = new Runner({ + appName: 'travel_assistant', + agent: memoryAwareAgent, + sessionService, + memoryService, +}); +``` + +### Configuring `VertexAiRagMemoryService` for cross-process persistence + +Pass a `VertexAiRagMemoryService` to `Runner` when sessions must persist in a managed Vertex AI RAG corpus across process restarts. + +```ts +import {VertexAiRagMemoryService} from '@google/adk'; + +const ragMemoryService = new VertexAiRagMemoryService({ + ragCorpus: + 'projects/my-cloud-project/locations/us-central1/ragCorpora/travel-memories', + similarityTopK: 5, + vectorDistanceThreshold: 0.7, +}); +``` + +## Limitations + +- **`InMemoryMemoryService` matches alphabetic tokens only**: `InMemoryMemoryService` extracts tokens with `/[A-Za-z]+/g` and does not index numeric literals or perform semantic embedding similarity. Use `VertexAiRagMemoryService` or `VertexAiMemoryBankService` for semantic search. +- **`VertexAiRagMemoryService` corpus listing budget**: `searchMemory` lists at most `10` pages of `100` files (`1,000` files total) when pre-filtering `ragFileIds` by tenant. If a shared corpus exceeds `1,000` files or listing fails, retrieval runs against the whole corpus and `tenantSource` still filters every returned chunk on `sourceDisplayName` so another user's memories are never returned. +- **Text-only indexing in `LOAD_MEMORY`, `PRELOAD_MEMORY`, and RAG transcripts**: `serializeSessionTranscript`, `LoadMemoryTool`, and `PreloadMemoryTool` read only `part.text` values from `Event.content.parts`; binary `inlineData` and `fileData` payloads belong in `BaseArtifactService`. + +## Related samples + +- [`samples/memory/`](../../../samples/memory/README.md) - Runnable sample exercising `InMemoryMemoryService`, `addSessionToMemory`, `searchMemory`, tenant isolation, and `EventActions.artifactDelta`. diff --git a/samples/memory/README.md b/samples/memory/README.md new file mode 100644 index 0000000000..561d69b2ff --- /dev/null +++ b/samples/memory/README.md @@ -0,0 +1,42 @@ +# Memory Sample (`BaseMemoryService`, `InMemoryMemoryService`, and `MemoryEntry`) + +This sample demonstrates how ADK agents ingest completed sessions into `ctx.memoryService` (`BaseMemoryService`) via `addSessionToMemory(session)`, retrieve matching `MemoryEntry` items across sessions via `searchMemory({appName, userId, query})`, and persist a JSON digest of the recalled memories through `ctx.artifactService` with `EventActions.artifactDelta`. + +## Overview + +`MemoryShowcaseAgent` is a deterministic `BaseAgent` subclass that seeds two completed sessions into `ctx.memoryService` (`InMemoryMemoryService`): one belonging to the active `{appName, userId}` with Lisbon flight, companion, meal, and hotel details, and a second belonging to `'other-user'` to verify tenant isolation. On each turn it calls `searchMemory({appName: ctx.appName, userId: ctx.userId, query})`, saves `memory_digest.json` via `ctx.artifactService.saveArtifact`, and returns the recalled `MemoryEntry` items (`author`, ISO 8601 `timestamp`, and `content`). + +## Sample Inputs + +- `Who am I traveling to Lisbon with and what meal did I request?` + + _Searches the active user's indexed sessions in `ctx.memoryService`, saves `memory_digest.json` via `ctx.artifactService`, and returns the matching Lisbon flight and hotel memories while excluding `'other-user'`._ + +- `Which Lisbon hotel did we confirm and what check-in note was saved?` + + _Exercises `searchMemory` for the hotel and check-in terms in `adk web` and surfaces both the `recallUserMemories` tool chip and the `memory_digest.json` artifact chip._ + +## Running the Sample + +Run the self-contained `InMemoryRunner` script directly to inspect each emitted `Event` and confirm that an unindexed user sees zero memories: + +```bash +npx tsx samples/memory/agent.ts +``` + +Or run the exported `rootAgent` interactively through the ADK CLI after building the workspace: + +```bash +npm run build +npm run sample -- samples/memory/agent.ts +``` + +`samples/` is not an npm workspace, so it is type-checked separately: + +```bash +npm run ts:check:samples +``` + +## Related Guides + +- [Memory](../../docs/guides/memory/index.md) - `BaseMemoryService`, `InMemoryMemoryService`, `VertexAiRagMemoryService`, `VertexAiMemoryBankService`, and the `LOAD_MEMORY` / `PRELOAD_MEMORY` tools. diff --git a/samples/memory/agent.ts b/samples/memory/agent.ts new file mode 100644 index 0000000000..db825da3e3 --- /dev/null +++ b/samples/memory/agent.ts @@ -0,0 +1,302 @@ +/** + * @license + * Copyright 2026 Google LLC + * SPDX-License-Identifier: Apache-2.0 + */ + +/** + * Memory (`BaseMemoryService`, `InMemoryMemoryService`, and `MemoryEntry`) + * ../../docs/guides/memory/index.md + * + * A deterministic `BaseAgent` that ingests completed prior sessions into + * `ctx.memoryService` (`InMemoryMemoryService`), queries tenant-scoped + * `MemoryEntry` items via `searchMemory({appName, userId, query})`, and + * persists the matched memory digest to `ctx.artifactService` while recording + * the revision in `EventActions.artifactDelta`. + * + * This sample exports `rootAgent` for `adk web` and `npm run sample`, and also + * includes a direct `InMemoryRunner` driver (`main()`) because the CLI only + * prints conversational text — running directly lets `main()` inspect the + * structured `MemoryEntry` fields (`author`, ISO 8601 `timestamp`, and + * `content`) and verify that searching under a second `userId` returns zero + * memories from the first user's sessions. + * + * Run (offline, no API key): + * npx tsx samples/memory/agent.ts + * npm run sample -- samples/memory/agent.ts + */ + +import { + BaseAgent, + createEvent, + createEventActions, + createSession, + Event, + getFunctionCalls, + getFunctionResponses, + InMemoryMemoryService, + InMemoryRunner, + InvocationContext, + isFinalResponse, + MemoryEntry, + stringifyContent, +} from '@google/adk'; +import {fileURLToPath} from 'node:url'; + +const MEMORY_DIGEST_FILENAME = 'memory_digest.json'; + +function extractEntryText(entry: MemoryEntry): string { + return ( + entry.content.parts + ?.map((part) => part.text ?? '') + .filter((text) => text.length > 0) + .join(' ') ?? '' + ); +} + +class MemoryShowcaseAgent extends BaseAgent { + private readonly fallbackMemoryService = new InMemoryMemoryService(); + + constructor() { + super({ + name: 'memory_showcase_agent', + description: + 'Demonstrates cross-session memory ingestion and tenant-scoped search via BaseMemoryService.', + }); + } + + protected override async *runAsyncImpl( + ctx: InvocationContext, + ): AsyncGenerator { + const memoryService = ctx.memoryService ?? this.fallbackMemoryService; + const userQuery = + ctx.userContent?.parts + ?.map((part) => part.text ?? '') + .join(' ') + .trim() || 'Lisbon vegetarian hotel'; + + const priorTripSession = createSession({ + id: 'prior-session-lisbon', + appName: ctx.appName, + userId: ctx.userId, + events: [ + createEvent({ + author: 'user', + timestamp: Date.parse('2025-02-10T09:15:00.000Z'), + content: { + role: 'user', + parts: [ + { + text: 'We booked flights to Lisbon with Nadia for October 14 and requested vegetarian meals.', + }, + ], + }, + }), + createEvent({ + author: 'model', + timestamp: Date.parse('2025-02-10T09:16:00.000Z'), + content: { + role: 'model', + parts: [ + { + text: 'Confirmed Lisbon hotel near Rossio Square with late check-in.', + }, + ], + }, + }), + ], + }); + + const otherTenantSession = createSession({ + id: 'prior-session-other-user', + appName: ctx.appName, + userId: 'other-user', + events: [ + createEvent({ + author: 'user', + timestamp: Date.parse('2025-02-11T14:00:00.000Z'), + content: { + role: 'user', + parts: [ + { + text: 'Secret Lisbon itinerary belonging to another user.', + }, + ], + }, + }), + ], + }); + + await memoryService.addSessionToMemory(priorTripSession); + await memoryService.addSessionToMemory(otherTenantSession); + + yield createEvent({ + invocationId: ctx.invocationId, + author: this.name, + branch: ctx.branch, + content: { + role: 'model', + parts: [ + { + functionCall: { + id: 'call-memory-1', + name: 'recallUserMemories', + args: { + appName: ctx.appName, + userId: ctx.userId, + query: userQuery, + }, + }, + }, + ], + }, + }); + + const searchResult = await memoryService.searchMemory({ + appName: ctx.appName, + userId: ctx.userId, + query: userQuery, + }); + + const matchedSummaries = searchResult.memories.map((entry) => ({ + author: entry.author ?? 'unknown', + timestamp: entry.timestamp ?? '', + text: extractEntryText(entry), + })); + + const artifactDelta: Record = {}; + if (ctx.artifactService) { + const rev = await ctx.artifactService.saveArtifact({ + filename: MEMORY_DIGEST_FILENAME, + artifact: { + text: JSON.stringify( + { + appName: ctx.appName, + userId: ctx.userId, + query: userQuery, + matchCount: matchedSummaries.length, + memories: matchedSummaries, + }, + null, + 2, + ), + }, + }); + artifactDelta[MEMORY_DIGEST_FILENAME] = rev; + } + + yield createEvent({ + invocationId: ctx.invocationId, + author: this.name, + branch: ctx.branch, + content: { + role: 'user', + parts: [ + { + functionResponse: { + id: 'call-memory-1', + name: 'recallUserMemories', + response: { + appName: ctx.appName, + userId: ctx.userId, + query: userQuery, + matchCount: matchedSummaries.length, + memories: matchedSummaries, + }, + }, + }, + ], + }, + actions: createEventActions({ + stateDelta: { + lastMemoryQuery: userQuery, + recalledMemoryCount: matchedSummaries.length, + }, + artifactDelta, + }), + }); + + const recalledLines = matchedSummaries + .map((item) => `[${item.timestamp}] ${item.author}: "${item.text}"`) + .join(' | '); + + yield createEvent({ + invocationId: ctx.invocationId, + author: this.name, + branch: ctx.branch, + content: { + role: 'model', + parts: [ + { + text: `Recalled ${matchedSummaries.length} memory entries for ${ctx.userId} in ${ctx.appName}: ${recalledLines}`, + }, + ], + }, + }); + } + + protected override async *runLiveImpl( + ctx: InvocationContext, + ): AsyncGenerator { + yield* this.runAsyncImpl(ctx); + } +} + +export const rootAgent = new MemoryShowcaseAgent(); + +async function main() { + const appName = 'memory_sample_app'; + const userId = 'user-1'; + const runner = new InMemoryRunner({ + agent: rootAgent, + appName, + }); + + const session = await runner.sessionService.createSession({ + appName, + userId, + }); + + for await (const event of runner.runAsync({ + userId, + sessionId: session.id, + newMessage: { + role: 'user', + parts: [ + { + text: 'Who am I traveling to Lisbon with and what meal did I request?', + }, + ], + }, + })) { + const calls = getFunctionCalls(event); + const responses = getFunctionResponses(event); + process.stdout.write( + `${JSON.stringify({ + id: event.id, + author: event.author, + isFinal: isFinalResponse(event), + functionCalls: calls.map((c) => c.name), + functionResponses: responses.map((r) => r.name), + artifactDelta: event.actions.artifactDelta, + text: stringifyContent(event), + })}\n`, + ); + } + + const unindexedUserCheck = await runner.memoryService?.searchMemory({ + appName, + userId: 'unindexed-user', + query: 'Lisbon', + }); + process.stdout.write( + `Memories visible to unindexed-user: ${JSON.stringify(unindexedUserCheck?.memories ?? [])}\n`, + ); +} + +if (process.argv[1] === fileURLToPath(import.meta.url)) { + main().catch((err: unknown) => { + process.stderr.write(`${String(err)}\n`); + process.exit(1); + }); +}