AWS has published a complete reference architecture for running a context-aware personal assistant with OpenClaw on Amazon Bedrock AgentCore. The 6 October guide combines an open-source agent loop with managed runtime, memory, model access and AWS infrastructure, addressing a common weakness of assistants that answer individual questions but forget useful context between conversations.
The example is a gardening assistant called Sprout, but AWS presents the design as reusable for support bots, fitness coaches and internal help desks. A CloudFormation template deploys the system, and a skills manifest determines its domain capabilities. That makes the article more than a narrow tutorial: it is a practical pattern for teams deciding how to host OpenClaw with persistent, isolated memory.
A serverless home for OpenClaw
The agent runs in a container on AgentCore Runtime. A thin Python wrapper adapts the OpenClaw gateway to AgentCore’s HTTP contract, exposing health and invocation endpoints without requiring a fork of the framework. The wrapper also checks whether the gateway process is alive when a frozen container resumes and restarts it when necessary.
AWS recommends this wrapper approach because it preserves OpenClaw’s upgrade path. The surrounding stack uses API Gateway and Lambda for Telegram webhooks, EventBridge Scheduler for proactive reminders, Amazon S3 for workspace storage, KMS for encryption, Secrets Manager for the bot token and CloudWatch for logs and metrics. Consumption-based runtime billing means idle waiting does not require an always-on server.
Memory is split into two layers
AgentCore Memory stores short-term conversation events and asynchronously extracts durable long-term records. The sample configures strategies for explicit user preferences, semantic facts and episodic summaries. A gardener’s choice to avoid synthetic fertiliser can therefore be retrieved weeks later and used when the assistant proposes a treatment.
Namespaces isolate each user’s information. Sprout incorporates the Telegram chat identifier into separate long-term and episodic paths, giving every user a predictable boundary. AWS highlights this design decision because clean namespaces simplify testing, audits and deletion requests. Multi-tenant builders should settle the isolation scheme before storing any conversational data.
Retrieval should fail gracefully
For each message, the application searches relevant long-term records, ranks them and injects the useful context into the system prompt. Retrieval has a time budget; if memory is unavailable, the assistant answers without it rather than failing the entire turn. This treats memory as an enhancement instead of a hard dependency.
Extraction is asynchronous, so a newly mentioned fact may not appear in long-term memory during the same session. Short-term events cover the immediate conversation, while extracted records support later sessions. Product teams need to explain that distinction so users do not assume every detail is instantly durable or perfectly recalled.
Models are routed by task
The sample uses Claude Haiku 4.5 for frequent text conversations and Claude Sonnet 4.5 for less common image-analysis requests. Model identifiers remain configuration values, allowing the deployment to change models without rebuilding the container. Image turns bypass the OpenClaw gateway in this implementation because its bundled path dropped image URL content before the request reached Bedrock.
Both routes receive the same persona and retrieved memory, preserving continuity. AWS uses plant identification to show why this matters: a vision model can make a plausible but wrong identification from pixels alone, while the user’s saved plant inventory provides grounding that improves the result. Memory can therefore affect factual quality, not only personal tone.
Skills keep the pattern portable
Capabilities are declared as small entries in a community skills manifest. The example includes weather, reminders and plant notes. AWS recommends keeping each skill focused on one task that a user can name clearly, making selection, testing and replacement easier. Changing the persona and skill set can turn the same infrastructure into a different assistant.
Telegram is the sample interface because it supports webhooks, text and images without a custom client. Teams can substitute another channel, but they still need secure secret storage, request validation and a clear mapping between the channel identity and the memory namespace. Formatting and delivery failures also need explicit handling.
Cost and operational controls
Prompt caching reduces repeated processing of the stable persona and memory prefix. AWS advises placing stable content first, volatile user input last and keeping the memory order deterministic so requests actually share a cache prefix. The guide estimates light personal use at a few dollars per month, but recommends budget alerts because retries or heavy conversation can change consumption.
Production adopters should add evaluation, data-retention rules and user controls for viewing or deleting stored memories. They should log tool calls, constrain credentials and test how the assistant behaves when a connected service or model is unavailable. Personalisation is valuable only when people understand what is remembered and can correct it.
Why the reference matters
The guide gives OpenClaw users a concrete deployment path that combines managed compute, durable memory and Bedrock models while leaving the agent framework replaceable. It also surfaces practical decisions that short demos often omit: process recovery, namespace design, extraction delay, model routing, prompt caching and cleanup.
For AWS, the article demonstrates AgentCore as an operating layer for third-party agent frameworks rather than a closed assistant product. For builders, its strongest lesson is architectural: wrap the framework, isolate memory early, degrade gracefully and keep skills small. Those choices make a persistent assistant easier to operate, audit and evolve after the first successful conversation.