AI Memory & Agent Engineering — Field Notes from Maximem

Context Language Models: What the UW and Meta Paper Changes for Agent Builders, and What It Leaves to Memory
A Context Language Model is an existing model that rewrites its own context like a file. It manages working context inside a task better than summarisation, with less compute, and throws that context away when the task ends.
AI Agent Memory
See all 38 →
Short-Term vs Long-Term Memory in LLMs, and What "Context Memory" Actually Means
Short-term memory in an LLM is whatever sits in the context window on the current call: the system prompt, the recent conversation, tool results and any retrieved text, all re-sent by the application on every request and gone when the request ends.

How to Give an LLM Long-Term Memory, and Why RAG Alone Is Not It
You give an LLM long-term memory by adding a layer outside the model that decides what is worth keeping from each conversation, stores it against the right person, updates it when a fact changes, and puts only the relevant slice back into the prompt on the next call.

Knowledge Graph vs Vector Search for Agent Memory: When to Add a Graph, and Why Time Matters
For agent memory, neither wins on its own: vector search finds past statements that mean something close to the current question, and a knowledge graph connects facts through the same person or company and tells which of two conflicting facts is current.
Context Engineering
See all 7 →
LLM Context Windows in 2026: GPT-4o Is 128K, the Frontier Is 1M, and What Happens When a Chat Runs Out
GPT-4o has a 128,000-token context window and can write at most 16,384 tokens in a single reply, with the reply counted inside the same 128,000.

Agentic Context Management: Agent Memory Is Not Merely a Storage & Retrieval Problem, It Is an Architecture Problem
We argue in our latest paper, that agent memory and cost is a lifecycle and architecture problem

An Anthropic Leader Told a Room of Founders to Stop Worrying About Context Windows. Here's the catch
An Anthropic researcher told a room of founders to stop worrying about context windows. Here is the question I did not get to ask, and why a bigger window solves short-term memory with bad tradeoffs and does nothing for long-term memory.
RAG & Retrieval
See all 8 →
Best Embedding Models in 2026: Open-Source Picks, Semantic Search, and How to Choose for RAG
As of 26 September 2026, the best open-source embedding model for most teams is Qwen3-Embedding, which is Apache 2.0, reads 32K tokens and ranks 5th on MTEB English v2 and 4th on MTEB Multilingual v2.

What Is RAG? How Retrieval-Augmented Generation Works, and Where It Stops
RAG, short for retrieval-augmented generation, is a technique that lets a large language model answer from information it was never trained on: when a question arrives, the application searches a knowledge source for the most relevant passages, places them in the prompt next to the question, and the model writes its answer from that material.

What Is a Graph Database, and When Should You Use One?
A graph database stores data as nodes and relationships and answers questions by following those relationships instead of joining tables; use one when your main questions are about paths across connected data, and keep a relational database for records and reporting.
MCP & Agent Protocols
See all 5 →
What is WebMCP? When and How to Use WebMCP in a Browser Agent?
The decision is per step, not per site: a page can expose a clean search tool and leave its account settings as ordinary DOM controls, so an agent that picks one mode per domain gets the worst of both.

MCP 2026-07-28: 20 Breaking Changes and the Errors They Cause
MCP 2026-07-28 removed sessions, the initialize handshake, and the ability for servers to initiate requests at all. It is wire-incompatible in both directions, so nothing breaks until a client upgrades underneath you. This is the full diff, every error you will hit with the fix for each, the HTTP+SSE deadline that two official sources disagree about, and a scorecard of what the release left alone.

MCP Servers Explained: What They Are and How AI Agents Use Them
Starting as a niche experiment, Model Context Protocol (MCP) is now the universal "USB-C" for AI agent integrations. Let's demystify MCP’s architecture across hosts, clients, and servers and understand how tools, resources, and prompts work together. Essential reading for engineers navigating the massive ecosystem of 18,000+ servers and 97 million monthly downloads.
Agent Frameworks
See all 16 →
Harness And Agents
Every LLM call is stateless, and a model on its own cannot act; the software that turns it into an agent is the harness.

How to Add Conversational Memory to a LangChain App (2026 Guide)
To add conversational memory to a LangChain app today, pass a checkpointer to create_agent and reuse one thread_id per conversation; to remember a person across conversations, add a LangGraph store or plug in Maximem Synap through the maximem-synap-langchain package.

How to Add Memory to a Pipecat Voice Agent
Pipecat's LLMContext and auto summarization manage one pipeline run; for recall across calls its docs list Mem0MemoryService. Synap ships the same pattern as two frame processors, SynapMemoryProcessor and SynapRecorder, at 92% LongMemEval, 93.2% on LoCoMo accuracy.
Agent Evals & Observability
See all 3 →
AI Agent Evals: The Judge Is The Part Nobody Measures
An LLM judge scoring the same agent output three times will often hand back three different answers, and almost no guide to agent evaluation asks whether the judge agrees with itself or what it costs to run on everything rather than a sample.

Most Agent Eval Frameworks Are Wrong. Here's What Actually Works
Your agent is silently degrading. Move beyond static benchmarks to master AI agent evaluation. This guide explores how to design frameworks that measure reasoning, tool-use, and reliability to bridge the gap between experimental prototypes and production-ready systems.

$4K Courses Will Teach You Agent Evals. Here's a Free Guide.
Move beyond static benchmarks to master the art of AI agent evaluation. This guide explores how to design frameworks that measure reasoning, tool-use, and reliability to bridge the gap between experimental prototypes and production-ready systems.
Claude Code & Coding Agents
See all 5 →
9 Essential Claude Skills for AI Engineers Building Production Agents + 1 Bonus Skill
Nine Claude Skills we run in production, plus one bonus, with the install command for each and what it saves you. Updated September 2026.

How To Use Jev In Your AI Agent: A Map Of Decision Seams
Jev cannot write a single word, which is the reason to care about it, so the question worth asking is never whether it should be your agent but which decisions inside your agent it should own.

Best Claude Code Skills in 2026: 12 New Skills and What Changed
Important Claude Code skills shipped since February 2026, what each replaces, install commands, and the security problem that came with the marketplace.
Voice Agents
See all 4 →
I Spoke to 500+ Voice AI Builders in India Over 3 Months. Here Is What I Found.
Field notes from 500+ Voice AI builders in India. An exploration of outbound dominance, the "too good" TTS problem, and why memory is the final infrastructure hurdle for production agents.

Voice Agent Stack: The Right Tools for Production Voice AI in 2026
Build production-ready voice agents with the right stack. Compare MCPs, frameworks (CrewAI, AutoGen, Swarms), APIs (Deepgram, ElevenLabs, Vapi), and learn cost-effective patterns for voice AI in 2026.
LLM Cost & Production
See all 4 →
How to Reduce LLM Token Costs in Long Conversations: What Caching Saves, and Where It Stops
A long conversation costs far more than its number of turns suggests, because every request carries the system prompt and the entire history again, so total input grows with the square of the number of turns.

Claude API Pricing in 2026: Every Model per Million Tokens, and What Pro and Max Cost
The Claude API costs between $1 and $10 per million input tokens and between $5 and $50 per million output tokens on Anthropic's current models.

Your AI Agent Is A Cash Guzzler. Here's a Framework for Thinking About It.
Most founders misjudge agent costs, focusing only on token price. In reality, stacked expenses from context accumulation and infrastructure can explode bills 10x at scale. Learn to identify the actual growth curve in your billing stack and why smart context management is the only viable path to sustainable unit economics.
Research & Product Updates
See all 7 →
Maximem Synap Updates: Higher Scores, 17 Integrations, and a Live Playground
Synap updates: 92% LongMemEval (up from 90.2%), 93.2% LOCOMO, 17 framework integrations, a browser playground, public pricing, and a free accuracy eval on your own agent.

How Synap Works Under the Hood
We launched Maximem Synap today. Here's a peek into how it is built.

Synap Scores 92% on LongMemEval, 93.2% on LoCoMo: What the Numbers Mean
Synap outperforms existing memory systems by redesigning context management on leading benchmarks; delivering higher accuracy, lower latency, and stable performance at scale through structured, domain-specific architectures.
Beyond our own writing, see our curated directory of engineering blogs.