How to build a customer context graph for AI agents
Short answer. To build a customer context graph for AI agents, store every customer message verbatim with its source and author, group messages into topics by meaning, resolve people and companies across your tools, link topics to the issues and PRs that address them, and give agents a read path that is scoped, bounded and honest about gaps. You do not need a graph database: a relational database with a table of typed links and a vector index covers it. The hard part is running three jobs forever: placing messages on topics, matching people, and keeping links current. If you would rather not run them, Modem (our product) does this for customer conversations. It reads Slack, Discord, support tools, call transcripts and trackers, groups messages into topics, links them to the people and companies behind them and to issues and PRs, and exposes the result to agents over an MCP server, a REST API and an Ingest API. It covers customer conversations only, and it is not a graph database or a vector store.
Agents give generic answers about customers for a few reasons that are easy to mix up. The data is scattered across tools. Retrieval finds text that looks similar but cannot say who said it or what they pay. Systems disagree with each other. And a request gets over-read: a customer says "I need X" and the model treats it as demand, though they may already have X. A context graph is the structure that addresses the first three. This guide builds it in the order you would do the work, then covers where buying fits. For the concept, and how it compares with a vector database or knowledge graph, see what is a customer context graph.
"Context graph" is not a settled term. Neo4j uses it for a memory architecture with decision traces, Zep uses it for a graph whose facts have validity windows, and Modem uses it for customer conversations linked to people, companies and work. This guide means the last one: a customer-scoped store an agent can query.
Step 1: write the questions the agent must answer
Everything else follows from these, so write them first. Five that a product or engineering agent usually needs:
- What are customers saying about this feature, with quotes?
- Who asked, and which companies are they at?
- Which of those companies pay us, and on what plan?
- Is there already an issue or PR for it, and what state is it in?
- Did this come up before, and what happened?
Each question tells you what to store and which join it needs. Question 3 needs billing data joined to companies. Question 4 needs issues and PRs linked to topics. Question 5 needs history. If a question needs data you will never connect, cut it now. These five also serve as your acceptance test at the end.
Step 2: separate conversations from context
Decide which sources produce feedback and which only describe the customer. The split matters more than the list.
- Conversation sources hold what customers said: Slack and Discord channels, support tickets, email, call transcripts, and the issues and PRs in your trackers. These get grouped into topics.
- Context sources describe the customer: billing, CRM, product analytics, docs. These should not create topics. A subscription event is not feedback, and mixing it in pollutes the counts.
Modem draws the same line. Its integrations either capture conversations that become topics, or give the agent read-only data or tools: Stripe, Salesforce, Notion, PostHog and Mixpanel never create a topic. They are queried when someone asks, not copied into the graph.
Start with support and community, then add account data. Calls and engineering sources multiply what is already there. Within each source, read only what you choose: specific channels, repos or boards. Start narrow and expand.
Step 3: design the entities and links
Keep the model small. Seven record types cover the questions above.
- Container: one unit of conversation in a tool (channel, ticket, thread, call, issue, pull request), with a kind, a URL and, for trackers, a state such as open, in progress or done.
- Message: stored verbatim with author, source ids, thread links and attachments. You cannot quote or re-derive from a summary alone.
- Concern: one independently trackable problem or request extracted from a stretch of messages. A thread that raises three problems yields three concerns. A concern is never a whole conversation.
- Topic: what concerns roll up into. It has a title, summary, issue type, priority, lifecycle state, and first and last evidence times.
- Person, with every identity they appear under (emails, Slack ids, Discord handles).
- Company, with the people currently at it.
- Work item: an issue or PR, which is another container whose state you read.
Then the links, which carry the meaning. A concern cites the message it came from. A concern is a member of a topic. A topic can be a recurrence of or superseded by another topic. A person authored a message and works at a company. A topic's evidence is a walk (topic, member concerns, cited messages), and "who is affected" is a walk through authors to people to companies.
Three choices save pain later:
- Store links as rows with history, not columns. When a concern moves from one topic to another, close the old link and open a new one, so you can see what happened and keep the previous placement.
- Compute counts at question time. Do not store "number of affected people" on the topic. Walk the links when asked, so the number cannot go stale when someone is merged or changes company.
- Record who or what changed something. Every change to a topic (member joined, priority changed, merged) should note whether a service or a person did it.
In a relational database this is a handful of tables plus one table of typed links with a validity range. A graph database also works, and it is the better call if your agents will run deep multi-hop traversals. The alternatives section below covers when.
Step 4: build ingestion that survives production
An ingest path is a connector per source, and each has its own auth, webhooks or polling, retries and backfill. The details that matter:
- Process each container in order. Queue messages per container and drain them in order, so a reply never lands before the message it answers. Accept first, process after: a successful receipt means "accepted", not "analyzed".
- Backfill with a bounded window. Modem imports about 30 days by default for most sources, with named exceptions such as Canny (full history, up to 10,000 posts) and Front (90 days). Pick a window per source and record it, because an agent should know what history it can see.
- Filter bots, read attachments, scrub early. Drop bot-authored content unless you flag it as human. If screenshots carry error text, extract it. Run redaction before storage, embedding or any model call, because text you already embedded is not redacted.
- Keep a custom path. Whatever connector you do not write should be pushable through your own ingest endpoint. Modem's Ingest API exists for the same reason: conversations from sources you build can be sent in and aggregated with the built-in integrations.
- Do not promise latency. Processing is queued. Do not tell an agent or a teammate how fast a message becomes a topic unless you have measured it.
Step 5: resolve identity conservatively
"Who said this" is the join that makes everything else useful, and it is where hand-built graphs quietly fail. Use a ladder that prefers exact keys:
- If you have seen this platform account before, reuse its person.
- Otherwise, if its email exactly matches an existing person, attach the account to that person.
- Otherwise, create a new person.
Do not match people by name or by embedding. Natural keys are weaker than they look (shared inboxes like billing@, changed addresses), and a false merge corrupts every count it touches. Never auto-merge two existing people because an email collides; record the collision and let a person decide. Some platforms give you no email at all. Discord is the common case, so those people stay separate until someone merges them by hand. That is a real cost, and it is better than a wrong merge. Keep an audit trail for every manual merge, and since merges are hard to reverse, build the undo before you need it.
Companies come from two signals: the workspace a user belongs to (a Slack workspace or Slack Connect partner, for example) and the work email domain, skipping free providers such as gmail.com. Attach your own teammates to your own company so "external customers only" is a filter, not a guess.
Step 6: group messages into topics
This is the step that turns a pile of messages into counts. A workable pipeline:
- Extract concerns. Have an LLM read a window of messages and emit one concern per problem that would be fixed by a different change. A shared theme is not enough to merge two concerns: stale results and wrong grouping are both "quality feedback", but fixing one leaves the other standing.
- Find candidates. Embed each concern and retrieve similar recent ones. Embeddings are a retrieval signal here, not the decision.
- Decide join, new, or split. Have a model compare the candidates with the evidence in front of it and decide whether the concern joins an existing topic, starts a new one, or means an existing topic should split.
- Update, never duplicate. When you re-read a window, update the concerns it continues instead of inserting new rows.
- Give humans the override. Merge, archive, re-prioritize, change issue type. LLM grouping can over-merge and over-split, and a wrong merge silently skews every count it touches.
Pick your similarity threshold by hand-labelling a sample of your own pairs and measuring, not by copying a number from a blog. Keep a replay set of labelled examples for each model step so you can see when a prompt or model change makes grouping worse. Skip one-off remarks and anything not about your product; only substantive discussion should become a topic.
Step 7: link topics to the work that resolves them
A topic is more useful when it points to the issue or PR that addresses it. Treat issues and PRs as containers whose text, state and URL you ingest, then place them under topics by what they say, using the same grouping step. Two rules help:
- A PR can join a topic but should not start one. Anchor topics in what customers said.
- Read state from the work item. An open PR is in progress; a merged one is done. Merged is not shipped: a deploy, flag or staged rollout may still sit between the merge and the customer. Do not let an agent tell a customer something is live because a PR merged.
Linking by content rather than by a hand-maintained ticket id lets "which topics did this PR ship for" run backwards over the same data.
Step 8: attach account context, and keep it honest
Mirror billing and CRM data read-only and join it when a question needs it. Match billing customers to people and companies using the same identity ladder. Record when each mirror last synced and show it with every answer: stale account data is worse than missing data, because the agent keeps answering confidently from last quarter.
Do not store revenue on topics, and do not rank topics by it unless you mean to. Revenue is a property of companies, and a topic reaches revenue through the companies behind it. "Which paying customers reported this" is a walk from topic to concerns to authors to people to companies to plan.
Step 9: build the retrieval layer
Agents need three kinds of access, and a single vector search covers only one.
- Meaning-based search: "conversations about export timeouts". Embeddings are good at this.
- Exact and structured lookup: a topic by id, a person by email, everything for one company. Do not use embeddings for these.
- Joins: affected companies, plan, linked issues, recurrence. These are walks over the typed links.
You can expose these as fixed endpoints, or as a natural-language question that a model turns into read-only queries. If you build the second, put the guardrails in the database, not the prompt: a read-only role, tenant isolation enforced by row-level rules, query timeouts, row and size caps, vector columns stripped from results, paging with a continuation token, and flags that say whether a result is complete, truncated or partial.
Two habits make answers trustworthy. Return the verbatim quote and its source next to any summary, so facts and interpretation stay separate and the agent can show its work. And when systems disagree, return each system's value with its source and timestamp instead of letting the model pick one.
Step 10: permissions
Decide these before the first agent connects.
- Scope by organization from the credential, never from the request. The caller should not be able to name a different tenant.
- Respect source access. Do not ingest what you should not index. Modem reads only the Slack channels it is subscribed to, never DMs, and a private channel only if a member shares it, with no history from before that. An agent cannot see what was never stored.
- Separate read from act. Use a read-only scope for search and a different scope for anything that writes or starts an agent run. Run write calls as the signed-in user, under their existing role, so an agent never exceeds the human behind it.
- Make destructive tools announce themselves. Merges and bulk updates should be annotated so clients ask first.
- Rate-limit per tool and per key, and keep a cheap read path separate from a paid, slower agent path.
- Ask before saving memory. If your agent keeps notes between sessions, have it propose each one and wait for a yes. A single shared memory for everything blurs who is allowed to see what.
Step 11: keep it fresh
Freshness breaks in four places.
- Ingestion lag. Webhooks first, polling where there is no webhook, ordered queues, and a backfill when you connect a source.
- Mirrors. Account data refreshes on a schedule (Modem's Salesforce sync runs about every 15 minutes, for example). Say so in answers.
- Topic state. Lifecycle should follow the state of the linked work. A common failure is an agent recommending work that already shipped, because nothing moved the topic when the PR merged.
- Identity and merges. A merge changes counts everywhere. Because counts are computed when asked, they correct themselves.
Keep a change history on every topic, including who or what changed it. When an answer looks wrong, you will want to see how it got that way.
Step 12: let agents reach it, and tell them it exists
Expose the read path where your agents already work. An MCP server is the common choice for interactive agents such as Claude Code or Cursor, and an HTTP API suits pipelines. Then add a line to the agent's instructions saying the graph exists and when to use it. An agent that was never told does not query it. For worked examples, see how to give AI agents customer context and the tools for giving Claude Code customer context.
The alternatives, and what each leaves for you
You do not have to build every layer from scratch. The tools below cover different layers. Mostly they are storage, retrieval or memory, so you supply the customer model, the connectors and identity resolution.
| Tool | Layer it covers | What you still build |
|---|---|---|
| Neo4j | A general graph database with vector indexes in the same store. A neo4j-graphrag Python package supports building graphs from text. | Ingestion, grouping, identity resolution and freshness. We found no identity resolution in the docs we checked. |
| Graphiti and Zep | Graphiti is an open-source temporal graph library with Neo4j, FalkorDB or Amazon Neptune as backends. Zep sells a hosted platform on top of it, with agent memory and an MCP server. Facts carry validity windows. | Your customer schema, your connectors, and identity resolution across tools. Strongest when facts change over time. |
| Mem0 | Memory for apps and agents, with an Apache 2.0 open-source core and, on the platform, optional graph memory that links entities (Pro plan). | Ingestion from your support and chat tools, and identity resolution across Discord, Zendesk and Salesforce. |
| Pinecone | A managed vector database, now with a knowledge-engine layer called Nexus. | Entities, links, identity and permissions, modeled on top. |
| Letta | Stateful agents whose documented primitives are memory blocks and archival search. | The customer data model and ingestion. |
| Graphlit | A developer API that ingests 30+ feed types (including Slack, Gmail and GitHub), extracts entities and offers hybrid search and an MCP server. | Your customer-specific model and identity rules. Check whether its feeds cover your support tools. |
| Glean | Enterprise search across company apps that inherits source-system permissions, licensed per seat. | It is built for employees finding company knowledge; we did not find feedback grouping in its docs. |
Check each vendor's pricing page before you plan around a number: Zep, Mem0, Neo4j, Pinecone, Graphlit. For a broader comparison, see the tools guide.
Build it yourself when you have a platform team, a domain the off-the-shelf options do not model, or data that cannot leave your infrastructure. Those are legitimate reasons. If none apply, Steps 4 to 11 are a service you would be running indefinitely.
What you would be building: the component list
For planning, this is what a working customer context layer consists of. We are not estimating effort, because we have no sourced figure for it.
- A connector per source, with auth, signed-webhook verification, retries and backfill.
- A store that keeps verbatim messages with source, author, thread and container state.
- Concern-level extraction that updates instead of duplicating.
- Topic placement over time: join, create, split, plus recurrence and replacement lineage.
- A link model with history, and a rule for how issues and PRs join topics.
- Identity resolution for people and companies across platforms.
- A meaning-based index alongside exact search.
- A query layer that is read-only, scoped per tenant, bounded, and flags incomplete results.
- Agent-facing auth with separate read and run scopes, rate limits and idempotency keys.
- A safe write path: actions as the user, destructive annotations, partial-success reporting.
- Tool guidance shipped to outside agents so they pick the right tool.
- Redaction before storage and embedding.
- Evaluation sets for each model step.
- A cost model that separates a cheap read path from a paid agent path.
Where buying fits: what Modem does and does not cover
We built Modem because customer signal is scattered, duplicated and unresolved, and an agent given raw access re-reads threads and still misses half the story. So Modem builds the grouped, linked picture at ingest, before the agent asks, and the agent reads findings instead of threads. Here is how that maps to the steps above.
What it stores and links. Modem reads conversations from connected sources and keeps each message verbatim. It splits conversations into the separate concerns inside them and groups those into topics by meaning. Topics link to the people and companies behind them (worked out when you ask) and to the issues and PRs that address them, by what those say. A topic's change history is kept, and any topic traces back to the exact messages behind it. It is stored in a relational database, not a graph database or a vector store. Embeddings are one retrieval signal for meaning-based search, while people are matched by exact email or platform account, never by embedding.
What it does not do. "Context graph" is our site's name for the whole; Modem's own docs say topics, sources, people and companies. It covers customer conversations only, not a logistics network or a general knowledge base. Plan and revenue come from Stripe and sit on the companies behind a topic, not on every topic. Salesforce is a read-only sync the agent can query when asked; it is not a stored link on companies. There is no HubSpot integration. Accounts without an email, such as Discord, stay separate until merged by hand. History is bounded by each source's backfill window. Modem does not read source code, and Notion, PostHog and Mixpanel are fetched live and never become topics. Processing is queued, and we make no promise about how fast a message becomes a topic.
How agents reach it.
- MCP server at
https://mcp.modem.dev/mcp, currently in beta. It uses OAuth, so there is no API key to paste, and it exposes 16 tools.search_modemtakes a natural-language question and returns a short answer plus the matching rows, with flags forcomplete,truncatedandpartial. It is read-only and spends no agent credits.modem_skillsreturns instructions that tell your client which tool to use, andmodem_docssearches Modem's public docs. A token with only thedata:readscope sees just those three. With theagent:invokescope you also getmodem_agent_invoke,modem_agent_get_run,modem_agent_send_messageandmodem_agent_cancel_runto start and manage a durable Modem agent run, plus record tools such asupdate_people,merge_people,create_companies,update_companies,merge_companiesandadd_people_to_company, which act as the signed-in user under their role. Agent runs started through MCP have no approval prompt. The limit is 20 calls a minute per organization per tool. See the MCP docs. - REST API at
https://api.modem.dev/v1, with an organization API key. It reads topics (including optional semantic search), people, companies, groups and channels. It does not return a topic's linked issues or PRs, and it does not accept messages. The limit is 120 requests a minute per key. Usesearch_modemor an agent run for the links. See the API docs. - Ingest API, write-only: it sends conversations in from sources you build, where they are aggregated with the built-in integrations. A 200 means accepted, not processed. See the Ingest API docs.
Memory and skills. The Modem agent keeps memories, which are notes it proposes and saves once a person agrees, and your team can write Org Skills, instructions the agent loads when relevant. Both persist across sessions, but neither is exposed over MCP or REST. An outside agent benefits from them only by running the Modem agent. A memory is a saved note, not the graph.
If this fits, the integrations list shows what you can connect and pricing shows what it costs. If you need a graph of something other than customer conversations, you are back in the build path above.
An acceptance test before you scale it
Take the five questions from Step 1. If you build, start with one source and one week of data, and check that grouping holds up against real duplicates before you write a second connector. If you buy, connect two channels and ask the agent a question whose answer you already know. In both cases the answer should include names, quotes and companies you would defend in a roadmap meeting, a note of anything missing, and it should still be right a month later without anyone tending it.
FAQ
Do I need a graph database to build a context graph?
No. A relational database with a table of typed links and a vector index handles the questions in Step 1. A graph database is worth it if your agents will run deep multi-hop traversals or you need temporal facts. Neo4j, and Graphiti on top of it or another backend, are the usual choices. You still write extraction, identity resolution and freshness yourself.
How do I keep a context graph from going stale?
Ingest through webhooks and ordered queues, compute counts when asked instead of storing them, read topic state from the linked issue or PR, record when each mirrored dataset last synced, and keep change history. Stale account data, and topics that never moved when the work shipped, are the two failures to test for.
How do I handle permissions in a context graph for agents?
Take the tenant from the credential, never from the request. Ingest only what you are allowed to index. Use a read-only scope for search and a separate one for writes, run writes as the signed-in user, and rate-limit per tool. Ask before an agent saves a memory.
Can an agent query Modem directly?
Yes. Connect the MCP server and the agent can call search_modem with a plain question. It is read-only, spends no agent credits, and flags incomplete results. The REST API reads topics, people and companies but does not return linked issues or PRs. The agent only sees what Modem has stored from the sources you connected.
Is Modem a CRM or a vector database?
Neither. It stores and links customer conversations to people, companies and work, and it uses embeddings as one retrieval signal. Your CRM stays the system of record for accounts.
