What is a customer context graph, and when do you need one?
Short answer. A customer context graph stores what customers said as linked records: messages grouped into topics, tied to the people and companies behind them and to the issues and PRs that address them. It answers questions that cross those links, such as "who else hit this, which accounts are they at, and is anyone working on it?" A vector database finds text that is similar to a query. A knowledge graph stores typed entities and relationships. A memory layer lets an agent remember across sessions. A CRM is the system of record for accounts, contacts and deals. They overlap, and most real systems combine two or more. Pick the one that matches the question your agent has to answer. Modem (our product) builds a customer context graph over customer conversations: it reads Slack, Discord, support tools, calls and trackers, groups messages into topics, and links them to people, companies and work. It is stored in a relational database, uses embeddings as one signal for meaning-based search, and is not a CRM, a graph database or a vector store.
We build Modem, so read the Modem sections as a vendor describing its own product. The comparison sections stick to what each vendor says about itself and to the structure of the problem.
"Context graph" does not have one definition
Treat the term carefully, because vendors use it for different things.
| Who | What "context graph" means there |
|---|---|
| Neo4j | A three-tier memory architecture for AI agents: long-term knowledge, short-term conversation memory, and "reasoning memory" made of decision traces. Neo4j's blog frames the knowledge graph as defining what a relationship means and the context graph as applying that to the task at hand. |
| Zep and Graphiti | A temporal graph built from conversations, business data and documents, where each fact has a validity window: when it became true and, if ever, when it was superseded. |
| Modem | Customer conversations stored as topics, people and companies linked together, with the original quotes attached. The site's name for what the product docs call topics, sources, people and companies. |
No source we checked shows an industry-agreed definition. Several vendors, including Neo4j, Graphlit, Mem0 and Zep, say they combine vector search with graph traversal, so the terms in this guide are not clean categories. If an answer to "context graph vs knowledge graph" sounds confident, check whose definition it uses.
The five tools, and when each is the right one
Vector database. It stores embeddings and returns items near a query. Pinecone describes its database as a managed vector database where writes are instantly searchable, and now also sells Nexus, a knowledge-engine layer above it. Use one when you have a lot of unstructured text and the question is "find things like this": help-center retrieval, document Q&A, "what is our policy on refunds". Vendors such as Neo4j and Mem0 point out where it falls short alone: it gets weaker when the question depends on connections between entities, such as who, which account, which issue.
Knowledge graph. It stores entities and typed relationships, sometimes with a taxonomy or ontology that says what the relationships mean. Neo4j is the general-purpose option, and it keeps vector indexes alongside the graph. Glean uses a knowledge graph to help employees search across company apps while inheriting source-system permissions. Use one when the relationships are the point, you need multi-hop questions, or you already have clean entities. You supply extraction, deduplication and identity resolution unless the product you pick provides them; we did not find identity resolution in Neo4j's docs.
Memory layer. It lets an agent keep and recall facts across sessions. Mem0 offers memory with an Apache 2.0 core and, on its platform, graph memory that links entities. Letta has memory blocks that stay in the context window plus archival search. Zep tracks how facts change over time. Use one when the agent has to remember a user, a preference or a decision from last week. It does not by itself ingest your support tools and decide that three tickets are the same problem; you build that.
CRM. It is the system of record for accounts, contacts, opportunities, owners and the activities people log. Use it as the source of truth for who the account is, who owns it, and what deals are open. It does not replace the places where customers actually talk, and it is usually not where "twelve people raised the same export bug across three tools" lives. A customer context graph sits beside the CRM and adds what customers said and how it connects. It does not replace the CRM, and Modem is not one. Modem has a read-only Salesforce sync the agent can query when asked, and no HubSpot integration.
Customer context graph. Use it when the question crosses the links: what did customers say, who said it, which accounts are they at, what do those accounts pay, and what work addresses it. Its scope is narrow on purpose. A graph of your logistics network, or of everything your company knows, is a different build.
Which one for which question
| The question your agent faces | Reach for |
|---|---|
| "What does our docs say about SSO setup?" | A vector database over your docs |
| "Which entities in this contract relate to each other?" | A knowledge graph |
| "What did this user tell us last week?" | A memory layer |
| "Who owns this account and what deals are open?" | Your CRM |
| "Who else reported this, which companies, and is there an open issue?" | A customer context graph |
These stack. A customer context graph can use embeddings for the "find conversations about X" step, a CRM for account ownership, and a memory layer for what the agent has learned about how your team works.
What a customer context graph links
Most feedback tooling stores feedback as rows: a message, a tag, maybe a linked ticket. A customer context graph stores relationships, and three of them do most of the work.
The same issue across channels. A complaint about export timeouts arrives as a Discord thread, two support tickets and a line in a sales call. In a list, that is four items. In the graph, it is one topic with four sources, and the count is the point. A list that stores duplicates as separate rows undercounts everything that arrives through more than one door.
The same person across identities. The user who emails support and the user who posts in your shared Slack channel are often one person with two identifiers. Linking them is what makes "who asked" answerable, and closing the loop possible.
The person to the company, and the topic to the work. A request is more useful tied to the account and to the issue or PR addressing it. That turns "12 people asked" into "12 people asked, at these companies, and this PR is open for it." For revenue context, the revenue-weighted view joins billing data to the companies behind a topic.
How Modem's version works
Modem reads conversations from connected sources and keeps each message verbatim with its author, source and thread. It splits a conversation into the separate concerns inside it, so a thread that raises three problems is tracked as three, and groups those concerns into topics by meaning, across sources. Each topic keeps a trail back to the exact messages behind it and a change history.
Topics link to the people and companies behind them, and to the issues and PRs that address them by what those say. "Who is affected" is worked out when you ask, by following authors to people to companies, which is why a merge or a company change corrects the counts. People are matched by exact email or platform account, never by embedding or by name. Accounts without an email, such as Discord, stay separate until merged by hand.
Under the hood it is a relational database with typed links between records. It is not a graph database and not a vector store. Embeddings are one retrieval signal: they help find conversations about a concept and propose candidate topics, while exact matching handles identity and typed links handle which company and which issue.
Account data is narrower than the term suggests. Stripe plan and billing data attach to the companies behind a topic. Salesforce is a read-only sync that the agent can query when asked, not a stored link. Revenue is not on every topic and does not set a topic's priority.
Agents reach it over an MCP server (16 tools, including a read-only search_modem that spends no agent credits), a REST API, and a write-only Ingest API for custom sources. The build guide covers each in detail.
An illustrative scenario
This is an illustrative scenario with no real customer behind it. Suppose a coding agent is asked to fix a customer-reported export bug.
- With a vector database alone, the agent gets the messages most similar to "export bug". It cannot tell that two came from the same person, which companies are involved, or that an issue already exists.
- With a CRM alone, it has account records and notes, but the support threads and Discord messages that describe the failure live in other tools.
- With a customer context graph, a question like "what do customers say about export timeouts, who, and is there an open issue?" returns one topic with its quotes, the people and companies on it, and the linked issue or PR if one exists. With Modem,
search_modemreturns a short answer plus the matching rows, with flags that say whether the result is complete, truncated or partial.
The structure does not guarantee the answer is right. Grouping is probabilistic, so a model can over-merge or over-split, and people can merge or re-prioritize topics by hand. Counts are of concerns, not of conversations or people.
Why a tag taxonomy is not a graph
Tags label rows; they do not connect them. Two tickets tagged exports still do not know they are from the same customer, and the tag carries no memory of what was said. When someone asks "what specifically breaks for enterprise customers in exports?", a tag system returns a pile of tickets to re-read. A graph returns the topic, the people on it, their companies and the quotes, because the joins were made at capture time and not reconstructed at question time.
That is also why building one by hand decays. The joins have to happen on every new message: group it with existing topics, match the sender to a known person, link the person to the account. Done manually, it is the triage work that slips first in a busy week, and a graph with stale joins degrades back into a list.
What having one changes
- Counting reflects every channel, not only the one your tooling watches, which changes what quantified feedback ranks.
- Insight questions become lookups. "What are the top complaints from customers on the enterprise plan since the last release?" is a query, not a research project across four tools. The plan comes from Stripe on the company, joined when the agent is asked.
- Agents get somewhere to stand. An agent pointed at raw channels re-reads threads and misses the half that lives in another tool. An agent pointed at a graph over MCP starts from the merged topic. See the MCP servers for customer feedback for the comparison.
- The loop can close. Because requesters are linked to topics and topics to work, "who asked" is a lookup. Modem's opt-in Close the loop automation posts an internal Slack note naming who asked when a PR directly linked to a topic merges. A merge is not a release, so time any message to customers to your own release.
Try it by hand first
You do not need software to think in graph terms. Pick your top five feedback themes. For each, write down every place it appeared and every named person who raised it, with their company. That is a customer context graph with five nodes, maintained by hand, and it will change your next prioritization conversation, because you will argue from counts and accounts instead of anecdotes.
The catch is maintenance. The hand-built version stops being true the week you stop updating it. If you want to build the automated version yourself, the build guide walks through entities, identity resolution, ingestion, linking, retrieval, permissions and freshness, and the tools guide compares the options.
FAQ
What is the difference between a context graph, a knowledge graph and a vector database?
A vector database retrieves items whose embeddings are near a query. A knowledge graph stores typed entities and relationships. "Context graph" has no single definition: Neo4j uses it for an agent memory architecture, Zep for a graph whose facts carry validity windows, and Modem for customer conversations linked to people, companies and work. Several vendors combine vector search and graph traversal, so they are not mutually exclusive.
What is a customer context graph and how is it different from a CRM?
A CRM is the system of record for accounts, contacts, deals and the activities people log. A customer context graph adds what customers said across support, chat, calls and trackers, grouped into topics and linked to people, companies and work. They sit side by side. Modem is not a CRM; it has a read-only Salesforce sync the agent can query when asked.
Do I need a graph database for a customer context graph?
No. Modem's version is a relational database with typed links between records, and embeddings are one retrieval signal. A graph database is worth considering if you need deep multi-hop traversals or temporal facts.
Is a customer context graph the same as agent memory?
No. Agent memory, such as Mem0, Letta or Zep, keeps what an agent learns across sessions. A customer context graph is built from your customers' conversations. Modem's agent has its own saved memories and Org Skills, which persist across sessions but are not exposed over MCP or REST, and a memory is a saved note, not the graph.
Which should I use to give AI agents customer context?
Match the tool to the question. Use a vector database for "find text like this", a knowledge graph when relationships are the point, a memory layer to remember a user, your CRM for account ownership and deals, and a customer context graph when a question crosses who said it, which company, and what work exists. Most teams end up combining two.
