Back arrowAll guides

Why AI agents answer generically about your customers, and how to give them context

Pixel art of a teal robot receiving a glowing gold document across a dark starfield
Talton Figgins•••18 min read

Short answer. An AI agent gives generic answers about your customers because it can't see four things: who the customer is, what plan they're on, what they've said to you before, and which reports from different tools are the same issue. It answers from what it can reach, which is usually your repo, your docs, and its training data. A longer prompt doesn't fix that. What fixes it is a context layer: conversations stored verbatim with their source, people and companies resolved across tools, billing joined to those companies, reports grouped by the concern they describe, and all of it exposed to the agent through something it can query mid-task, such as an MCP server. You can build that yourself from the steps below. Modem (our product) is a managed version for software teams: it reads conversations from the sources you connect (Slack, Discord, support tools, calls, issue trackers), groups them into topics by meaning, links each topic to the people and companies behind it and to the issues and pull requests that address it, and gives agents that data through an MCP server with 16 tools. Its limits: it only knows what it has read from sources you connected, it doesn't read your source code, and plan data appears only if you connect Stripe.

We build Modem, so weigh our take on it accordingly. The diagnosis and the build-it-yourself path apply whatever you use.

Why AI agents give generic answers about customers

Ask a coding agent "what do customers say about our export feature" and you get a confident essay assembled from code comments and general knowledge. Ask a support or product agent "is this account at risk" and you get advice that would fit any account. The agent isn't being lazy. Each session starts blank, and the customer half of the picture lives in places it can't reach or can't connect.

Four pieces are usually missing:

  • Who the customer is. The person in a Discord thread, the requester on a Zendesk ticket, and the attendee on a sales call may be the same person at the same company. Nothing in the raw data says so. Without that link, the agent can't tell one loud customer from five quiet ones.
  • What plan they're on. Plan, billing status, and spend live in your billing system or CRM, not in the conversation. Without them, "someone wants SSO" never becomes "three paying accounts on the top plan want SSO."
  • What they've said before. History is spread across chat, tickets, email, and calls. An agent that sees only the current ticket can't know the customer reported the same thing a month ago, or that the team already fixed it.
  • Which reports are the same issue. Customers describe one bug five different ways in five tools. Without grouping, the agent either overweights whichever channel is noisiest or treats each report as new.

Two more problems sit on top of those. Records are not meaning: a CRM can say "active" while billing says "past due" and support says "escalated," and an agent handed all three either picks one or goes vague. And context goes stale: a summary written last month may recommend work the team has since shipped.

Why a longer prompt or plain RAG doesn't fix it

A bigger system prompt or a "master context" file is push: someone assembles context ahead of time and pastes it in. It works for a short, stable list (product conventions, known limitations). It goes stale the day it's written, and someone has to rebuild it every time.

Retrieval over documents (RAG) answers "what is our refund policy" well, because the answer is a passage of text. It answers "who hit this bug and what do they pay us" badly, because that answer is a join across people, companies, billing, and many conversations, not a passage.

A vector store on its own finds messages that read like the question. It can't tell you two messages came from the same person, that a fact has changed since, or which company an author works for. Those are relationships, and you have to model them.

Raw access to every tool (an MCP server for Slack, another for Zendesk, another for your CRM) gives the agent reach without synthesis. It re-reads threads and tickets in every session and still has to work out which ones match. In our own test of Modem's MCP server (one question, 3 runs per cell, tokens counted as what entered the context window), querying any single source directly cost at least 2.4 times the tokens for a partial answer, and querying all three sources cost 4.6 times. That is one test on one question, not a general benchmark, but the direction matches what you'd expect: synthesis done once at ingest is cheaper than synthesis redone in every session.

What customer context an agent actually needs

Think of it as a customer context object the agent can ask for. A workable version has these parts:

PartWhere it comes fromWhat it lets the agent answer
Identity: person, emails, platform accounts, companySupport tools, chat workspaces, work-email domains"Is this the same customer who wrote in last month?"
Account: plan, billing status, amountBilling system (Stripe), CRM"Is this a paying customer, and on which plan?"
Conversation history, verbatim, with linksChat, support tickets, email, call transcripts, issue comments"What exactly did they say, and where?"
Grouped concerns: which reports describe the same issueA grouping step over the conversations"How many companies hit this, and since when?"
Work state: linked issues and pull requestsLinear, Jira, GitHub, GitLab"Is someone already fixing it?"
Product usageProduct analytics (PostHog, Mixpanel)"Do they actually use the feature they're asking about?"

Two rules make that object trustworthy.

Keep facts separate from interpretation. Store the customer's words and the billing record as facts. Treat summaries, labels, and priority as interpretation that cites the facts. When an agent says "this is a top request," you want to trace it back to the messages that support it.

Write down decision rules. Which system wins when they disagree (billing for plan, the CRM for account owner, the support tool for ticket status). What counts as "affected." Which actions need a person to approve them. Agents follow written rules far better than implied ones.

Build it yourself: what data, joined how, exposed how

You can build this without buying anything. These are the components, in the order that pays off.

1. Collect conversations verbatim

Pull messages from the places customers talk: support tools, Slack or Discord channels, email, call transcripts, GitHub or Linear issues. Keep each message with its author, source, thread, timestamp, and a link back to the original. A summary alone can't be quoted or re-checked later. Most vendors' APIs and webhooks support this; several also ship their own MCP servers (see our list of MCP servers for customer feedback).

2. Resolve people and companies

Match people across tools by exact email first, then by platform account. Derive companies from work-email domains, skipping free providers such as gmail.com. Platforms that don't give you an email (Discord is the usual one) stay separate until someone merges them by hand. Fuzzy name matching is tempting and produces wrong merges; if you add it, make it a suggestion a person confirms.

3. Join account data

Match billing customers to people by email and to companies by domain, so plan and status hang off the company. Keep it read-only and refreshed by webhook or on a schedule, so the agent never reasons from last quarter's plan.

4. Group by concern, not by message

One support thread can raise three separate problems. Split conversations into the concerns inside them, then group concerns that describe the same issue by meaning. Embeddings help find candidates; a model or a person confirms the match. Keep a link from every group back to the messages it came from, and keep history when a report moves from one group to another.

5. Expose it to the agent

Pick how the agent pulls context when it needs it:

  • An MCP server is the most direct route for Claude, Claude Code, Cursor, ChatGPT, and other MCP clients: the agent sees your tools and calls them mid-task. Give it a read-only search tool first. Add write tools later, scoped to what the signed-in user could do anyway.
  • A retrieval endpoint works for your own agents and pipelines: a natural-language search over the grouped concerns that returns rows plus the evidence behind them.
  • A context file (CLAUDE.md, .cursor/rules, AGENTS.md) still has a place for stable facts and for one instruction: "query the customer context server before answering questions about what users want."

Whichever you pick, bound it: per-organization scoping, read-only by default, rate limits, page sizes, and a flag on each result that says whether it is complete or truncated, so the agent knows when it is guessing.

6. Keep it fresh, permissioned, and tested

Ingest continuously rather than in batches someone has to remember. Respect source permissions: a private channel your bot wasn't invited to should stay out. Then test with the questions your team actually asks (see the checklist below), and when someone corrects the agent, record the correction where the next session will see it.

Which existing tool covers which layer

No single product category does all six. Each of these is the right tool for part of the job:

Tool typeExamplesGood atYou still build
Vector databasePineconeFast meaning-based recall over text; Pinecone now also sells Nexus, a knowledge layer above its databaseIdentity, account joins, grouping, freshness
Graph databaseNeo4jTyped relationships and traversal, with vector indexes in the same storeExtraction, identity resolution, freshness
Temporal memoryZep and GraphitiFacts with validity windows that expire when superseded; Zep sells a hosted platform with an MCP serverYour customer schema and source connectors
Agent memoryMem0, LettaPer-user or per-agent memory; Mem0 offers graph memory on its Pro planIdentity resolution across Discord, Zendesk and your CRM
Enterprise searchGleanEmployees searching company apps with source permissions enforcedFeedback grouping and billing joins
CRMSalesforce, HubSpotThe system of record for accounts and deals; both run official MCP servers with read and write toolsThe conversations, and the grouping of what customers said

A CRM and a context layer coexist. The CRM stays the record for accounts and deals; the context layer adds what customers said and which of those reports are the same issue. Our guides on what a customer context graph is and how to build one go deeper on the data model.

The managed version: Modem

Modem does steps 1 through 5 for you, for the sources it supports. Here is what it stores and how agents reach it.

What it stores. Messages from the sources you connect, kept verbatim with their author and a link back. Native sources include Slack, Discord, Microsoft Teams, Intercom, Zendesk, Pylon, Plain, Front (beta), Canny, email (by forwarding to an inbound address), Gong and Fathom calls, Linear, Jira, Jira Service Management, GitHub, GitLab, and Sentry user feedback; anything else can be pushed through the Ingest API. Modem splits conversations into the separate concerns inside them, groups those into topics by meaning, and classifies each topic as a bug report, feature request, complaint, praise, or discussion. Each topic links to the exact messages behind it, to the people and companies who raised it, and to the Linear, Jira, or GitHub issues and GitHub or GitLab pull requests that address it, matched by what they say. Our site calls this whole structure a context graph; underneath it is a relational database with typed links and change history, and embeddings are one search signal, not the store.

Who the customer is. People are matched by exact email and platform account. Companies come from the Slack workspace (for Slack Connect, the partner's workspace) or the work-email domain. Discord authors and others without an email stay separate until someone merges them, and merges can't be undone yet.

What plan they're on. Connect Stripe and Modem matches Stripe customers to people and companies; the agent joins billing to topics when you ask, for example "which paying customers reported bugs this month?" Salesforce is a read-only sync the agent can query; it isn't stored on companies. There is no HubSpot integration. Plan data is not on topics, and a topic's priority is an AI-assigned score that weighs severity, how many companies are affected, recurrence, and more, not revenue.

How agents reach it.

  • MCP server at https://mcp.modem.dev/mcp, in beta. OAuth sign-in, no API key. It has 16 tools: three read-only ones (search_modem, which answers a plain-language question with a short answer, matching rows, and flags for partial or truncated results, and spends no agent credits; modem_docs; modem_skills), four tools that start and manage full Modem agent runs (modem_agent_invoke and its companions, which run asynchronously and use agent credits), and nine record tools that edit people, companies, and topics as the signed-in user, under that user's role. A token with only the data:read scope sees just the three read-only tools. Calls are limited to 20 per minute per organization per tool. Agent turns started over MCP have no approval prompt, so a run acts on what you ask with the integrations your organization has connected.
  • REST API at https://api.modem.dev/v1 with an organization API key (120 requests per minute) for topics, people, companies, and history. It doesn't return a topic's linked issues or pull requests; use search_modem or an agent run for those. The @modem-dev/cli package is a command-line client for this REST API, so it sees what the REST API sees, not what the MCP server's search sees.
  • Ingest API for pushing conversations from sources Modem doesn't connect natively. It writes only.

Add it to Claude Code with one command:

claude mcp add --transport http modem https://mcp.modem.dev/mcp

Setup for Cursor, VS Code, Codex and others is in the MCP docs. In ChatGPT, install the Modem plugin, sign in, and pick Modem with @ in a chat. Then add one line to your agent instructions, such as "query Modem for customer feedback before answering questions about what users want," so the agent knows to ask.

Why we built it this way. Customer signal is scattered, duplicated, and unresolved, and an agent given raw access re-reads threads in every session and still misses half of it. So Modem does the synthesis when messages arrive, not when the agent asks: the agent reads topics with the evidence attached instead of re-deriving them. The read-only search is kept separate from the paid agent, and gated by its own scope, so the common question ("has anyone reported this?") costs no credits and can't change anything.

Limits. Modem answers only from sources you connected and from what their backfill windows imported (30 days for most sources). In Slack it reads only channels it's subscribed to, and never DMs. It doesn't read your source code. Agent memories and the Org Skills your team writes persist across Modem conversations but aren't exposed over MCP or REST. No time from message to topic is promised. When someone asks, the Modem agent can also hand a fix to Claude Code, Cursor, or Devin through each vendor's own API, with a brief it writes from the topic; that is a separate path from MCP, covered in how to have your coding agent fix user-reported bugs.

When Modem is the wrong pick. If your agent's questions are about deals, pipeline, and account owners, connect your CRM's own MCP server. If you're building per-user memory into your own app, a memory layer like Mem0, Zep, or Letta is the right shape. If you need employees to search documents across the company, use enterprise search such as Glean. And if one tool holds nearly all your feedback, that tool's own MCP server is less to set up.

An illustrative example: one question, with and without context

This scenario is illustrative, not a single customer's story. An engineer in Claude Code asks: "Customers keep complaining about export timeouts. Find out who's affected and propose a fix."

Without customer context, the agent searches the codebase for "export" and "timeout," finds a plausible slow query, and proposes an index. It may be right. It has no way to know.

With Modem connected, the agent first calls search_modem with the question. It gets back a short answer and the matching rows: a topic grouping reports about export timeouts from a Discord channel and Zendesk tickets, the companies behind them, the customers' own words (including one that mentions the row count where exports start failing). If the result is partial, the response says so. Because Stripe is connected in this example, the agent then asks the Modem agent (modem_agent_invoke) which plan each of those companies is on, since joining billing data is the agent's job. The agent starts from the reported failure mode instead of a guessed one, and the engineer can see which customers to tell once the fix ships.

Test it with the questions your team actually asks

Whatever you build or buy, check it against real questions before you trust it:

  1. "Who reported this bug, at which companies, and what plan is each on?"
  2. "Has this come up before, and did we already fix it?"
  3. "What are the top requests from paying customers this month, with quotes?"
  4. "Is this one issue or several that look alike?"
  5. "Which issue or pull request is addressing this topic?"

If the agent can't answer one, find which layer is missing (identity, billing, history, grouping, or work links) and fix that layer. When a person corrects an answer, put the correction where the next session will read it.

FAQ

Why does my AI agent give generic answers about my customers?

Because it can't see customer-specific data. It doesn't know who the customer is across your tools, what plan they're on, what they've said before, or which reports are the same issue, so it falls back to general knowledge. Fix it by giving the agent a queryable context layer, not a longer prompt.

Is RAG or a vector database enough to give agents customer context?

Not on its own. Retrieval finds text that reads like the question, which works for documents and policies. Customer questions need joins: the same person across tools, a person to a company, a company to a plan, and many reports to one issue. You need identity resolution and grouping alongside the vector search.

How is a customer context layer different from a CRM?

A CRM is the system of record for accounts, contacts, and deals. A context layer holds what customers said, grouped by issue and linked to the people and companies behind it. They work side by side: Modem, for example, reads Salesforce as a read-only source the agent can query and doesn't replace it.

Can I give Claude Code or Cursor customer context without Modem?

Yes. Connect the MCP servers your tools already ship (Intercom, Linear, Sentry and others), keep stable facts in CLAUDE.md or .cursor/rules, and paste exports for one-off tasks. What you give up is the cross-tool join: no single-tool server knows that a ticket and a Slack thread describe the same bug. Our guides to Claude Code and Cursor compare the options.

Does Modem put revenue on every topic?

No. With Stripe connected, billing data sits on the matched people and companies, and the agent joins it to topics when you ask. Topic priority doesn't use revenue. Salesforce data is queried by the agent on request rather than stored on companies.

Can an agent connected to Modem over MCP change data?

Yes, if its token has the agent:invoke scope: the nine record tools act as the signed-in user with that user's role, and merges can't be undone. Agent runs started over MCP have no approval prompt. For read-only access, use a token limited to the data:read scope, which exposes just search_modem, modem_docs, and modem_skills. If you're weighing what customer data a coding agent should see at all, read is it safe to give your coding agent access to customer support data?