How to have your coding agent fix user-reported bugs
Short answer. You can automate every step from a user's report to an open pull request. You should not automate the merge, and you should not let an agent write the fix until the bug has been reproduced. The pipeline has seven steps: capture the report with its environment details, merge duplicates and size the cluster, decide whether an agent should touch the bug, reproduce it and write a failing test, brief the agent, review the PR against that test, and tell the people who reported it. Reproduction is the gate. With no reproduction and no failing test, the agent is patching a guess, and a guess that compiles is harder to catch than one that doesn't. Modem (our product) covers the first two steps and the briefing. It reads reports from Slack, Discord, email, support tools and GitHub issues, groups them into topics, and when a teammate asks, its agent writes a brief and hands it to Claude Code, Cursor or Devin. Modem does not reproduce the bug, read your code, write the patch or review the PR. The coding agent and your team do that.
"Automatically fix bugs reported by users" is how most people phrase this, and the accurate version is "automatically up to a reviewable PR." The fix is rarely the hard part now. What breaks these pipelines is the input: a report in a chat thread that says "export is broken," with no version, no steps, no logs, and no idea whether three other customers hit it too.
This guide is vendor-neutral. Every step works by hand or with tools you already run, and the table in step 5 compares the real options. We build Modem, so weigh its parts with that stake in mind. It appears where it fits and stays out of the steps it doesn't touch.
The pipeline at a glance
| Step | What happens | Who decides |
|---|---|---|
| 1. Capture | The report arrives with version, steps, environment and logs, or someone gets them | Automated, plus a person for missing details |
| 2. Dedupe and size | Duplicates merge into one cluster with a count of reporters and accounts | Automated |
| 3. Choose | Is this a bug an agent should touch at all? | A person |
| 4. Reproduce | A failing test or a recorded reproduction exists | Agent or person. No repro, no fix |
| 5. Brief and delegate | The agent gets the evidence, the failing test and the limits | Automated, or a click where approvals apply |
| 6. Review | A person reads the PR against the failing test | A person |
| 7. Tell the reporters | The people who reported it hear what happened | A person triggers it, after it ships |
1. Capture the report with what an agent will need
Several AI answers to this question start from an in-app bug reporter that captures the page URL, console errors, failed network calls, app version and a redacted account ID. If you have one, wire it to your tracker and you have the best input there is.
Many teams don't, especially developer-tool companies whose users report in a shared Slack channel, a Discord thread or a support ticket, often with some version of "it works on my laptop." Those reports carry none of that data. The pipeline still works, but step 1 has to say who gets the missing pieces and how. A short intake checklist is enough:
- Version and environment. App or SDK version, OS, browser, region or deployment.
- Steps, expected and actual. The exact sequence, and what the user thought would happen.
- Evidence. Error text, a screenshot or recording, a request ID, a link to the error-tracker event if there is one.
- Who and how often. The account, and whether it happens every time.
Decide who asks for the gaps. A support or customer-facing teammate replying in the original thread is the usual answer, and it keeps the reporter involved from the start. For bugs that throw an error, an error tracker such as Sentry holds the stack trace and context the user never will, so link the event to the report.
Also sort the report: is it a bug in your product or in how the customer integrated it? The second kind is a support answer first, not an agent task.
2. Merge duplicates and count who is affected
"Uploads fail on large files" in Slack, "attachment upload spins forever" in a ticket and "can't attach anything since the update" in an email are one bug. Handed to an agent as three tasks, they become three PRs circling one root cause. Collapse them first and keep the count: how many people, how many accounts, since when. That count is also your best evidence for step 3, because "five accounts since Tuesday" and "one report, one account" deserve different treatment.
By hand, this means searching the tracker for the error string or the screen name before filing, and adding each new reporter to the existing issue instead of opening another. How to handle duplicate feature requests covers the routine, and it applies to bugs too. If the repo is public, keep customer names and private links out of issue comments and in your own log.
Modem does this grouping as reports arrive. It reads conversations from connected sources, groups reports that mean the same thing into one topic, and keeps the original messages and the people and companies behind them. It classifies the topic as a bug report, feature request, complaint, praise or discussion, not each message. Priority on a topic is an AI-assigned score that weighs severity, how many companies are affected and whether it recurs, among other factors, so it is not a raw count of reporters. Feedback aggregation is the general category if you are comparing tools.
3. Decide whether an agent should touch it (human gate)
Choosing which bugs get agent time is prioritization, and a person with context makes that call better than a threshold. A well-reproduced bug with three affected accounts is a clean delegation. A one-line report that hints at data loss is a conversation first.
Some categories are worth keeping human-led however clear the report is:
- authentication and authorization
- payments and billing
- data migrations and schema changes
- security-sensitive fixes
- incidents that affect many customers at once
An agent can still help on these by gathering context or drafting a test, but a person owns the change. Put this list in your repo instructions file (AGENTS.md, CLAUDE.md or the equivalent) so the agent sees it too.
4. Reproduce it first (the gate that matters)
This is the step the usual pitch skips, and it is where most failed attempts go wrong. A patch that makes the error disappear is not evidence the bug is fixed. It can hide the symptom, miss the root cause or add a side effect, and the reviewer sees a tidy diff with no way to tell.
The rule is no verified reproduction and no failing test, no autonomous fix. In practice:
- Have the agent, or an engineer, reproduce the bug in a sandbox from the report's steps and environment, and write a test that fails because of the bug.
- If it cannot reproduce the bug, stop. The agent's output is a list of what is missing, which goes back to the reporter through whoever owns the thread. That is the pipeline working, not failing.
- Only with a failing test in hand does the agent write the patch. Done means that test passes and the existing tests still do.
The failing test also gives your reviewer something concrete to check in step 6. It is cheap insurance, and it is the line between "an agent fixed a bug" and "an agent changed code near a bug."
How much of reproduction is automated depends on the tool. Error-driven tools such as Sentry's Seer start from stack traces and telemetry, which helps when the bug throws an error. For bugs with no event, reproduction comes from the report's steps. Modem does not reproduce bugs. It supplies the report and the evidence around it.
5. Brief the agent and delegate
An agent handed a thin task guesses at the gaps, and it guesses confidently. The brief does the load-bearing work. A usable one has:
Problem: the bug in the users' words, with 1 to 3 short quotes
Impact: how many people and accounts, since when
Environment: versions, platform, config, error text, request IDs
Reproduction: the steps, and the path to the failing test
Done when: the failing test passes; named existing tests still pass
Non-goals: files, behavior or areas the change must not touch
Open a pull request. Do not merge.How to give coding agents a backlog they can execute covers this contract in more depth, and how to give Cursor context about a specific customer bug report shows one tool's version.
Two cautions before text goes to an agent. First, a user's report is untrusted input. It can contain text that reads like an instruction, so tell the agent to treat quoted reports as data, and review what it changes. Second, redact secrets and personal data from the report before it goes into the brief, because the task text goes to the agent vendor's servers.
Where the work can start
Each coding agent has its own front door, and several can start from a ticket or a chat thread directly. This table reflects the vendors' own docs as checked on 2026-10-05. Plans and pricing change, so check each vendor's page before you commit.
| Agent | How a report reaches it | What it does with it |
|---|---|---|
| GitHub Copilot cloud agent (formerly the coding agent) | Assign a GitHub issue to Copilot, mention @GitHub in Slack or Teams, or assign it in Jira, Linear or Azure Boards | Plans, then opens a draft PR it keeps updating. The person who asked for the PR cannot approve it, and Actions workflows on its PR wait for someone with write access to approve them |
| Cursor Cloud Agents (formerly Background Agents) | @cursor in Slack (it reads the thread), delegating to Cursor or mentioning @Cursor in Linear, GitHub events, the API, or automations | Works in a cloud environment. Over the API a PR is opened only if autoCreatePR is true. The default pushes a cursor/... branch |
| Devin | @Devin in Slack, assigning or labeling in Linear or Jira, or Automations (GitHub, GitLab, Slack, schedules, webhooks) | Opens a PR and posts the link back on the ticket. Engineers review before merge, under the same branch protections as a person |
| Claude Code | @claude in an issue or PR comment, or in a new issue's title or body, through a GitHub Actions workflow you add. Also @Claude in Slack and cloud sessions | The Action can push commits and open PRs. In the Slack and web flows, a person clicks Create PR. Nothing watches your issues until you add the workflow |
| Sentry Seer | An error issue in Sentry. Automation triggers on issues with 10 or more events in the last 14 days and a sufficient fixability score | Root cause, a plan or a drafted PR, at the stopping point you set. It never merges, and it can hand off to Claude Code, Cursor or Copilot |
| Linear's agent | Delegate an issue to the agent, or use triage rules to hand issues over on arrival | A coding session runs, opens a PR and shows the diff on the issue. Needs Linear's Business or Enterprise plan |
A few points that matter for the bug-report case:
- Seer starts from errors, not from user words. A report that never produced an event, or produced only a few, will not auto-trigger it. It is strong where the bug throws and weak where it doesn't.
- Claude Code runs only if you set it up. The GitHub Action responds to
@claudeor a new issue once the workflow is in your repo. Routines can run on a schedule, an API call, or GitHub pull request and release events, but their documented GitHub triggers do not include new issues. - Most of them can be steered mid-run. Devin subscribes to the Slack thread it posts in, so replies reach it. Cursor takes follow-up guidance through an
@Cursorcomment on the Linear issue. Claude cloud sessions accept queued follow-ups. Check the vendor's flow before you rely on it. - No vendor reviewer is documented to check a fix against the customer's original report. Cursor's Bugbot, Devin Review and Claude Code Review examine a diff, and none of the docs we checked say they compare it to the report. That comparison is your reviewer's job, which is why the report and the failing test belong in the PR.
Handing off from Modem
When a teammate asks, Modem's agent can write the brief and make the handoff, from the dashboard or from Slack. Before delegating it can query Modem for the relevant reports, error patterns and affected users, and it typically carries that evidence into the task text. The brief is free text the agent writes, so what it includes varies. Ask for the sections above.
The agent then calls each vendor's own API with your key (this is not MCP). An admin sets an organization Anthropic API key for Claude Code, plus a GitHub token for tasks that touch a repository. Each member adds their own key for Cursor and Devin. Claude Code and Cursor are set up to open a PR. Devin receives only the prompt, with no repository field, and works in whatever repos it has access to itself.
What you can do afterward:
- Check status from the conversation that started the task. The PR link appears when the vendor reports one, and for Claude Code only if the agent printed it.
- Follow up. Claude Code and Devin accept follow-up messages through Modem. Cursor has no follow-up tool in Modem, so steer a Cursor run in Cursor.
- Cancel or edit in the vendor's tool. A task can't be recalled from Modem once it is sent.
Modem does not link the delegated task to the topic. If the repo is connected to Modem's GitHub integration, the PR the agent opens is read like any other, and Modem tries to connect it to the topic by what it says. Naming the issue in the PR description helps, since Modem follows explicit references one hop. Treat that as best effort, not a guarantee.
Why a person sees it first. Modem is the briefing and orchestration layer, not a coding agent, and handing work to an outside agent is the kind of action a person should see before it happens. So in Slack and, by default, in the dashboard, Modem shows the task and waits for an approve or deny click. The tradeoff is that delegation there is a click, not a hands-off pipeline. Automations, Discord, Microsoft Teams and agent runs started over MCP have no approval step, so what you ask for runs as written. We haven't confirmed that an automation can start a coding agent end to end, so this guide doesn't promise it.
Your coding agent can also pull the evidence itself. Modem's MCP server (in beta) has a read-only search_modem tool that answers questions like "who reported the export timeout, and which accounts?" without spending agent credits. A token with only the data:read scope sees only the read tools. Calls are limited to 20 a minute per organization per tool. Setup is documented for Claude Code, Cursor, VS Code with Copilot and Codex.
6. Review the PR against the failing test (human gate)
An engineer reads the agent's PR before merge, the same as anyone else's. The earlier steps are what make that review cheap: a failing test that now passes, a "done when" list, and non-goals to check the diff against. Look specifically for:
- changes to files the task didn't need to touch
- a test edited to pass instead of code fixed so it passes
- a fix that handles the reported case and misses the root cause
- anything in the human-led categories from step 3
Branch protection enforces this, not goodwill. Copilot blocks the requester from approving its PR, Seer never merges, and Devin follows the branch protections a person would. For Claude Code, Cursor and the rest, your own rules are the control, so require a review and passing CI on any branch an agent pushes to. If you let an agent auto-merge anything, restrict it to the narrowest class of low-risk, well-tested fixes you can name, and widen it only after you have watched it work.
7. Tell the people who reported it
The reporters are attached to the topic or issue you built in step 2, which turns "tell them it's fixed" into a message instead of a search. Two cautions. A merged PR is not a shipped release, so wait until it is live before anyone writes to a customer. And the reply should come from a person, or from a draft a person approves. How to close the feedback loop with customers covers the mechanics.
In Modem, an opt-in automation template posts an internal Slack note naming who asked when a PR directly linked to a topic merges. That is all the merge does. Customer replies are written by the agent when a teammate asks, and in Slack and, by default, in the dashboard they wait for approval before going to a customer-facing channel.
An illustrative run
This is a made-up example, not a recorded result. A user writes in a shared Slack channel that CSV export fails for their workspace. A teammate replies asking for the app version and the size of the export. Two more reports of the same failure arrive in a support ticket and an email, and all three sit on one topic with three accounts attached. An engineer decides it is a good agent candidate: no auth, no billing, no schema. The agent reproduces it against a large test workspace and writes a test that times out. The brief carries the quotes, the three accounts, the failing test and the instruction to open a PR and not merge. The agent changes the export to stream, the test passes, and an engineer reviews the diff against the test and merges. After the release goes out, a teammate asks the agent to draft a reply to the three reporters.
FAQ
Can an AI coding agent fix user-reported bugs fully automatically?
It can automate everything up to an open pull request. The vendors' own defaults stop short of merging: Seer never merges, Copilot blocks the requester from approving its PR, and Devin follows your branch protections. Fully hands-off report-to-merge is possible to wire up, but we advise against it beyond a narrow class of low-risk, well-tested fixes.
What if the report came from chat and has no logs or version?
Then step 1 is a person's job. Reply in the thread and ask for the version, steps and error text before the report goes near an agent. If the bug throws an error, link the error-tracker event. If it can't be reproduced, the right outcome is a list of what is missing, sent back to the reporter.
How do I stop the agent from patching a guess?
Make reproduction the gate. Require a failing test before any patch, define done as that test passing without breaking others, and have a person review the PR against the test. Without a reproduction, the agent's job is to ask for more information.
Is it safe to feed raw user reports to a coding agent?
Treat them as untrusted input. A report can contain text that reads as an instruction, and it can contain secrets or personal data. Redact before briefing, tell the agent to treat quoted reports as data, and remember the task text goes to the agent vendor's servers.
Which coding agent should I use?
The one your team already uses and can review. Copilot cloud agent suits teams on GitHub who want independent approval enforced. Cursor and Claude Code suit teams already working in them. Devin suits ticket-first teams in Slack, Linear or Jira. Seer suits bugs that throw errors, and Linear's agent suits teams whose tracker is the system of record. None of them captures or dedupes customer reports on its own, which is what steps 1 and 2 are for, whether you do them by hand or with Modem.
Does Modem fix the bug?
No. Modem groups the reports, and when a teammate asks, writes the brief and hands it to Claude Code, Cursor or Devin. It does not read your source code, reproduce the bug, write the patch or review the PR.
