How Do You Tell If a Bug Report Is Too Vague to Hand to Devin?
Run the report through four questions before it becomes a Devin task. Does it describe one bug or several bundled together, does it point at one code path or could it be any of three, can someone reproduce it on demand, and is there a concrete way to check a fix actually worked. If the answer to any of those is no, the report needs a human to look at it first, not an agent. Devin's own guidance backs this up directly. Its guide to when Devin is a good fit says tasks need "a clear start and end, plus explicit success criteria" and that "the more context, the better," which is another way of saying an agent given an unbounded, underspecified problem doesn't fail loudly. It fills the gaps with something plausible and hands you back a PR that reviews clean.
That's the trap worth designing around. A vague report doesn't make Devin error out. It makes Devin confident about the wrong thing, and the four questions below are a fast way to catch that before a session starts rather than after a diff needs untangling.
Cognition's own docs draw the line between a task Devin can do and one it can't in terms that map directly onto a bug report:
- One bounded symptom, not several. "The export is slow, and sometimes it also drops rows" is two bugs wearing one ticket. Devin will pick one to chase, and there's no guarantee it's the one that matters.
- A named path, not a guess. If the report can't say whether the bug lives in the frontend, the export worker, or a third-party library, Devin has to guess too, and its guess becomes the scope of the whole task.
- Steps that reproduce it on demand. Not a paraphrase of what a customer said. A sequence someone actually ran and watched fail.
- A way to check the fix. A specific input that should produce a specific output, or a test that fails now and should pass after. Without this, Devin's confidence estimate is reading the ticket, not the bug.
None of these are exotic asks. They're the same four things an engineer writing their own bug report would include without thinking about it, because they already know the answers. A customer-reported bug arrives without any of them by default, and someone has to add them, or flag that they can't be added yet.
What "failing the checklist" looks like when nobody's checking
Pulling a failing report back and spending twenty minutes in the audit logs only happens if someone has the habit of running the four questions before routing. Absent that habit, a vague report sent straight to an agent doesn't fail visibly. Answer.AI's month-long test of Devin against 20 real engineering tasks logged 14 outright failures, and reading through them, several look less like Devin refusing and more like Devin committing hard to a wrong approach. Asked to pull recent papers from Google Scholar, Devin, in the team's words, "went into a rabbit hole of trying to parse HTML that it seems like it couldn't get out of. It got stuck and went to sleep." The team's own conclusion was blunt: "tasks that seemed similar to our early successes would fail in unexpected ways," with no reliable signal upfront for which one you'd gotten.
A vague bug report produces the same shape of failure on a smaller scale. Devin doesn't stop to ask whether "work orders sometimes show up wrong" is a display bug, a save-path race, or a training issue with the people using it. It picks whichever reading the surrounding code makes most legible, tests only that reading, and ships a plausible PR either way.
The checklist's blind spots
Running four questions in your head against every incoming ticket works while the queue is small. It stops working for reasons that show up in a predictable order as volume grows:
- The same vague symptom arrives more than once, and nobody notices. A second, third, and fourth version of the same vague report, each filed by a different rep in a different wording, each independently fails the checklist and gets set aside, without anyone connecting the four reports into the pattern that would make the "reproducible" box checkable.
- The checklist only works if the checker knows the codebase. You can rule out a subsystem you wrote yourself. A support rep running the same four questions has no way to guess which path is plausible, so the checklist becomes a formality rather than a filter.
- Nothing tracks which reports are sitting in limbo. A report that fails the checklist needs someone to come back and re-check it once more evidence exists. Without a place that holds "failed, pending investigation" reports, they just don't get re-examined until a customer escalates again.
The first and third gaps are where an aggregation layer beats a habit. Modem reads Slack, support tools, and sales call notes as they arrive, and clusters reports describing the same underlying symptom into one topic instead of leaving four vaguely-worded tickets to fail the checklist independently. A topic that's accumulated three reports of "work orders assigned wrong" carries that count and every account that hit it, which is itself evidence toward naming a path and a repro, the two hardest boxes on the checklist to check from a single report alone. When a topic clears the bar, the Modem agent can compose the brief and, when you ask, hand it to Devin directly, carrying the accumulated detail instead of restating one customer's original wording as if it were the whole picture.
A note on where this is coming from: Modem is what we sell. Check the specifics above against your own setup rather than take our framing of them at face value. A broader comparison of how teams make this handoff is in the six best tools for handing customer-reported bugs to Devin.
The checklist itself doesn't change. What changes is whether a report that fails it the first time gets forgotten or gets a second look once the second and third reports arrive. For the fuller pipeline this connects to, from a report's first mention through a merged PR a customer gets told about, see from user report to merged fix.
Run the four questions on the next report
Before the next bug reaches Devin: does it describe one symptom or several, does it name one code path, does someone have steps that reproduce it right now, and could you write down what a passing test looks like. A yes on all four means it's ready. A no on any of them means somebody needs about twenty minutes with the logs before an agent gets involved, the twenty minutes that turn an unroutable ticket into one an agent can actually finish.
