← All writing

How to detect bot-created GitHub issues when the [bot] check misses them

You already skip every author whose login ends in [bot], and a nightly job still filed forty issues into your tracker last night, one every few seconds, under a maintainer's personal account.

The filter is not broken. It is answering a narrower question than the one you have.

What GitHub's own signal covers

Two fields in the webhook payload carry it. The actor's user.type is the string Bot, and the actor's login ends in [bot]. Together they catch GitHub Apps posting as themselves, which is most of the obvious traffic: github-actions[bot], renovate[bot], dependabot[bot]. Check this first. It costs nothing and it never flags a person: GitHub sets both fields itself, and a login with brackets in it cannot be registered by hand. What it does instead is miss things, which is the rest of this page.

One detail worth deciding rather than discovering. In the implementation this page describes, the type comparison ignores case and the suffix check is a plain ends-with, so it is case-sensitive. Those two halves do not have to agree, and nothing tells you which way you chose until a login with different casing walks through.

Skipping is the right outcome, but not in silence. An owner who sees nothing happen on forty issues will assume it is broken. Post one notice explaining why the robot was skipped, scoped to the repository rather than the issue, at most once a month, and reserve real silence for traffic the author never sees anyway, such as your own comments and pull request threads. If your tracker is the thing filling up rather than the filter you are building, the maintainer's version of this is the same problem from the other side.

The case it misses

A human account running a script. A personal access token in a cron job, an Actions workflow authenticating as a maintainer, a one-off migration importer someone wrote on a Thursday. The account type is User, the login has no suffix, and the avatar is a photograph of a person who is asleep.

Nothing in the account metadata separates that from a real report, because there is nothing to separate: it is a real account. What separates them is pace. In our own data the bulk filers are easy to see once you look at gaps rather than identities. Three accounts on the repositories we watch filed between 45 and 117 issues each, and between 76 and 92 per cent of each account's issues arrived less than a minute after that same author's previous one. Three accounts is a small population and the figure is ours rather than a fact about GitHub, but no amount of reading those profiles would have shown it.

Three signals, two points, one decision

Scoring beats a single test, because every individual signal here has an honest counter-example. Three signals, one point each, two points calls it a batch.

  • Earlier issues by the same author inside a thirty-minute window, one point each, capped at two. The cap matters: it means a burst of three is decisive on its own, and a burst of thirty is not more decisive than that.
  • The issue was created through a GitHub App on the author's behalf, which the payload reports as performed_via_github_app being present and not null. One point, never two, because plenty of integrations post through an app for a real person.
  • An automation phrase in the body. One point, for the same reason.

A burst of two is a person having a bad morning. An app alone is a person using a tool. Only the combinations are worth acting on.

Where the burst line sits

These are the numbers a future change will move by accident, so pin each one with a test that fails loudly. Named as they are in the suite: aFirstIssueIsConversational, because a first issue is always a person. oneEarlierIssueIsStillAPerson, because two bugs found in one sitting is ordinary. threeInTheWindowIsABatch, which is the line itself. issuesOutsideTheWindowDoNotCount and theWindowEdgeCounts, which fix the window at exactly thirty minutes and settle the boundary rather than leaving it to whichever comparison operator got typed. And issuesFromTheFutureAreIgnored, which is the one nobody writes until it bites: your webhook host and your database do not share a clock, and an arrival time a minute in the future will otherwise manufacture a burst out of nothing.

The automation marker, and the sentence that must never match it

The pattern covers the phrasings that actually appear: auto-generated and autogenerated, generated automatically, generated by a script (or a bot, tool, workflow, pipeline or action), created automatically, an automated issue, automated report or automated ticket, and the long form, this issue was created, opened or filed automatically, or by a script, bot, workflow or action.

It has a real limitation worth knowing before you copy it. The generated-by branch requires an article, so generated by a script matches and generated by our CI pipeline does not.

Write the false-positive test before you write the pattern, because the failure mode is expensive and specific. A person describing a generator is not a generated issue, and they have usually written you exactly the kind of report a bug report guide asks for: a concrete artefact, a trigger, a browser. That test is ordinaryProseIsNotFlagged, and it is the one to write first.

What the pattern should catch

Auto-generated from failing test run #4412. Do not edit.

What a naive match on "generated" does instead

The PDF generated by the export button is blank on Safari.

Now the honest part, which is the most useful thing on this page. Across 238 stored real issue bodies on the repositories we watch, that pattern has never once matched. Not a false positive: no positives at all. performed_via_github_app has never been the difference on its own either. Pace is the only signal that has ever decided anything in our data. Build all three, because the other two cost a regex and a null check, but do not expect them to earn their place.

Why a burst cannot see itself

This is the bug that makes the whole thing look like it works while it does nothing.

If you count an author's recent issues by querying whatever your system stores for them, you are querying records written after your reply posted, half a minute or more after the issue arrived. Across 190 burst issues in our own records, 78 per cent arrived within thirty seconds of the same author's previous one, with a median gap of twelve seconds. The entire burst lands before the first stored record exists. Every issue in it looks like a first issue, the classifier returns the friendly answer forty times, and every log line is correct.

The fix: a filing record written before anything slow

One tiny document per issue, written the moment the webhook lands: author, repository, issue number, arrival time. Before any analysis, any model call, any token fetch. It is the only write in the request that has to be fast, because it is the only one the next issue in the burst depends on.

Four details that turn out to matter. Key it on repository plus issue number, so a redelivered webhook overwrites rather than double counts and a retry cannot invent a burst. Query it as a range on arrival time alone and filter the author in memory, so it needs no composite index and no migration when you add a field. Give the records an expiry of a day when the window is half an hour, so the collection stays small on its own. And classify per author, never per installation: a repository shared by a release script and a colleague should still answer the colleague like a colleague. The tempting alternative, a list of known robot account names, is a maintenance job that grows with every customer and stops working silently the first time somebody renames an account.

After the record landed first, burst detection caught 54 of one prolific author's 69 issues in our data. The fifteen misses are the first one or two of each burst, before the window has anything in it, which is the expected shape rather than a fault.

When the lookup fails, and what happens next

A failed read must never cost the reply. Catch it, log it, and fall back to the human treatment rather than the automated one, because the two mistakes do not cost the same. A script treated as a person gets one warm message that nobody reads. A person treated as a script gets a curt reply to the report they sat down and wrote, and remembers it.

The one thing WhatProblem uses this score for is choosing the register a reply is written in, conversational or stripped back to the substance, before the prompt is built at all.

What the separation is actually for is what happens after it. Machine-filed issues are uniform, complete-looking and missing intent; human-filed ones are short and ambiguous in the specific way that "it doesn't work" is ambiguous. A batch tends to resolve to actionable or not ours in bulk, without a clarifying round trip each, which is the cheapest end of the order a triage pass should run in. Mixed together, the batch buries the few issues a person wrote by hand and waited for an answer to.

WhatProblem asks these questions for you

It reads new GitHub issues and asks what is missing, in the issue thread, before anyone on your team has to.

Install from GitHub Marketplace