nicolas-dev

From 2,244 open issues to under a thousand in about three weeks: what I take from the Next.js case.

Nicolas Gonzalez · 4 min read
  • claude-code
  • coding-agents
  • nextjs
  • subagents
  • automation
  • mcp
  • ai
  • productivity
Leer en español →

On August 10, 2026, the Next.js backlog stood at 2,244 open issues. About three weeks later, by the team's own account, it was under a thousand, and 218 new reports had come in between those two dates.

The numbers are in a post published on September 4, whose title says "one month". Do the subtraction (2,244 minus 1,462 closed plus 218 new) and you get exactly 1,000, so I take the figures as an order of magnitude.

The short explanation, per the post, is a research agent and maintainers who read the evidence behind every result, plus some closures made outside that review.

I read it as someone who uses agents every day and looks first at what the agent was allowed to do, and what it wasn't. In my own workflow I usually let an agent do a first review pass before I open a PR. The call is still mine.

Seeing that same logic applied to thousands of issues seemed like a good case to take apart.

What exactly did the agent do?

The repository gets an average of 36 new reports a week, and the backlog had peaked at 3,109 open issues in January 2025. By their account, coding agents made it easier to file detailed reports, and that means a much higher volume to review.

They had already tried marking issues stale and closing them after inactivity, first at two years and later at 18 months. It helped bring the backlog down, but inactivity turned out to be a poor stand-in for relevance: they closed some reports they should have kept.

So they built closability, an agent that runs on eve, Vercel's open-source agent framework. For each issue it works in a clean sandbox with the repository, Node, Playwright and Chromium.

It reads the conversation and the supported versions, searches related issues, PRs, commits, releases and documentation and, when needed, tries to reproduce the bug on the reported version, the latest stable release and canary. At the end it looks for evidence that contradicts its first conclusion.

It returns structured data: a confidence score, the main reason, a summary, the evidence and the references. Each investigation took about 30 minutes on average, and to cover the whole backlog they gradually raised concurrency until 200 sessions were running at once.

One detail I appreciated: the run over the full backlog used GPT-5.6 Luna with reasoning effort set to max. The post doesn't say which models the other agents use. What you can copy is the design, not a specific model.

The result was 1,462 issues closed, with the note that the count includes a few closures made outside this review.

The limits they set

What struck me most was what they didn't let it do. closability is one of several agents in their "Maintainer Agent" (others reproduce, verify, bisect, write e2e tests and prepare fixes). Outside its sandbox it can only read: it doesn't comment, close, push or deploy.

They also configured it to ignore instructions it finds inside issues or the repository, a defense against prompt injection. And its confidence score is conservative: failing to reproduce a bug isn't enough to recommend closing it.

Per the post, maintainers read the evidence behind every result in a close queue. In most cases it was enough to read the summary and check the sources the agent had already gathered.

By the agent's own classification, of the closed issues 37% were already fixed, 19% were duplicates and 16% were expected behavior. Another 17% landed in "other".

Before starting, they added a GitHub Action so an issue can be reopened if closing it was a mistake, and later used it to check the results. I also read it as what gave them room to move fast: if they got one wrong, there was a way back.

The autonomy came later. As of the post, they let agents close the clearest cases on their own: the first agent reviews and, if its score is 80 or higher, a second agent looks for reasons to keep the issue open. If both recommend closing, the second one picks the reason and writes the comment, and a separate GitHub Action closes the issue.

They started with up to 25 a week, and people can still reopen an issue if the call was wrong. Every Monday, the agent also researches up to 100 issues that have had no activity for at least 30 days. Every code change still goes through human review before it lands in Next.js.

What do I take for my own work?

Simon Willison wrote on September 24 that agents make engineering even harder: getting their full potential takes extraordinary discipline and knowledge. To me, Next.js shows it well: a good part of the result came from the system they built around the agent.

My small version of all this is this very blog. It has an MCP where an agent creates the articles: I tell it what I want to write and the draft shows up in my panel, in Spanish and English, linked together. This article arrived that way.

Deleting only works on drafts, and the code enforces that. Publishing is my call.

Sources

Nicolas Gonzalez

Full Stack Developer with 6+ years on product teams in the US and Latin America. I write about applied AI, legacy migrations and React.

Let's connect on LinkedIn ↗
Keep reading
See section →