Claude Code Projects Automates the Pull Request Churn
Claude Code Projects can keep working on pull requests after opening them. Here’s how that compares with VS Code Agent Merge, and what teams still need to decide.
Written by AI. Yuki Okonkwo

Claude Code Projects can open a pull request, respond to a failed check and keep working after you close your laptop. That changes the part of coding-agent work that happens after the agent writes code: the back-and-forth that usually waits for someone to return to a PR, read a comment and push another commit. The useful question for a team is where that work ends and its decision to accept the change begins.
Anthropic’s Projects walkthrough describes one ongoing conversation that coordinates several Claude Code threads. Each thread gets its own session and branch in the cloud, though a user can ask to run a thread locally when it needs local files or tools. In Anthropic’s example, the user and Claude agree on tasks after exploring a website’s performance problems; Claude then starts separate threads. Users can inspect, steer or stop them. Projects is in beta, with a gradual rollout to Pro and Max plans.
A thread with a concrete change can open a PR, then watch it, push a fix when continuous integration (CI, the automated checks run against a change) fails and address review comments. The main conversation surfaces a PR’s next step, and the user can ask Claude to merge when ready. That’s an appealing division of labor: delegate the recurring follow-up without keeping every branch and check status in your head. Anthropic also warns that running several full Claude Code sessions at once uses a plan’s limits faster than running one. Parallel work has a cost even before anyone reviews it.
The Workflow Underneath the Agent
The branch is an older, useful piece of this story. A branch gives work a separate line of development; merging brings its changes into another branch. The Git book’s branching example shows a developer putting a hotfix on its own branch while other work continues. When histories diverge, Git may need to combine them, and conflicting edits can stop the merge until someone resolves them. Think of two people editing copies of the same recipe: keeping the copies separate is easy; deciding which instructions survive in the shared version takes another step.
On GitHub, a pull request gives collaborators a place to propose, discuss and merge branch changes. GitHub’s PR documentation lists required reviews and auto-merge as separate options; auto-merge can complete a PR once its configured requirements are met. Projects works on top of those existing branch and PR mechanics. The new bit is who keeps tabs on the unfinished work: a Claude thread can return to its branch when CI or a reviewer asks for a revision.
That history helps explain a small but consequential trap in the word merge. Git can combine histories, a platform can enforce configured requirements, and a reviewer can decide whether the change does what the product needs. Those are different jobs in one workflow. An agent resolving a conflict may get checks running again, but the conflict resolution is itself a code change to examine. A green check tells you that the checks passed for the version they ran on; deciding whether those checks cover the intended behavior requires knowing what the change was supposed to accomplish.
Two Routes Through the Same PR Queue
VS Code’s Agent Merge demonstration offers a close comparison. In its example, a user opens a PR, then enables Agent Merge to monitor review comments, CI failures and conflicts with the branch being merged into. The agent acts on a comment, runs a build and tests, pushes changes, resolves a conflict and pushes again. The demonstration also shows Copilot flagging a high-priority review recommendation for the user to examine. Its flow begins with a PR the user chooses to open; Projects emphasizes a coordinator that splits a larger goal into threads, some of which can open PRs themselves.
Both workflows can take on the maintenance loop after a PR exists. Their demonstrated merge controls differ: Anthropic tells the Projects user to ask Claude to merge when ready, while VS Code makes automatic merging an option the user can turn on if the PR meets its requirements. That describes the controls shown in each product demonstration, not a ranking of which produces better code. It also leaves a practical configuration question: what does a team require before its repository permits a merge?
Consider a performance fix. One thread changes an endpoint, another changes the web app, and both produce PRs. Each PR might have passing tests and comments that an agent has addressed. Someone still needs to ask whether the endpoint and web app now agree on the behavior users will encounter. That question follows from Projects’ ability to coordinate work across repositories: dividing a goal into tasks makes the pieces easier to work on, while the original goal remains the standard for judging the combined result.
Review can use automation, too. A proposed verification design suggests keeping checks and acceptance criteria outside the code-writing agent’s authority, recording which candidate ran against which checks, and sending qualified changes onward for human review. His illustrated case asks whether software reports an unreadable stored-data error and preserves the original data. A check for the error message alone could pass code that clears the data. It’s a proposal with an illustrative example, not a measured result for Projects or Agent Merge, but it supplies a sharp question for any PR: could a defective implementation pass the checks being used to approve it?
The strongest case for PR-follow-up agents is straightforward. A developer shouldn’t have to repeatedly refresh a page to discover a broken build or manually carry every review comment into another editing session. An agent that notices the failure, proposes a repair and shows its work could give reviewers more time to assess the change. The corresponding risk is that a fast repair loop also produces more revisions to inspect. If an agent changes code after a review comment, the reviewer needs to know which revision the approval covers and whether the relevant checks ran on that revision. GitHub’s configurable review and merge requirements offer a place to draw that boundary; neither product’s walkthrough demonstrates that one configuration suits every repository.
For someone evaluating either tool, the questions are concrete: Who can start threads and open PRs? What may an agent change after review begins? Which checks must run on the final revision, and who chooses those checks? Can the agent trigger a merge, or must someone explicitly request or approve it? Claude Code Projects makes it easier for work to continue while its user is elsewhere. A team’s answer to those questions determines what ready to merge means when the user returns.
More Like This
Claude Fable 5.1: Seven Prompting Changes That Actually Move Results
AI Labs tested Claude Fable 5.1 on real client work. Here's what their seven tips reveal about effort settings, hidden model switches, and token costs.
Six Claude Code Skills That Change How Agents Decide
Six Claude Code skills covering task memory, marketing flows, Karpathy's agent rules, web automation, UI variation, and pre-build validation.
Visual Plans for Claude Code Change Agent Reviews
Builder.io's Steve Sewell introduces visual-plan and visual-recap skills for Claude Code, turning AI-generated markdown walls into interactive MDX diagrams and wireframes.
Claude Sonnet 5.5's Coding Gains Put Review Costs in Focus
Anthropic says Claude Sonnet 5.5 improves coding at the same token price. For open-source maintainers, review time, retries and governance shape the real cost.
Claude Opus 5.5 Turns the AI Model Race Toward Price
Anthropic cut Claude Opus 5.5 prices, but workload cost depends on tokens, cache use and safeguards. What buyers should test before switching.
MiniCPM5 Shows the Promise and Fragility of Local AI
MiniCPM5-2B posts striking coding scores on local hardware, but benchmark gaps and fragile sampling defaults complicate claims that it rivals larger models.
GPT 5.6 Sol vs Fable 5: Early Numbers, Real Tradeoffs
GPT 5.6 Sol is half the price of Fable 5 — but is it half as good? Early benchmark comparisons, alignment regressions, and the politics reshaping who gets access.
Claude Fable 5 Prompting Habits That Actually Matter
Nate Herk distilled Anthropic engineer insights into six Claude Fable 5 prompting habits. Here's what holds up, what's wild, and what it means for how you work.