AI Engineering 6 min read Oct 2, 2026

Fixing the PR Bottleneck in OroCommerce: What I Learned from Matt Pocock's Talk

AI agents now write code faster than humans can review it.

Written by Mukesh · Reviewed Oct 2, 2026

Post

AI agents now write code faster than humans can review it. In Paris, Matt Pocock argued that AI coding agents have made the pull-request review bottleneck worse by flooding repos with more code. In his words, a "software factory" without brakes becomes a "slop cannon." daily

His fix is not faster humans. It is a system that raises code quality before a human sees it. Here is the talk, and how an OroCommerce team can use it.

1. The core idea: three brakes

He proposes a three-layer pipeline: cheap automated checks, automated review by a separate sub-agent loaded with team-specific coding standards, and human review reserved for high-risk changes. Each layer makes the one above it cheaper. daily

flowchart TB
    A[Agent writes code] --> B[Layer 1: Deterministic checks<br/>lint, static analysis, tests]
    B --> C[Layer 2: Reviewer sub-agent<br/>reads CODING_STANDARDS.md and commits fixes]
    C --> D[Layer 3: Human review<br/>focused on one-way-door changes]
    D --> E[Merge]
    D -. Retro skill turns every review comment<br/>into a new check or standard .-> B
    D -. .-> C

2. Green CI can lie

Checks are cheap because they cost CPU, not tokens. But Pocock showed three ways agent-written tests pass while proving nothing. Here is each one in PHP terms:

FailureOro/PHP version
Tautological testThe service has MAX_LENGTH = 280 and the test asserts assertSame(280, Service::MAX_LENGTH). It re-states the code.
Structure-sensitive testA test reads services.yml or a template as text and asserts that line X appears before line Y. It breaks on refactors and proves no behavior.
Test that can't failEvery collaborator is mocked, including EntityManager and the D365 client, so the assertions only check the mocks.

He noted the agent isn't cheating. It follows instructions and writes tests too tied to structure. Automated review and human review are the lie detectors for the automated checks.

3. Design your way out: deep modules

Pocock's design remedy comes from John Ousterhout's A Philosophy of Software Design: deep modules hide complex behavior behind simple interfaces, which produces fewer structure-sensitive tests because tests sit at the interface. Your job is to force the agent to use the small interface instead of reaching into the internals. His codebase-design skill gives shared vocabulary for this: locality, leverage and seams. biggo

In Oro this looks like:

  • Shallow: twelve helper classes for D365 sync, with controllers and listeners calling each one directly.
  • Deep: one ErpGateway with two or three methods. It hides auth, retries, field mapping and error translation, and your tests target only that interface.

4. Implement, then review: standards belong in the reviewer

Most people get this wrong. The implementer agent's context is already full of exploring, editing and debugging. Adding coding standards makes it worse. Review is the underloaded step.

"...implement, you make it work, and then code review, you actually make it good..." — Matt Pocock

So:

  • Keep standards out of AGENTS.md and out of the implementer's prompt.
  • Put them in a CODING_STANDARDS.md that only the code-review sub-agent reads.
  • Have the reviewer commit fixes instead of leaving comments. Comments create more work for the human, so reserve them for real open questions.

He also advises building your own reviewer rather than relying only on generic tools. A generic reviewer is too broad to know your stack, and an over-specific one stops being reusable.

diffcommits fixesreal questions onlyImplementer agentcontext: explore, edit,debugReviewer sub-agentfresh context +CODING_STANDARDS.mdClean branchHuman

5. Make human review fast: doors and blast radius

Not every PR deserves the same attention. Classify each PR as a one-way door (hard to reverse, such as expensive migrations, data loss, or an email blast to 60,000 people) or a two-way door (easily revertible), and summarize the blast radius and merge danger at the top. daily

Oro one-way doors to review hard:

  • Schema migrations on large tables.
  • Data fixtures that change prices or customer data.
  • ERP/D365 write-back syncs.
  • Price list or search reindex changes.
  • Anything that triggers mass notifications.

A localized two-way door can get a light review. For comprehension, his PR skill favors pseudocode and Mermaid/UML diagrams over walls of text, and he credits the "Show Me" skill from the Human Layer repo as an influence. He admitted the PR skill is still in progress.

6. The retro loop: never write the same comment twice

A human review is also a review of the system that produced the code. The new Retro skill takes a session, a PR with its session, or a week of PRs and reviews. It then suggests new automated checks and standards. It also looks at navigation pointers, token-efficiency of tools, and bloated skills or steering files. Over time, each review makes the next one cheaper.

7. How to adopt this in OroCommerce (5.1 or 6.1)

Step 1: Install and set up.

npx skills add mattpocock/skills
/setup-matt-pocock-skills

Check the repo for the current state of code-review, retro and pr, because the talk announced them just as v1.3 shipped.

Step 2: Write docs/CODING_STANDARDS.md. Make every rule checkable, and pin your Oro version. Example rules to start with:

  • Custom code lives in your own bundle, and core classes are extended, never edited.
  • Every schema change has an Oro migration.
  • Controllers stay thin and business logic lives in services.
  • Every action has ACL, using the attribute or annotation style your version supports.
  • Datagrids, layouts and entity config follow Oro's extension points, with no core template overrides when an extension point exists.
  • No tests that re-assert constants, read source files, or mock everything.
  • External systems are accessed only through the gateway class.

Step 3: Strengthen Layer 1. Run php-cs-fixer, PHPStan or Psalm, and PHPUnit (including functional tests with fixtures) before any agent review.

Step 4: Build an oro-review flow. Run /code-review main with CODING_STANDARDS.md as the Standards axis and the ticket as the Spec axis. Let the reviewer commit fixes.

Step 5: Add a PR template. Put these at the top of every PR:

markdown
## Door type: One-way / Two-way
## Blast radius: <what can break, who is affected>
## Merge danger: Low / Medium / High
## Why: <2-3 lines>
## What changed: <pseudocode or a Mermaid diagram>
## How to test: <steps>

Step 6: Run Retro weekly. Feed it last week's PRs and review comments, then promote repeated comments into checks or standards.

8. My takeaways

  1. Don't try to one-shot good code. Implement first, then review.
  2. Put standards in the reviewer, not AGENTS.md.
  3. Add more deterministic checks, since they cost almost nothing.
  4. Spend human attention only on one-way doors.
  5. Turn every repeated review comment into automation.

Source: Matt Pocock, "Fixing the PR Bottleneck," AI Engineer Paris 2026; skills at aihero.dev/skills and github.com/mattpocock/skills.

About the author

Mukesh is the developer behind InfoMukesh, writing practical notes from hands-on work with PHP, Laravel, e-commerce platforms, AI, and web applications.

Related reading

Stop Leaking Secrets: How to Catch Security Flaws Before Your Pull Request

Finding security vulnerabilities and exposed API keys during pull request reviews or CI/CD pipelines causes unnecessary rework and security risks. This guide explains how to implement "Shift-Left Security" by running static application security testing (DevSkim) alongside dedicated secret scanning (Gitleaks) directly on your local workstation using Git pre-commit hooks. By catching insecure coding patterns and credentials the moment you commit, you eliminate awkward PR reviews and prevent leaks before they ever enter Git history.

From Lost Paper Receipts to One-Tap Profit: How I Built a Custom App for a Cab Driver

Independent taxi drivers juggling multiple aggregators like Uber, Ola, Rapido, and offline private bookings face daily accounting chaos, often relying on paper slips that get misplaced. This case study documents the end-to-end development of SD Travels (Driver Portal)—a lightweight, driver-centric mobile app engineered to deliver instant net-profit visibility, quick single-tap entries for fares, CNG refills, and maintenance, and reliable offline-first local data persistence. Within its first week of real-world use, the app completely eliminated month-end bookkeeping guesswork and provided effortless, real-time daily profit tracking

The AI Code Review Bottleneck Nobody Warned Us About

AI coding tools have made writing an app faster than ever, but that speed didn't remove the bottleneck — it just moved it downstream to code review. This post breaks down why AI-generated code tends to skip architecture and testing by default, how that shows up as cascading production failures, and why the actual engineering work now happens at review time instead of at the keyboard.