October 2, 2026
AI in the Client’s Hands
The most dangerous code isn’t the code that doesn’t compile. It’s the code that compiles perfectly but has bugs.
I want to share a story from one of our Adobe Commerce Cloud projects. A client’s QA engineer began using AI to fix bugs, and the first results were excellent. Then a seemingly small change taught all of us something about the limits of clean code, code review and the context we give these tools.
A QA Engineer Starts Using Codex
At one of our clients, a QA engineer was responsible for manual testing. He is bright and motivated. He had started learning PHP and JavaScript with the goal of moving into test automation, and he became interested in Magento as well. I do not know exactly what prompted him to try fixing bugs himself, but he took the initiative.
He installed OpenAI Codex and started working on bugs he had found. His first pull request addressed how PAID / NOTPAID status was determined for purchase order (PO) orders. The second fixed an issue introduced by an update to a third-party product-label module: a label sometimes covered the price, making it look as though the price had disappeared.
Both fixes worked. The code was clean, followed Magento standards and met the project’s requirements. He achieved this without deep prior knowledge of Magento or a complete understanding of the project architecture. Our role was to review his pull requests. I am not advertising Codex, but it did a brilliant job generating the code.

Then Came a “Simple” Feature
Next, the client asked him to add a feature that sounded simple: a button in the back office, under order management, to create an additional order status for a reorder.
The generated code again looked nearly flawless, and the new feature worked. We conducted our usual code review, and the QA engineer manually tested the order flow. We went live. Everything seemed fine.
A few hours later, we discovered that the change had broken the standard reorder functionality. The new feature still worked perfectly; the existing one did not. We rolled back the change and then implemented the feature with the standard reorder flow in mind.
.png)
The Missing Context
That is the part of the story that stays with me. The code was clean. A human had reviewed it. Manual testing had passed. Yet an existing function broke in production.
An AI agent works with the task, files, patterns and instructions it is given. It may miss a dependency that no one identifies or tests. We had checked whether the new button worked, but we missed its effect on the standard reorder flow. The problem was not that the code looked bad. The problem was that our review and testing did not cover enough of the system.
This is where an experienced technology partner can provide value. We know the surrounding system, the history of its customizations and the business consequences when an established workflow stops working. We also have to put that knowledge to work in review and testing.

What We Learned
- Review the impact, not just the code. Ask what a change touches as well as what it adds. Clean code can still change behavior elsewhere.
- Test related workflows. A change to order statuses or reorders should include checks of the existing order flows, however small the request sounds.
- Give the agent context. Explain the relevant architecture, constraints and related behavior. Better instructions help, though they do not replace testing.
- Assign a release owner. A person must be responsible for the decision to ship and for what happens in production, regardless of how the code was written.
- Encourage people who want to grow. This QA engineer showed initiative and learned quickly. The gap was in our shared process, not in his ambition.
- Measure time to stability. Count the time spent reviewing, testing, diagnosing and reworking a change, not only the time saved writing it.
What Changes for Developers and Agencies?

Should developers be afraid that AI will take their jobs? The concern is understandable. I see our profession changing again. We have moved from webmasters to backend developers, from backend developers to full-stack engineers, and now toward more work directing and checking code. We have always had to keep learning.
AI can be a highly capable assistant. It is like replacing a manual drill with an electric one: the tool helps you work faster, but you still have to decide where to drill the hole.
Should agencies worry that clients will do more development themselves? Yes. That risk is real. We need to rethink our relationships with clients and the services we provide. When generating a feature becomes easier, business consulting, domain expertise and technology judgment matter more. We can help clients use these tools while looking at the consequences across the whole system.
Clients are already using AI to write code. Our job is to help them use it well. The AI assistant is already on the client’s QA engineer’s workstation, and it has already opened its first pull request.


