As a non-technical PM, when I first started building an application with Replit last year, I didn’t know what I was doing. I was amazed at the output of my vibe-coded minor complexity application. At the same time, I was frustrated with the output. The ‘minor complexity application’ had turned into a huge, monolithic, complex application that the agent kept breaking.
It would make changes that weren’t needed, touch code that was already working, and break flows that I had already tested and passed. I was perplexed by the amount of tokens I was burning on changes I didn’t ask for. This one time it broke the login experience by refusing to accept the generated two-factor token, and I couldn’t get past the login screen. It also made up the customer testimonials on the home page without me asking.
I got better at prompting specifically: only make changes to this feature; do not make any changes to this, this, and that. Nevertheless, the application became huge, and then I learned to modularize and componentize it so I could make changes to smaller things.
Eventually I created git branches to make changes. This helped create a separation between working software and recent updates, and to do a diff of the changes to see exactly which file had changed to review and discard the changes where necessary.
That’s the thing though. AI coding tools hand non-technical builders, like myself, the smallest part of the job and make it look easy and hide the rest of the work. This is a great way to learn to build. But building to ship takes up time for activities around the code - like rearchitecting the application, prompting to narrow the scope, regression testing, and GitHub (this has been a learning curve for me). Some of these you can learn quickly; others may take some time.
An average engineer only spends 12-18% of their time actually coding. They spend the rest of their time in meetings and communication (12%), coordination, debugging (9%), and code reviews (5%).
One report suggested that teams with high AI adoption completed 21% more tasks and merged 98% more pull requests. But PR review time went up 91%, PRs got 154% bigger, and bugs rose 9%.
The rhetoric around AI coding is that AI can code better and non-technical PMs don’t need to worry. This advice is fine when you’re learning to build, but incomplete and dangerous when building to ship. A non-technical PM may not be able to read the code and confirm AI is building what is asked. Deploying incomplete or buggy code to production can have real ramifications in terms of user experience, business value, and customer satisfaction.
The tasks an engineer does, besides coding, are far more technical than what we ascribe to a PM. We can expect PMs to code a requirement 50-60% of the way, but then we need engineers to do thorough peer review, and I wonder how much of the efficiency gains will come out of it for the team overall.
Complex legacy architectures, coding culture (heck, culture in general!), how old the codebase is, how technical the product manager is (or were they rebranded from product owner or business analyst) all factor into how efficiently we can use AI to code. So before we start claiming efficiencies from AI in software development, and look at teams like Anthropic where more than 80% of the code is written by Claude, shouldn’t we look at which type of building we are asking PMs to do, and whether this model of AI software development works in our context?

