9 Tunnels

Notes on building, leading, and the journey between milestones. By Angelo Rodriguez.

I was fixing a bug while waiting to board my flight

Angelo standing in a Delta boarding area in a grey zip jacket, reading his phone with a carry-on bag at his feet, other passengers and a gate counter in the background.

We were in the group about to board. My wife looked over, took this picture, and asked what I was doing.

“Fixing a bug for one of our clients.”

She gave me the look. Not a suspicious look, more of a “you are on vacation and also that is a phone” look. And honestly, it felt weird saying it out loud. Twenty-five years in engineering leadership and some part of my brain still insists that fixing a bug requires a desk, two monitors, and a long uninterrupted stretch of deep focus. The kind of session where you load the whole problem into your head and stay in the zone until it gives.

But that is the world we live in now. I was standing at the gate waiting to board, holding the same device I use to check the weather, and real work was getting done on a client codebase.

What I was actually doing

Here is the part that matters. I was not typing code on a six-inch screen. Almost nobody wants to work that way, and I am certainly not going to start.

I was reading a failure, describing the correct behavior, and reviewing what came back. The work happened at our home base, literally on Mac Studios set up to run our development and test harnesses. The phone was a control surface, not an editor. It is closer to defining and approving a change from your phone than to writing one.

Which means the thing that made this possible was not the phone at all. It was everything we had already built behind it.

The people who got here before me

Boris Cherny, who created Claude Code, has been writing software from his phone for a while now. When I first saw that, my reaction was the honest engineer reaction: cute demo, not a workflow.

And I understand that reaction, because the question underneath it is the right question. How could you possibly believe code written this way? Where is the design discussion, the review, the person who can explain the change six months from now when it breaks?

I did not believe it either. Not from a demo, not from a thread, not from anyone telling me it worked. I believed it when we built our own development and test harnesses and ran real client work through them ourselves. That is the only thing that moved me, and I suspect it is the only thing that will move you.

So do not take my word for it. Look at what companies that live or die by their code are actually doing.

Anthropic published its own internal numbers in June. More than 80% of the code merged into its production codebase in May 2026 was authored by Claude, up from low single digits before Claude Code launched in research preview in February 2025, and in the second quarter of 2026 the typical Anthropic engineer was merging eight times as much code per day as in 2024. This is a company last valued at roughly $965 billion, close to a trillion dollars, which has filed to go public, and it builds the product that carries all of that value this way.

And if you want a number closer to 100%, it exists. Boris Cherny has said that across a thirty day stretch, 100% of his contributions to Claude Code were written by Claude Code, and in June he said he had not written a line of code by hand in about eight months. Anthropic has described something in the neighborhood of 90% for Claude Code’s own codebase, and has put its company-wide figure at 70 to 90 percent, although that one counts code written with Claude Code assistance rather than authored outright. These are descriptions offered in interviews, not audited metrics, and the definitions shift between them.

Still, the man writing software on his phone is not doing a party trick. He is running the tool he built, at the limit of what it can do, on the product his company sells.

The figures reported elsewhere are lower and older. Satya Nadella said in April 2025 that maybe 20 or 30 percent of the code in Microsoft’s repositories was written by software. Sundar Pichai put Google at well over 30% of new code around the same time, up from more than a quarter the previous October.

Here I have to stop myself, because the tempting move is to line all of these up and call it a trend line. They do not line up. Google is counting new code where an engineer accepted an AI suggestion. Anthropic is counting lines merged and attributed to Claude. Cherny is counting his own contributions to one product. Those are different denominators measuring different things, and stacking them on one scale would be exactly the sloppy reasoning this article is arguing against.

What they support together is narrower, and still worth your attention: at several serious engineering organizations, a large and growing share of production code now originates from a model rather than from a person typing. Whether that code is any good is a separate question. It is the question the rest of this piece is about.

Coding is dead

Earlier this year, my friend Arup, founder and CEO of BlastAsia and co-founder of Xamun, told me flatly that coding is dead.

There was no back and forth. I just was not ready to accept it.

What I could not get past was accountability. How could you possibly validate code that no engineer is accountable for?

The answer turned out to be simpler than my objection. The engineer is still accountable. What they are accountable for moved. They own the definition of the requirements and whether it is actually right. They own whether the result is verified against that definition and properly tested. They own whether it is secure, whether it performs under real load, whether it handles data the way the law and the customer require, whether it fails safely and can be rolled back. They own whether the system is observable enough to tell you what it is truly doing in production rather than what you hoped it would do. And they own the last call, which is evaluating the output and being able to prove it is correct rather than assert it.

Not one of those got easier. Every one of them got more important. What went away was the typing, and the typing was never the part I was trusting in the first place.

That is the part Arup is right about, and it is bigger than most engineers want to admit. Typing code is no longer the bottleneck. It has not been for a while. What is not dead, and what I would argue just became the whole job, is deciding what should exist and proving that what got built actually does it.

The harness

At Qandaba we have spent this year building a development harness. It runs on Mac Studios at our home base and it holds ten or more coding sessions at the same time.

We have not really tried to find its limit. There is a lot more headroom in the machine than we are using. What actually limits us is how fast we can define work properly and evaluate what comes back, and frankly how big we are willing to imagine.

That number sounds like a flex. It is not, and it is not the interesting part either. The interesting part is what those sessions need from us to be useful, because ten agents writing code with no definition and no verification is not ten engineers. It is ten times the review backlog and a very expensive mess.

So the harness is built around the two things that are actually scarce:

Definition. What are we building, what does correct look like, what are the acceptance criteria, what are the edge cases, what must not change. Written down before anything runs.

This is not waterfall. Definition can be iterative or phased, and ours usually is. It just has to be written down and kept current.

Verification. Does the thing that came back match the definition. Not “does it look right.” Does it demonstrably do what we said, checked by specialized passes that each know one thing well, with humans approving risk at defined checkpoints.

Generation sits in the middle and it is the cheap part now. Everything expensive moved to the ends. This is the same argument I made in Control of your codebase is an illusion, except I am no longer describing it as a direction. We are running it.

Why verification stopped being optional

I spent most of my career as an engineering leader, and my job description, stripped down, was this: I had to trust that every one of my engineers wrote code correctly.

I would like to tell you I never trusted that blindly. That would not be true, and I paid for it once.

In one of my roles, a critical piece of code was reported as working correctly. That report came from the top engineers on the team and from their engineering leaders, who were engineers themselves and still deep in the code. These were the most trusted technical advisors in the company, and I want to be precise about what that means. It was not that I trusted them. Everyone trusted them. If you had walked the building asking whose word you would take without checking, you would have heard the same handful of names from every desk, including mine. So nobody checked. I certainly did not.

It was not working correctly, and the price for that was mine to pay.

The lesson had nothing to do with their carelessness or my naivety. It was narrower and more useful than either. Trust is not a verification method. It never was, no matter how senior the people are or how right they have been before. Everything I built after that, I built with verification and validation underneath it. Code review, test coverage, acceptance criteria, traceability, staging, canaries. The discipline existed because humans are fallible in known ways, and because a human who writes a line of code carries intent and accountability along with it. You can ask them why. They remember. They care whether it works next quarter.

AI changes that in three ways, and all three point the same direction.

There is no author to ask. The model does not carry intent from one line to the next the way a person does. It carries pattern.

The output is confident in a way that defeats skimming. Human bugs usually look like mistakes. AI bugs frequently look like well-formed, plausible, idiomatic code that is simply wrong about your business.

And the volume is the real problem. When one engineer can produce what used to take a team, the review capacity that was already your constraint becomes the thing that snaps first.

So here is where I have landed. Definition and verification of AI-generated code are mandatory now.

They were always required. That is not the new part. The new part is that AI removes every excuse for skipping them and punishes you faster when you do. Substitute trust for verification and you will ship code nobody read. Skip the definition, let the model infer what you meant, and you will get slop, at a volume no review process you currently have can absorb.

Verification and validation used to be the thing you asked for in every planning cycle and never fully got funded. Now it is load-bearing.

The good news is that a great deal of verification work is itself automatable, so checks we could never afford to run deeply, on every change, we can run. The depth of checking that used to be a budget argument is now an engineering decision. Most teams have not noticed, because they spent the entire AI budget on the generation half.

What we see in our own work

The most useful evidence I have is not mine. One of our architects, Ryan, has been running client work through this for months, and his assessment carries more weight with me than my own because he is in it every day. With Qandaba’s combination of models, skills, and process, he finds the output near perfect, provided he defined it correctly. What he keeps coming back to is how accurate the harness is. It still surprises him, and he is the one watching it work.

Which is a smaller claim than it sounds, because that condition is carrying most of the weight.

This is first-party observation from a small consultancy. No control group, no independent reviewer, no defect rate per thousand lines to hand you. Weigh it as what it is.

When we did go back and trace the defects that made it through, most of them were not model failures. They were requirement gaps. Something we did not say. An edge case nobody wrote down. An assumption that lived in someone’s head and never made it into the definition. Often it was not that we asked for the wrong thing. It was that we never defined it at all, so the model inferred what we wanted, and inferred it incorrectly.

That is a humbling result, and it is also the most useful result we have gotten all year, because it tells you precisely where to spend your effort. If your AI output is disappointing, the first place to look is not the model, the tool, or the prompt. It is the clarity of what you asked for.

For the skeptics

I know there are engineers reading this who think all of it is overstated. I have real respect for that position. I held a version of it, and there is a lot of noise in this space that deserves the skepticism.

Three things I would offer.

First, this is the worst it is ever going to be. Everything you are judging right now is a floor, not a ceiling. Whatever the tools cannot do today is a snapshot with a very short shelf life. Capability, price, and regulation could all move in unhelpful directions, and I will grant that. But for two years now the models have only gotten better.

Second, if you evaluated this a year ago and put it down, your data is stale. Even three months ago is stale. For me the change was datable. It came with Claude Sonnet 4.6 and GPT-5.3-Codex, when things that had been demos became things I would put in front of a client. If your opinion was formed before that, it was formed on different evidence.

Third, and this is the one I would happily argue with you about over a drink, look at what the job has always actually been. A software engineer’s role covers the entire development lifecycle. Understanding the problem. Defining the requirements. Designing the system. Building it. Verifying it. Operating it and living with it afterward. Writing the code was always one phase among several, and we let it become the one we identified with because it was the part that felt like the craft.

AI did not delete any of those phases. It shifted where your effort goes. And it shifted it squarely onto the three that get under-budgeted on almost every project I have ever seen: requirements definition, software design, and verification.

Think about your last project honestly. How much time went into defining what you were actually building and proving that it worked, compared to how much went into typing, and then into fixing what the typing got wrong? On most teams I have led or advised, it is not close. Those three phases got whatever was left over after the schedule was already spent.

If that ratio bothers you, and it should, then this is not an argument against AI. It is the argument for it. The work is finally being pushed toward the parts that were always underfunded, and the parts we were quietly shortchanging are now the parts that decide whether the output is any good.

Where we landed

We drank the kool-aid. I will say that plainly, because the skeptics are going to say it anyway.

The difference is we have the receipts. Client work shipping through the method. A harness running ten or more sessions at once. A defect trail that points at our requirements rather than at the model. And a founder waiting to board his flight, closing out a real bug on a phone, which three years ago would have been an absurd sentence.

Two invitations to close.

If you want help building this in your own organization, the definition layer, the verification layer, the harness that holds it together, that is what my team at Qandaba does. Book a free assessment or reach us at info@qandaba.com.

And if you think I am wrong, I would genuinely like to hear it. Not the version where we talk past each other about whether AI is good. The specific version. Tell me where the method breaks, what you tried, what happened. I have changed my mind on this once already, and I would rather do it again than spend another year being wrong in public.