A dark title card with a jagged score line and a question mark on the left, and five checkup rows on the right, four passing and one flagged

AI Writes Most of My Code. I Treat Codebase Health Like a Checkup, Not a Score.

Aug 27, 2026

On this page

I’m the sole developer of a note-taking app, and AI writes most of its code. As the codebase grew, reviewing every line stopped being realistic. I needed another way to tell when the codebase was getting harder to maintain and extend.

A codebase is healthy if the cost of confidence per change doesn’t rise as the codebase grows.

Cost of confidence is what it takes, after adding or changing one feature, to conclude that it actually works. This cost is an axis you can separate from code size. The codebase could grow tenfold while the cost of verifying one change stays the same.

So the question became: is there a number that goes up or down as the cost of confidence rises or falls, in other words, as the codebase gets healthier or sicker?

I never found one. I’m not claiming no such metric exists. But in my codebase, an Electron desktop markdown app written in Svelte and TypeScript, every candidate I tried had a problem I couldn’t get around.

The search for a single metric

Verification time per commit

The first candidate was the time spent verifying each commit. It’s the most intuitive option: record the time, watch the trend, and read a rise as decay.

But verification time tracks the difficulty of whatever that commit did. A hard change takes longer to verify than an easy one, and that says nothing about the health of the codebase.

Full E2E suite duration

The second candidate was the runtime of the full E2E suite. E2E tests launch the real app and click and type the way a user would, and they’re the longest-running verification in my repository.

This number naturally grows as features and tests are added. The measurement environment turned out to matter too. My valid measurements from August 2026 ranged from 29.6 to 39.9 minutes, a spread of 10.3 minutes across runs. Once my computer went to sleep mid-run and a single test was recorded at 17.8 minutes. Twice I started a measurement while another app was using over 90% of the CPU.

When this number rises, you can’t tell whether the codebase decayed, features grew normally, or the machine was busy.

Failure-cause breakdown

The third candidate was the composition of test failures by cause: product defects, stale tests that hadn’t kept up with product changes, and tests that flake from run to run. Track the share of product defects, and read a rising share as decay.

But to classify one failure as a product defect, you must have already investigated it. Classifying 28 failures took four full suite runs and dozens of individual investigations. As a solo developer, that isn’t a metric I can afford to maintain.

What the three had in common

All three were proxies for the same total, cost, expressed as a single continuous number, and each proxy was contaminated: by difficulty, by normal growth, by the measurement environment, by the cost of measuring.

The cost of confidence is a real cost, but it turned out not to be a value you can measure the way you measure blood pressure.

Check for the causes instead

So I changed the question.

Instead of measuring how much the cost is, check whether any cause of rising cost is present right now.

This is how a health checkup works. A checkup doesn’t produce one overall health score; it looks for a set of known risk signals. The list of signals is never guaranteed to be complete, but a state with zero known problems and a state with several can still be told apart, and that distinction is enough to decide it’s time to take care of the codebase.

Looking back over the past four months through this lens, I could pick out the causes that had actually raised my costs. And each checkup item answers yes or no, is independent of code size, and has a fixed cost to check.

A full medical checkup can’t happen daily, but a cheap codebase check can. You can’t get a brain MRI every day; you can check your blood sugar every night.

Translated to a codebase, checking blood sugar before bed becomes running a small health check on every commit. And Git hooks do the checking, not a person.

There are five checkup items. The first three are the blood-sugar checks that run on every commit. The fourth is the full checkup you schedule. The fifth is more like tracking body composition: not a value at one moment, but what has accumulated.

Checkup item 1: Does a new test produce a false green?

Some tests I added along with a change looked like they verified it, but didn’t actually verify the change. So I made this mandatory: deliberately break the behavior the test claims to cover and watch the test fail, a manual form of mutation testing.

Any commit that adds or modifies a test must describe that mutation check in its commit message. If the message doesn’t have it, a Git hook rejects the commit, so a test change that skipped the check bounces back until the check is done.

Checkup item 2: Do the docs point at code that no longer exists?

My codebase runs to 160,000 lines, and it includes documentation that helps the AI get oriented quickly and work in a consistent way. It saves tokens and time. But as code keeps changing, the docs fall behind. I had rules requiring doc updates in the harness, the set of rules and verification tools the AI works inside, and they still slipped sometimes. And stale docs sometimes steer the AI the wrong way: it reads them and acts on them.

So I built an automatic check that verifies that every file a document points at actually exists. A Git hook runs it on every commit alongside lint and type checking, and a dead reference blocks the commit.

Checkup item 3: Did a new test land in a more expensive layer than necessary?

When I first started developing with AI, I was struck by how cheap writing code had become, and I put every test at the E2E layer. Testing in the environment closest to real use seemed most accurate. Four months later the full suite took 40 minutes.

The same behavior costs different amounts to check depending on the layer. A unit test that runs just a function or a component is fast. An E2E test that boots the real app and drives the screen is comparatively slow. You need E2E when what you’re checking crosses the operating system or a process boundary; if you’re checking a function’s inputs and outputs, a unit test is enough.

A commit that adds a new test file must state which layer it went in and why. Without that line, a Git hook blocks the commit. What’s enforced is not that the judgment is correct but that the judgment happens. The verdict can be wrong, but tests no longer pile up at the E2E layer without anyone having asked the question.

Checkup item 4: Are any tests failing right now?

In health terms, the first three items are like checking your waist, your blood sugar, and your weight: quick enough to do all the time, so Git hooks enforce them.

But the way you set aside a day each year for a full checkup, you also have to run the entire suite, unit tests and E2E together, and see whether anything is red.

All green is best, but even short of that you can set a baseline. If 3 of 300 tests are red and you know exactly which 3, developing on top of that is a different activity from developing without knowing. With a baseline, you can tell whether the codebase got worse, got better, or held steady.

A full run takes tens of minutes, so it can’t run on every commit. If you’re maintaining a working system with CI/CD, running everything every time is right. Building a product alone, I can’t. (Before a release, of course, the full suite runs.)

So I stopped tracking the timing myself: a Git hook counts the churn in the source, insertions plus deletions, since the last full run, and past 20,000 lines it tells me it’s time to think about running one.

The 20,000 came from measurement. There was a stretch of four months without a single full run, in which 111,393 lines changed and 28 red tests piled up, roughly one per 4,000 lines. So 20,000 lines is about five reds’ worth, and that felt like the right point to stop and look. The value is provisional and will be corrected by future measurements.

When the warning fires I can run the suite or defer it. Deferring requires writing the reason in the commit message, or the commit doesn’t pass. A deferral resets the counter: that commit becomes the new starting point for the line count, and the count builds again from zero, because an alarm that keeps ringing gets switched off. But the fact and the reason remain in the commit message, so one git log command counts how many times I deferred and what my excuses were.

Checkup item 5: Is the sum getting worse while every change is justified?

The first four items all examine whether the verification system itself is sound: tests are real, docs aren’t rotten, tests sit in the right layer, nothing is red. But there’s a way for code to get worse while passing all of it. The code handling Korean text input once grew to ten state flags over eight months. Five bug fixes put them there, every fix small in scope, none doing refactoring I hadn’t asked for. Open the commits one at a time and there is nothing to point at. A check that looks at one commit can hardly catch this, because each change was sound and it’s the combined structure that got complicated.

Mechanical counting wasn’t enough either. Static analysis can surface candidate mutable fields, but whether a field is internal state or a reference the object holds so it can hand work to another object only becomes clear by reading the code.

So this is the one item where the checker is an AI. Every 30 commits that touch product code, a notification fires, and the AI sweeps the codebase with a checklist and records the results in a commit message. The checklist looks for the kind of buildup a single diff rarely shows: growing mutable state, second copies of the same logic, entries piling up in exception lists. Unlike the other four, this item doesn’t block commits. Commits already made can’t be blocked, and the point here is not gating but keeping count of what’s accumulating.

Summary

To examine something every time, the judgment has to be mechanical and the cost has to be low. Anything slow gets bypassed.

What a script can verify directly is forced to run on every commit. What takes too long, like the full suite, gets a threshold and a forced judgment. And what can only be judged by reading the code is left to a periodic review.

A mechanism that demands records in commit messages can’t stop false records. Nor do these five items cover every cause of declining codebase health; the cost of finding the code you need, for instance, is not checked here.

What the current setup guarantees goes exactly this far: the fixed questions don’t get skipped, and each decision, to run or to defer, leaves a record. Causes not on the list will have to be added as new checkup items as I run into them.