Skip to main content

What we stopped paying for after agents started writing our code

Stefan Kirsch, CTO & Co-Founder

16 min read

Our engineering team is three people, and in the last four weeks we merged 708 pull requests. Infrastructure and tooling, AI subscriptions included, cost us about €1,500 a month at list prices. Most of the operational work behind those numbers, from Kubernetes upgrades and CI runners to log triage and code review, is now done by Claude and Codex, with one of us steering. That changed which services are still worth buying.

I’m Stefan Kirsch, CTO and co-founder of amaiko. I first worked with AI during my physics studies and have been a believer ever since. I’ve been coding for 40 years, and I’m fine with never typing another semicolon. I’ve been CTO at small startups and at large enterprises, and of all the teams I’ve led, this one is the smallest and the most efficient. The other two are Markus Heigl, a senior developer with decades of experience, and Florian Schauer, our senior SRE.

Prices are October 2026 list prices in euros, converted from US dollars at 1.10 where needed. Estimates are marked as such.

Servers

We rent 16 Hetzner Cloud servers in Germany and Finland. Eleven of them form three Kubernetes clusters on Talos Linux: production, an internal cluster for our tooling, and a small development cluster. Staging is a vcluster inside production. The other five run CI. OpenTofu provisions the servers, and ArgoCD deploys everything else from git.

Inside the clusters we run a lot of things that would otherwise be separate lines on a cloud bill: Postgres via CloudNativePG, Keycloak, a NetBird VPN, Gatus for uptime checks, Renovate, backups with k8up, and our monitoring.

Florian runs all of this, with Claude doing most of the hands-on work, from Talos upgrades to figuring out why a rollout won’t go healthy. Claude also creates servers through the Hetzner API. It gets the token through its shell environment, so nobody has to paste a key into a conversation. Production changes go through git and ArgoCD, and an agent that wants to restart, delete or sync anything in production needs an explicit go from one of us.

At today’s list prices the cluster side (11 nodes, 11 load balancers, roughly 1.5 TB of block storage and some object storage) costs about €480 a month.

For comparison, I priced the same shape on AWS in Frankfurt: three EKS control planes, Graviton instances matched by vCPU and memory, 11 application load balancers, NAT gateways, gp3 volumes and an assumed 2 TB of egress a month. That comes to about €3,700 a month on demand, €3,100 with a one-year Savings Plan and €2,400 with a three-year one, or about €4,500 on demand with x86 instances. Hetzner comes out roughly eight times cheaper on demand and about five times cheaper against the three-year commitment.

Two caveats. Hetzner’s CX servers have shared vCPUs, while a Graviton vCPU is a full physical core. Our workloads run fine on shared cores, but the comparison isn’t like for like. And Hetzner raised its prices twice this year: by about 30% in April, existing servers included, and by 113 to 176% for new orders of its CPX and CCX lines in June.

Our Windows CI machine runs on a more expensive server type because the cheaper one was out of stock, and one of our recent orders came back with resource_unavailable. Hetzner’s cheapest line can’t be ordered at all at the moment. I’m not confident Hetzner could triple our capacity next week.

We’re already an OVHcloud customer, for backups and for our first experiments with self-hosted models. Neither is counted in this post’s totals. Our whole setup is code, and I think we could move to another provider in a day or so. All a provider has to give us is a server with an SSH shell. Talos and everything else comes after that.

Monitoring

Metrics go to Prometheus and Mimir, logs to Loki via fluent-bit, traces to Tempo through an OpenTelemetry collector, and Grafana sits on top. All of it runs on the internal cluster, with long-term storage in Hetzner’s object storage. This stack replaced SigNoz in June. SigNoz was buggy and tedious to work with, and we had to write our own MCP server for it. Above all, Claude works much better with Grafana. I assume that’s because Grafana is so widely used and documented that the models saw plenty of it in training.

In the last seven days the stack ingested about 5 GB of logs and 7 GB of traces per day and held about a million active metric series. The internal cluster it shares with ArgoCD, Forgejo and the VPN costs €48 a month in servers. With volumes and the storage bucket, monitoring’s share is well under €100 a month; that one is an estimate, since nothing on the bill is split by workload.

At these volumes, hosted services with 30-day retention would cost:

Per monthAs is (~1M series)Series cut to 100k
Grafana Cloud Pro~€6,200~€660
Datadog~€48,000~€5,500

A single month of Datadog at our current volume would cover about eight years of our three Kubernetes clusters. Almost all of that is metric series: Grafana Cloud charges about €5.90 per thousand active series, and Datadog bills Prometheus series as custom metrics at about €4.50 per hundred. Logs and traces are about €110 of the Grafana figure. The 100k column assumes we would cut cardinality before moving to a hosted service. We have never had a meeting about which metrics we can afford to keep.

Our agents read all of this themselves. Grafana’s MCP server gives Claude Loki, Mimir and Tempo as tools, so debugging usually starts with Claude querying logs and traces instead of one of us pasting excerpts into a chat. A scheduled job goes through the production and staging logs, groups errors by fingerprint and opens an issue for each new one. The volume numbers above came out of the same connection while I was drafting this.

Hosted vendors have MCP servers too. For us the difference is cost and where the data sits. An agent chasing one bug can easily run fifty exploratory queries, and we don’t pay per query. The logs are stored on our servers, and only query results go into the model’s prompt.

Code review

CodeRabbit’s support once described our usage as “not normal”. We were on Pro+, its largest self-serve plan at the time. On a typical morning it reviewed the first two or three PRs. Then we hit the rate limit, and the unreviewed ones piled up. We don’t merge unreviewed code, so they waited, and so did we. In practice maybe one PR in five got its review when we needed it.

CodeRabbit’s limits are counted per developer and per hour. Three people with several agents each produce far more PRs and fix rounds than that allows. The money was never the problem: three seats on today’s top public plan cost just under €200 a month.

Reviewing after the PR is open had a cost of its own. Every round of fixes meant another push, which started CI again and used up another review from the hourly allowance. CodeRabbit also doesn’t list Forgejo among its supported platforms, which would have ended things a month later anyway.

Today the review happens before a PR exists. When Claude finishes a piece of work, it commits on a local branch and runs Codex against it:

codex review --base origin/main -c model_reasoning_effort="xhigh"

Codex reviews with GPT-6.1 Sol or GPT-6 Astra, always at the highest reasoning effort and never with a smaller model. Claude works through every finding: real ones get fixed, wrong ones get a written refutation. Then Claude resumes the same Codex session and asks for a verdict on each finding:

codex exec resume "$SESSION" \
  "We addressed your findings in commit <sha>. Finding 2 is not a bug because ...
   Judge each of your findings: is it resolved? Did the fixes introduce anything new?"

This repeats until Codex accepts every fix or refutation and reports nothing new. Then the branch gets pushed and the PR opened.

We resume the session so Codex still has its earlier findings, Claude’s answers and the fixes in context. A fresh review starts from zero and tends to raise the refuted findings all over again. Rules we would otherwise repeat in every round go into AGENTS.md, which Codex reads before it starts. The reviews are included in our OpenAI subscriptions.

Git hosting and CI

GitHub outages stopped our deploys several times. When there’s a critical bug in production, refreshing githubstatus.com is not how I want to spend the next hour. Our Actions bill was approaching €900 a month, which is a lot for three developers. And the hosted runners were slow, most of all for Windows and for our macOS and iOS builds.

Self-hosted runners would have helped with cost and speed and done nothing about the outages, since every job still depends on GitHub’s control plane. In December 2025 GitHub also announced a fee of about 0.2 cents per minute for jobs on self-hosted runners, starting March 2026, and postponed it within about 48 hours. I’d rather not budget around a vendor that has already tried to charge by the minute for my own hardware.

So we moved to Forgejo, GPL-licensed free software developed under Codeberg e.V., a German non-profit. It runs as one pod on our internal cluster, so its cost is already in the server numbers. It ran next to GitHub from August with our main branch mirrored into it, and it became our primary forge on October 7, with the full history: branches, tags, releases and every pull request under its original number. That was less than a week ago as I write this, so ask me again in six months. Our product colleagues file and follow issues in Forgejo too, without anyone buying them a seat.

CI runs on our own machines: four Linux servers with 24 parallel job slots between them, and a Windows machine that needed a custom install, because Hetzner Cloud has no Windows image.

Our macOS runner is an M2 Mac mini on a desk in the office, on the office Wi-Fi. NetBird connects it to the other machines, and a firewall keeps it apart from the rest of the office network.

Together that costs about €220 a month, with the Mac written off over three years. GitHub-hosted runners cost about half a cent a minute on Linux and 5.6 cents on macOS. Our four Linux servers cost as much as about 21,600 GitHub Linux minutes a month.

macOS is where the difference is largest. Spread over three years, a new Mac mini with an M6, 16 GB of memory and 512 GB of storage (€1,066 before VAT) costs about as much as 530 GitHub macOS minutes a month, roughly 18 minutes a day. An always-on M2 Mac on AWS in Frankfurt costs about €0.95 an hour with a 24-hour minimum, about €700 a month. That buys a new Mac mini every seven weeks or so.

The AI subscriptions

Each of us has two subscriptions: Claude Max 20x at about €180 a month for development, where Claude Code does the implementing, and ChatGPT Pro at about €90 for reviews with Codex. That’s about €270 per person and about €820 a month for the team. We have no metered API spend for coding, and the limits have held up with the five to nine agents I usually run at once.

The split follows the workload, since implementing burns far more tokens than reviewing. I want a second model family checking Claude’s work, and keeping both vendors gives us somewhere to move work if one of them changes its terms. That has already happened once: at its DevDay in September, OpenAI introduced a Pro tier at about €450 and lowered the usage included in the €180 tier at the same price. Our €90 plan wasn’t affected.

Changing terms are a real risk, because the subscriptions are priced far below the API. ccusage prices local Claude Code and Codex logs at API list rates. Over the last 30 days, my own usage came to about €9,000 for Claude Code and about €680 for Codex, against €180 and €90 in subscription fees. I don’t expect prices like that to last.

Headcount

Without agents, the setup above needs people: a platform team to run it, or managed services and a smaller platform team. For comparison I priced the managed route at about €5,000 to €5,400 a month, before macOS and Windows CI. Everything our €820 covers, from writing and reviewing code to running the platform and reading the logs, would be done by people. The question is how many.

Over the last four weeks the three of us merged 708 pull requests across our two repositories, not counting bots, release PRs or automated provisioning. That’s 177 a week, or about 59 per person, with everyone working regular 8-hour days. Of those, 511 were in the application repository, 128 a week.

LinearB’s 2026 benchmarks rate anything above 228 changed lines per PR as “needs improvement”. Our median in the app repository is 548, including tests and translations into five languages (25th percentile 180, 75th percentile 1,523).

For a baseline, DX measured a median of 1.4 merged PRs per week for developers who don’t use AI, 2.3 for daily AI users and around 4 for frequent Claude users. LinearB counts anything above 2 per week as elite.

Turning PR counts into headcount is crude. To keep the number down, I only counted the app repository, compared against elite developers at 2 PRs a week, well above the 1.4 median, and ignored the size of our PRs. That puts our app repository at the merge rate of about 64 such developers. I then divided by a haircut for everything a PR count misses and added one platform engineer to run the managed services:

HaircutModeled teamPeople beyond our 3Additional salaries per month
2×~3330~€255,000
3×~2219~€160,000
5×~1411~€94,000
10×~74~€34,000

The salary column assumes about €8,500 a month per senior engineer, a normal German senior salary plus the employer’s social contributions, without recruiting, equipment or management. Our own salaries are in that range.

Coordination isn’t in the model at all. A team of 22 spends real time on meetings, on keeping everyone in sync and on staff planning. We have a retrospective every two weeks and a regular meeting about our tooling. Anything else gets a short call when it comes up.

I’d pick the 3× row. That’s a judgment call about how much real work a merged PR represents, and you’re welcome to pick another one. At 3×, the model comes out at about 22 people and about €1.9 million a year in additional salaries, against about €9,800 a year for our subscriptions. The 10× row is still a team of about 7.

Monthly costs, using the 3× row:

Per monthUsModeled team
People3~22
ServersHetzner, ~€480AWS, ~€3,700 (~€2,400 with a 3-year plan)
CIown Linux, Windows and Mac machines, ~€220GitHub-hosted Linux, ~€570 at an assumed 10% utilization, plus macOS and Windows minutes
Monitoringself-hosted, in the server billGrafana Cloud, ~€660 with trimmed cardinality
Git hosting and issuesForgejo, in the server billGitHub, ~€80 (Team) to ~€420 (Enterprise)
Code reviewCodexpeople, counted in the headcount
AI€820none
Tools and infrastructure~€1,500~€5,000 to €5,400, before macOS and Windows CI

Review and testing

Each of my agents works in its own git worktree, on its own branch, and runs its own subagents, so I always have several features and fixes going at once. I follow them through their PR descriptions. I stopped reading the code a while ago. At this volume I couldn’t, and with our checks in place I trust code reviewed by agents more than code reviewed by people. In my years as a consultant I saw plenty of bad code that one person had written and another had approved.

The app repository has these checks:

  • 22 CLAUDE.md files with guidelines for the agents, each next to the code it covers
  • 22 project skills for recurring work such as the review gate, issue triage, log triage, dependency updates and releases
  • a Codex review of every change before it becomes a PR
  • a full build before every commit: linting, type checks, tests, and checks for complexity, duplication and dead code
  • a coverage rule: every commit needs well over 80% coverage on the code it changes
  • 3,596 test files, up from 1,254 on July 1

The rules are strict. There isn’t a single linter hint on our main branch, never mind a warning, and God forbid an error.

We spent months on this setup. We had been working with agents long before July, and as models and harnesses improved, they did more and more on their own. In July we merged about 100 PRs a week; the last three weeks were 231, 200 and 194. Over the same period the median PR in the app repository grew from 428 to 548 changed lines. DX followed more than 400 companies from November 2024 to February 2026, while their use of AI tools rose by 65% on average. Their median PR throughput went up by about 8%, and DX puts the gain most organizations get from current AI coding tools at 5 to 15%.

When we started, agents weren’t good enough for more than small steps. I set the initial architecture myself and wrote a lot of the early code by hand. I think that slow start is one of the reasons the setup works as well as it does. In a new project I’d go slowly at the beginning again before letting the agents loose.

Across a dozen PRs that needed four to eight review rounds, the defect found in a round usually came from the previous round’s fix. So we added a written checklist that the agents go through before committing any fix: which other places have to change with it, what the fix makes reachable that wasn’t before, and whether the finished diff contradicts itself.

Two other rules started the same way. We once merged a PR with five test shards still running. We hadn’t made those checks required, so auto-merge went ahead. Agents now merge through a small script that re-checks that every check on the head commit is green and then merges exactly that commit. And tests written for fixes kept passing for the wrong reason: one checked that a key was present and ignored its value. Every fix now gets checked by reverting it, and exactly the tests written for it have to go red.

Reverts stayed between 0.4% and 1.1% of merged PRs every month while weekly throughput doubled. October is at 0.6% so far.

Since July we’ve had one serious production incident. On September 11 a NestJS upgrade made our global rate limiter run on WebSocket messages too. It threw on every one of them, so our chat apps stopped answering for everyone already on the new release. The Teams bot uses HTTP and kept replying, and so did our scheduled tasks. With those replies still coming in, nothing looked wrong, and it took us almost two hours to notice. The fix was in production about 40 minutes after that.

Staging had been broken the same way for two days. We were sloppy and didn’t test there before the rollout. Things had been going well for a while, and we had gotten comfortable. Claude now has a CLI that drives our web app, so it can test staging on its own.

What’s next

We want to run open models on our own hardware for small, well-bounded fixes, end to end without a human in the loop. We’re experimenting with that now and have our first test results.

How we measured

  • Throughput: git log --first-parent on the main branches of our two repositories, counting squash-merged PRs by the three of us between September 14 and October 11, 2026. Bots, release PRs and automated provisioning are excluded. PR size is the number of changed lines from --shortstat.
  • Monitoring volume: seven-day averages of the ingest and active-series metrics in our own Mimir.
  • Prices: list prices as of October 11, 2026, excluding VAT unless noted, from the vendors’ pricing pages linked above and the AWS price list for eu-central-1. US dollar amounts are converted to euros at 1.10.
  • Benchmarks: LinearB’s 2026 Software Engineering Benchmarks Report, DX’s AI impact reports for Q4 2025 and Q1 2026, and DX’s longitudinal analysis of November 2024 to February 2026.
  • AI usage at API prices: ccusage daily over the Claude Code and Codex logs on my machine, September 11 to October 10, 2026.
  • This post: Claude wrote it from my notes and answers. Codex reviewed it in twelve rounds, for facts and for sentences that sounded like AI. The numbers come from the sources above. Same process as our code.

We build amaiko, an AI colleague for Microsoft Teams.

We build amaiko, an AI colleague for Microsoft Teams.