Hi, I'm Serhii, and today we're going to talk about how we can half your CI pipeline with some practical optimization strategies. In our organization, at around 50 merge requests a day and 400 pipelines run, including merge trains, one minute of overhead in your CI pipeline costs nearly six to seven hours of engineering time every single day. Now, you can think about your organization and scale this accordingly. This is a talk about how we cut our merge train pipeline from about an hour to 22 minutes. It's a 64% reduction using a structured ROI-prioritized approach.
Five areas in the order we use them, with actual numbers for our 400-person legal tech SaaS, where half of the people are developers. So this was our starting line. Merge train pipeline, about an hour. CI health, the success rate without retries is around 82%, $16,000 a month on AWS compute. And the dumb complaint on our developer experience survey was simply pipeline speed. So the number that mattered most wasn't the duration, it was the failure rate. Nearly one in five merge attempts failed.
Retries became muscle memory. People stopped trusting the system. Once developers stopped trusting CI, every other engineering metrics started sliding, too. Now, most teams approach CI optimization usually backwards. They start with easy quick wins because the quick wins can ship in an afternoon. I actually started different. So we started with a table, every potential optimization scored by expected time saving, effort, and risk. So instance migration was the largest single lever.
12 minutes, so it went first. Observability tweaks at one and a half minutes went later, even though they were the easiest. So we focus on the high impact first, quick wins in parallel, and fine-tuning last. Pick the right battles before you fight them. The biggest single win came from the compute instance migration for our feature, or end-to-end to test, how you call them, which are always considered slow, flaky, and heavy. I benchmarked three different configurations around 10 runs each because one slow pipeline can blow up your average. So the first instance was C7a AMD-based, averaged 25 minutes, but actually it ranged from 22 to 32 minutes, really widespread. The second option was C7i Intel-based, and in general, it was slower, 27 minutes, but the spread was just 24 to 30. So we actually picked up Intel. Why? In a merge train, variance is more damaging than mean performance.
Comments