Reducing CI Costs for Flutter Apps: Spot Instances and Resource Optimization
The $14,000/Month CI/CD Leak
When I first audited our engineering infrastructure, the CI/CD pipeline costs stood out as a massive, unoptimized hemorrhage. We were running 25 dedicated c5.2xlarge instances on AWS 24/7 to handle Flutter build agents for our iOS and Android build tracks. At an on-demand price point, that was costing us roughly $14,000 per month.
The issue wasn't just the sheer number of instances; it was the architecture. Our pipeline was treating build agents as persistent infrastructure rather than ephemeral compute tasks. In cloud cost engineering, persistence is a luxury that your P&L statement usually cannot afford. If you are paying for an instance while the build queue is empty, you are effectively burning cash to heat the data center. By shifting our Flutter build pipeline to a reactive, spot-instance-based architecture, we reduced that $14,000 monthly burn to approximately $3,800, creating an annual saving of over $122,000.
Diagnosis: The Idle Resource Trap
Most teams build their Flutter CI environment by spinning up a Jenkins or GitHub Actions Runner controller with fixed worker nodes. This creates a predictable but expensive outcome: your compute costs scale linearly with time, not with demand. If your team pushes code at 10:00 AM, the build agents are saturated. If those same agents sit idle at 2:00 AM, you are still paying for the full resource footprint.
Flutter builds are resource-intensive. Compiling Dart code to machine code for iOS (AOT compilation) and generating Android bundles (R8/ProGuard) requires significant CPU and RAM bursts. However, these bursts are intermittent. We diagnosed our usage and found that while our peak utilization hit 90%, our average utilization across the 24-hour cycle was closer to 22%. That 68% gap was our primary cost optimization opportunity.
Redesign: Moving to Ephemeral Spot Fleets
To capture those savings, we had to dismantle the persistent worker model. The solution was implementing an ephemeral runner architecture using AWS Spot Instances. Spot instances offer a discount of up to 90% compared to on-demand pricing, provided you can handle the eventuality of an interruption signal.
For a CI pipeline, interruptions are not a deal-breaker if you have proper orchestration. If a spot instance is reclaimed by AWS, the build fails, and the runner controller triggers a retry. The cost of a few retried builds is statistically insignificant compared to the 80% savings on total compute hours.
Step-by-Step Implementation Guide
- Orchestration Layer: Swap static nodes for an auto-scaling group (ASG) or a container-based orchestrator (like Kubernetes/EKS). We chose EKS with the Karpenter autoscaler because it scales up and down in seconds, not minutes.
- Spot Instance Mapping: Configure your ASG or Karpenter node pools to prioritize Spot instances. If Spot capacity is unavailable, allow for a small fallback pool of On-Demand instances to prevent complete pipeline stalls.
- Dockerizing the Flutter Environment: Build a custom Docker image that contains the Flutter SDK, Android SDK, and Xcode (if running on macOS nodes). This ensures the node is 'ready' the moment it boots.
- Lifecycle Management: Implement a cleanup script that shuts down the node immediately after the job finishes. Use an
agent-timeoutsetting of 5 minutes to prevent 'zombie' nodes.
# Example Karpenter Provisioner configuration for Flutter Builds
apiVersion: karpenter.sh/v1alpha5
kind: Provisioner
metadata:
name: flutter-build-pool
spec:
requirements:
- key: "karpenter.sh/capacity-type"
operator: In
values: ["spot"]
- key: "node.kubernetes.io/instance-type"
operator: In
values: ["c5.2xlarge", "c5.4xlarge"]
limits:
resources:
cpu: "100"
provider:
instanceProfile: BuildNodeProfile
The Quantitative Impact: Before and After
Before the migration, we were bound by static pricing. We used 25 instances at $0.34/hour (c5.2xlarge). Total cost calculation: 25 instances * 730 hours/month * $0.34 = $6,205 per environment (we ran multiple environments, totaling $14k).
Post-migration, our strategy shifted to using Spot instances which averaged $0.09/hour. We implemented a 'scale-to-zero' policy where our compute fleet drops to zero during non-working hours.
- Monthly Savings: $10,200
- Annualized Savings: $122,400
- Build Time Impact: An increase of 3% due to node cold-starts, which was mitigated by keeping a single 'warm' node during peak hours.
By treating CI/CD capacity as a commodity market rather than a fixed asset, we aligned our infrastructure spend with our engineering activity. Engineering managers often fear the 'interruption' aspect of spot instances, but when you quantify the cost of a build retry (approx. $0.15 in wasted compute) against the $10k/month savings, the math is overwhelmingly in favor of spot.
Pro-Tips for CI Optimization
- Pro-Tip 1: Cache Layers are Non-Negotiable. Even with spot instances, if your build takes 20 minutes to resolve Pub dependencies, you're wasting time and money. Use
pub getcaching in your build agents. We store our.pub-cachein an S3 bucket and mount it as a volume to our build containers. This reduced our average build time by 4 minutes per run. - Pro-Tip 2: Multi-Architecture Handling. Flutter developers often build for multiple targets. Do not use the same instance size for everything. Android compilation is CPU-bound; iOS compilation on remote runners is RAM-bound. Profile your actual builds and create specific node pools for 'Android-Heavy' and 'iOS-Heavy' tasks to avoid over-provisioning.
- Pro-Tip 3: The Interruption Handler. If you are using Jenkins, install the 'Amazon EC2 Spot Instances' plugin. If using GitHub Actions, ensure your self-hosted runners are wrapped in a script that catches the SIGTERM signal to clean up gracefully before the spot reclaim hits. This prevents corrupted workspace states that lead to expensive, non-deterministic CI failures.
- Pro-Tip 4: Clean Up Artifacts. Many engineers forget that build artifacts occupy expensive EBS storage. Set up an S3 Lifecycle Policy to expire artifacts older than 7 days. Storing 5 TB of old APKs/IPAs in high-performance EBS volumes is a silent budget killer.
Conclusion: Infrastructure as Code, Not Capital
Engineering efficiency is rarely about choosing the 'latest' tech; it is about choosing the most 'economical' tech. Flutter builds are deterministic tasks. They have a start, a duration, and an end. There is no reason to pay for capacity that isn't being utilized. By shifting to a spot-first infrastructure and utilizing autoscaling to hit zero during idle periods, we turned a major operational cost into a lean, highly efficient service.
If you take away one thing from this analysis, it is this: measure your average utilization, not your peak capacity. If your utilization is below 40%, you are effectively overpaying by at least double. Redesigning for spot instances and cold-start optimization isn't just about saving money—it's about building a culture where infrastructure spend is treated as a first-class engineering concern, equivalent to code complexity or latency. We moved the $122k we saved back into our R&D budget, hiring a senior engineer who has since focused on performance optimization of our app's core rendering engine. That is the leverage cloud cost engineering provides.