How We Migrated onto K8s in Less Than 12 months | Figma Blog (opens in new tab)
Figma migrated most of its core services from AWS ECS to Kubernetes in under 12 months because ECS was increasingly limiting its platform ambitions. Kubernetes offered better support for stateful workloads, Helm-based software, autoscaling, service networking, and the broader CNCF ecosystem. The migration was considered worthwhile because Figma had relatively few core services and had already containerized its workloads, making the transition more manageable.
Figma’s Existing Compute Platform
- By early 2023, Figma was already running all services in containers on Amazon ECS.
- ECS had enabled rapid adoption of containerized workloads, but Figma’s growing infrastructure team began evaluating a more capable long-term platform.
- Figma is not organized around thousands of microservices:
- A small set of powerful core services provides modularization and traffic isolation.
- New product capabilities are usually added to existing services rather than creating new ones.
- This limited service count made a Kubernetes migration more practical.
Limitations of ECS
- ECS lacked Kubernetes primitives needed for complex workloads.
- Running
etcdon ECS required fragile custom startup code to manage cluster membership because ECS does not provide StatefulSets or persistent pod identity. - Kubernetes StatefulSets provide stable identities and stateful networking for systems such as
etcd. - ECS did not natively support deploying groups of services packaged as Helm charts.
- Open-source tools such as Temporal would require manual conversion into Terraform configurations.
- This increased installation and maintenance effort.
- ECS also made routine infrastructure operations more cumbersome.
- For example, safely removing a malfunctioning EC2 instance was difficult.
- EKS can cordon a node and move its pods elsewhere while respecting graceful shutdown behavior.
Access to the CNCF Ecosystem
- Kubernetes would give Figma access to a larger ecosystem of open-source cloud-native tools.
- Autoscaling was a major motivation:
- Figma was provisioning services for peak demand, wasting resources during lower-traffic periods.
- Kubernetes tooling such as KEDA supports scaling based on CPU, SQS queue length, and custom Datadog metrics.
- Figma expected to adopt a service mesh eventually.
- Existing AWS load balancer routing created operational drawbacks:
- Network Load Balancers could take several minutes to register or remove targets.
- This slowed emergency deployments and increased incident remediation time.
- Envoy offered more customization than AWS load balancers, including custom filters for shedding load during incidents.
- Figma had already deployed standalone Envoy machines for a major service and saw Kubernetes ecosystems such as Istio as a path toward fleet-wide service-mesh adoption.
Figma’s experience suggests that Kubernetes was justified not simply as a replacement for ECS, but as a foundation for more capable operations and broader platform tooling. Organizations considering a similar move should first assess their workload complexity, existing container maturity, and whether Kubernetes capabilities will materially reduce infrastructure work.