Skydio Site Reliability Engineer to build, operate, and scale cloud infrastructure including Kubernetes and AWS at Skydio. Focus on production infrastructure reliability and scalability.
Responsibilities
We are looking for a hands-on Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure, including Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability.
You don't need to be an expert in every area, but you should have strong Kubernetes and cloud fundamentals with meaningful depth in at least one infrastructure domain. Our technology helps save lives. You’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most.
How you'll make an impact:
Build, operate, and troubleshoot production Kubernetes/EKS clusters.
Perform Kubernetes upgrades, node rollouts, and cluster maintenance.
Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.
Define and maintain infrastructure using Terraform.
Build and operate CI/CD and deployment infrastructure.
Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.
Build monitoring, alerting, and observability for critical infrastructure.
Participate in on-call rotations and respond to production incidents.
Identify and solve infrastructure scaling and reliability problems.
Qualification
Strong AWS fundamentals including VPCsHelm and GitOps experienceMulti-region infrastructure experience
Required
3+ years of experience as a Production Engineer, SRE, DevOps, or equivalent infrastructure role
Strong hands-on experience operating Kubernetes
Experience managing Kubernetes/EKS upgrades and production clusters
Strong AWS fundamentals including VPCs, networking, load balancers, EKS, IAM, and databases
Production experience with Terraform or similar infrastructure-as-code tooling
Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins
Experience diagnosing production infrastructure and networking problems
Experience solving meaningful scaling or reliability challenges
Preferred
Helm and GitOps experience.
Datadog or similar observability tooling.
PostgreSQL/database operations experience.
Multi-region infrastructure experience.
On-premises or disconnected deployment experience.
Streaming or high-throughput distributed systems experience.