We get excited about candidatesBachelor's degree in Computer ScienceStrong problem-solving skills
Required
We get excited about candidates, like you, because you possess:
Bachelor's degree in Computer Science, Information Technology, or a related field.
8+ years of experience in database operations, site reliability engineering (SRE), or a related role with a heavy focus on data platforms and core operational infrastructure.
Strong proficiency in Automation and IaC frameworks, with extensive hands-on experience in building database IaC
Proven track record in Database Production Support and Operations, with deep practical experience managing robust, highly available 24x7 runtime systems (prior experience handling large-scale database production support at scale is highly valued).
Extensive experience with the AWS cloud platform and hands-on implementation of CI/CD pipelines and DatabaseOps workflows.
Proficiency in monitoring and observability tools (e.g., Datadog, CloudWatch, DevOps Guru, Database Performance Insights) to track metrics, latency, throughput, and system errors.
Strong understanding of various system performance metrics at a low level (such as Disk/IO saturation) and experience identifying or eliminating operational bottlenecks.
Familiarity with managing large-scale database structures across various systems (SQL and NoSQL).
Strong problem-solving skills, ownership mentality, proactive communication skills, and a baseline capability to document clear incident response playbooks and operational requirements.
The Database Operations Engineering team is dedicated to ensuring the reliability, scalability, and performance of our data infrastructure. We focus on standardizing and implementing monitoring and alerting across all datastores to track key metrics like errors, latency, and throughput, and to ensure critical systems are covered. Our team leads horizontal efforts to keep databases up-to-date, implements Infrastructure as Code (IaC) for high availability and performance, and automates key processes to enhance operational efficiency.
We lead and evangelize the principle of 100% automation. Additionally, we define and document operational requirements, develop incident response processes, and automate monitoring and compliance checks to maintain a secure and reliable data environment. By continuously improving load testing and optimizing data governance practices, we support the overall health and efficiency of our data systems.