Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning in the CI/CD process
Implement and maintain policy-as-code to block unsafe infrastructure and Kubernetes changes at admission time
Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings
Partner with Security Engineer and broader Security & Reliability disciplines
Participate in and improve the on-call rotation; build the runbooks and automation
Coach team members and engineers across the org on reliability patterns and operational best practices