Key Responsibilities
Incident Management and Reliability
-
Lead or coordinate incident response for assigned markets and services, with independent ownership of routine incidents and supported ownership of complex multi-system incidents
-
Facilitate communication during incidents and drive timely resolution, including clean shift handoffs across global time zones
-
Contribute to post-incident reviews and root cause analysis activities
-
Ensure corrective and preventive actions are identified, tracked, and completed for owned services
Monitoring, Alerting, and Observability
-
Implement and continuously optimize monitoring, logging, alerting, and tracing solutions
-
Develop meaningful alerts based on service behavior, customer impact, and business priorities
-
Build and maintain dashboards that provide actionable insights into system performance and reliability
Platform and Market Owners
-
Own day-to-day SRE responsibilities for one or more assigned markets, platforms, or services
-
Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date
-
Assess platform health, identify reliability risks, and raise improvements before incidents occur
-
Partner with engineering teams to ensure new features and services meet reliability requirements before production release
-
Track reliability metrics for owned services, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
Platform Engineering, Automation and AI
-
Develop and maintain tools, scripts, and automation that reduce manual effort and improve operational efficiency
-
Identify and eliminate repetitive tasks through automation, with a bias toward self-service capabilities over ticket queues.
-
Contribute to auto-healing and auto-remediation capabilities: detection rules, remediation runbooks as code, and automated response workflows.
-
Build and extend internal platform tooling using infrastructure as code and CI/CD pipelines rather than manual configuration
-
Apply team best practices for the responsible use of AI within SRE workflows, including AI-assisted diagnosis, triage, and remediation
Mandatory Skills
-
2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles.
-
Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
-
Working knowledge of at least one major cloud provider (AWS preferred)
-
Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work.
-
Experience participating in incident response and on-call or shift-based operations.
-
Understanding of SLI/SLO concepts and reliability engineering fundamentals.
-
Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions.
-
Strong written and verbal English communication skills for cross-region collaboration.
-
Experience with Kubernetes, Docker, and container orchestration in production.
-
Infrastructure as code experience (e.g., Terraform, CloudFormation).
-
Experience building or contributing to auto-remediation or event-driven automation workflows.
-
Experience supporting distributed systems across multiple markets or regions.
-
Relevant certifications (AWS, CKA, or similar)