Lead or coordinate incident response for assigned markets and services, with independent ownership of routine incidents and supported ownership of complex multi-system incidents
Facilitate communication during incidents and drive timely resolution, including clean shift handoffs across global time zones
Contribute to post-incident reviews and root cause analysis activities
Ensure corrective and preventive actions are identified, tracked, and completed for owned services
Monitoring, Alerting, and Observability
Implement and continuously optimize monitoring, logging, alerting, and tracing solutions
Develop meaningful alerts based on service behavior, customer impact, and business priorities
Build and maintain dashboards that provide actionable insights into system performance and reliability
Platform and Market Ownership
Own day-to-day SRE responsibilities for one or more assigned markets, platforms, or services
Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date
Assess platform health, identify reliability risks, and raise improvements before incidents occur
Partner with engineering teams to ensure new features and services meet reliability requirements before production release
Track reliability metrics for owned services, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
Platform Engineering, Automation, and AI
Develop and maintain tools, scripts, and automation that reduce manual effort and improve operational efficiency
Identify and eliminate repetitive tasks through automation, with a bias toward self-service capabilities over ticket queues
Contribute to auto-healing and auto-remediation capabilities: detection rules, remediation runbooks as code, and automated response workflows
Build and extend internal platform tooling using infrastructure as code and CI/CD pipelines rather than manual configuration
Apply team best practices for the responsible use of AI within SRE workflows, including AI-assisted diagnosis, triage, and remediation
Team Contribution
Share knowledge through documentation, training sessions, and post-incident learning activities
Review runbooks and operational documentation to maintain quality standards
Contribute to the continuous improvement of team processes, standards, and ways of working
Your skills and experience
2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
Working knowledge of at least one major cloud provider (AWS preferred)
Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
Experience participating in incident response and on-call or shift-based operations
Understanding of SLI/SLO concepts and reliability engineering fundamentals
Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
Strong written and verbal English communication skills for cross-region collaboration
Experience with Kubernetes, Docker, and container orchestration in production
Infrastructure as code experience (e.g., Terraform, CloudFormation)
CI/CD and GitOps pipeline experience (e.g., GitLab CI, ArgoCD)