Overview
As a Senior Software Engineer within the SRE team, you will own the delivery of key infrastructure and reliability features for your team, applying a software engineering mindset to solve complex operational challenges. You will take full responsibility for the technical implementation of features you own — from initial design through to their behaviour in production.
Your focus will be on building robust, scalable, and observable platform components on our AWS and Kubernetes-based infrastructure, improving developer experience, and lifting the reliability of the services your team owns. You will set a high bar for engineering quality, mentor engineers around you, and ensure the reliability and performance of what you ship align with team and business goals.
Responsibilities
-
Take full ownership of the technical implementation of key platform and reliability features — from design through to production behaviour — ensuring solutions on our AWS and Kubernetes-based infrastructure are robust, scalable, and maintainable.
-
Design for operability by building in observability, graceful degradation, and recoverability so the systems your team owns are understandable and safe to run in production.
-
Define and maintain SLOs for services your team owns, working with product and stakeholders to set meaningful reliability targets, and make data-driven cases for system health improvements.
-
Lead incident response for your team’s systems — coordinating investigation, communicating user impact to stakeholders, and driving post-mortems through to completed corrective actions.
-
Actively reduce toil and incident risk by identifying and eliminating root causes, improving runbooks, and keeping on-call practices healthy (fair rosters, current documentation, clean handovers)
-
Champion a developer-centric platform with efficient CI/CD pipelines (GitHub Actions, ArgoCD) and internal tooling that removes bottlenecks and enables high team velocity.
-
Use AI-assisted engineering workflows (e.g. Claude Code) to improve delivery velocity, code quality, and developer experience.
Qualifications
-
5+ years in infrastructure or platform/SRE engineering roles, with hands-on delivery in high-scale production environments.
-
Deep hands-on expertise with AWS, Kubernetes (pref EKS), and Terraform, and modern CI/CD tooling (e.g. GitHub Actions, ArgoCD).
-
Solid understanding of Site Reliability principles — SLOs, error budgets, alerting, on-call, and the incident lifecycle - with proven experience running incident response and post-mortems.
-
Ability to own the technical implementation of features end-to-end, making sound independent decisions in the face of open-ended requirements.
-
Hands-on experience using AI-assisted development tools (e.g. Claude Code) in day-to-day engineering work.
-
Clear written and verbal communication - able to present technical arguments concisely and adapt the message to the audience
Nice to have
-
Experience with observability/incident management platforms such as Honeycomb.io or Incident.io.
-
Experience with messaging technologies such as Apache Kafka (or AWS MSK), RabbitMQ, AWS SQS etc.
-
Experience with databases such as RDS, MySQL, or PostgreSQL.