Senior Site Reliability Engineer

LinkedIn|CBTW APAC|Vietnam|10 Aug 2026
Apply Now →

JOB DESCRIPTION

Overview

As a Senior Software Engineer within the SRE team, you will own the delivery of key infrastructure and reliability features for your team, applying a software engineering mindset to solve complex operational challenges. You will take full responsibility for the technical implementation of features you own — from initial design through to their behaviour in production.

Your focus will be on building robust, scalable, and observable platform components on our AWS and Kubernetes-based infrastructure, improving developer experience, and lifting the reliability of the services your team owns. You will set a high bar for engineering quality, mentor engineers around you, and ensure the reliability and performance of what you ship align with team and business goals.


Responsibilities

  • Take full ownership of the technical implementation of key platform and reliability features — from design through to production behaviour — ensuring solutions on our AWS and Kubernetes-based infrastructure are robust, scalable, and maintainable.
  • Design for operability by building in observability, graceful degradation, and recoverability so the systems your team owns are understandable and safe to run in production.
  • Define and maintain SLOs for services your team owns, working with product and stakeholders to set meaningful reliability targets, and make data-driven cases for system health improvements.
  • Lead incident response for your team’s systems — coordinating investigation, communicating user impact to stakeholders, and driving post-mortems through to completed corrective actions.
  • Actively reduce toil and incident risk by identifying and eliminating root causes, improving runbooks, and keeping on-call practices healthy (fair rosters, current documentation, clean handovers)
  • Champion a developer-centric platform with efficient CI/CD pipelines (GitHub Actions, ArgoCD) and internal tooling that removes bottlenecks and enables high team velocity.
  • Use AI-assisted engineering workflows (e.g. Claude Code) to improve delivery velocity, code quality, and developer experience.


Qualifications

  • 5+ years in infrastructure or platform/SRE engineering roles, with hands-on delivery in high-scale production environments.
  • Deep hands-on expertise with AWS, Kubernetes (pref EKS), and Terraform, and modern CI/CD tooling (e.g. GitHub Actions, ArgoCD).
  • Solid understanding of Site Reliability principles — SLOs, error budgets, alerting, on-call, and the incident lifecycle - with proven experience running incident response and post-mortems.
  • Ability to own the technical implementation of features end-to-end, making sound independent decisions in the face of open-ended requirements.
  • Hands-on experience using AI-assisted development tools (e.g. Claude Code) in day-to-day engineering work.
  • Clear written and verbal communication - able to present technical arguments concisely and adapt the message to the audience


Nice to have

  • Experience with observability/incident management platforms such as Honeycomb.io or Incident.io.
  • Experience with messaging technologies such as Apache Kafka (or AWS MSK), RabbitMQ, AWS SQS etc.
  • Experience with databases such as RDS, MySQL, or PostgreSQL.