Lead and develop a team of Site Reliability Engineers (Level 6-7), initially 3 direct reports and growing as the Vietnam site consolidates, owning performance, coaching, retention, and day-to-day execution
Build individual development plans that grow Level 6 engineers toward independent Level 7 scope
Establish a high-ownership culture where engineers are accountable for outcomes, not tasks
Run team rituals: 1:1s, shift retrospectives, and development check-ins
Follow-the-Sun Operations
Own the Vietnam shift within GRE’s global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management
Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work
Serve as escalation point for the Vietnam shift during complex or high-severity incidents
Reliability Practice Execution
Own production reliability outcomes for the markets, platforms, and services within the Vietnam team’s scope
Drive SRE operating standards within the team: incident response rigor, SLO ownership, runbook quality, and post-incident follow-through
Ensure monitoring coverage, dashboards, and alerting remain accurate and effective across owned services
Enforce GRE-wide reliability standards, ensuring the Vietnam practice operates in alignment with the global model
Platform Engineering, Automation, and AI
Own the Vietnam team’s contribution to GRE’s reliability modernization roadmap: auto-healing, auto-remediation, and self-service issue mitigation
Treat automation as a delivery commitment, not a byproduct: plan, prioritize, and track toil-reduction and self-service work alongside operational coverage
Ensure the team builds with platform engineering practices: infrastructure as code, runbooks as code, GitOps workflows, and reusable tooling over one-off fixes
Champion AI-first engineering practices within the team, ensuring engineers develop with and through modern AI tooling
Partner with GRE’s Foundations and Intelligence pillars on observability, signal detection, and automated response infrastructure
Stakeholder Engagement
Represent the Vietnam team in East GRE planning and reliability forums
Communicate reliability outcomes, coverage status, and team health clearly to GRE leadership
Partner with the Associate Director on headcount planning, team scope, and Vietnam-specific delivery
Your skills and experience
5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles
1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership
Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
Experience operating in shift-based, on-call, or follow-the-sun coverage models
Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
Proficiency in at least one scripting or programming language sufficient to review and guide automation work
Understanding of SLI/SLO frameworks and reliability engineering fundamentals
Strong written and verbal English communication skills for cross-region collaboration with US and India teams
Experience building or standing up a new team, site, or shift operation
Experience managing engineers across the early-to-mid career range with a track record of promotions or level progression
Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
Experience in multi-region or globally distributed team models
Relevant certifications (AWS, CKA, or similar)
Why you'll love working here
Attractive Benefits:
100% salary during probation period
Annual Leave: 18 days/ year
2 days WFH/ week
Five “Recharge Days” – Extra days, in addition to company holidays.
Flexible Friday afternoon
Full salary insurance
13th-month bonus
1 day off for birthday
Advanced health insurance (Generali)
Regular engagement activities: sport clubs, internal event…