Middle DevOps Engineer/Site Reliability Engineer (Global Product)

Indeed|Logix Technology|Thành phố Hồ Chí Minh|7 Aug 2026
Apply Now →

JOB DESCRIPTION

LOGIX TECHNOLOGY is a distinguished software services company, specializing in the provision of professional software development services, fostering strong partnerships through the establishment of offshore development centers, offshore product development, and software testing with an infrastructure-oriented approach. We excel at simplifying and enhancing offshore outsourcing, not only by delivering innovative solutions to meet our clients’ exact requirements through high-performance teams but also by significantly contributing to our clients’ success by reducing time-to-market and improving quality.

At LOGIX TECHNOLOGY, we are committed to delivering the highest level of professional excellence and making a significant impact in the software development sector. If you are a dedicated professional looking to join a dynamic team working with global leaders, we encourage you to explore the exciting opportunities that await you with us.

Job Summary

We are looking for a Middle DevOps Engineer/Site Reliability Engineer (SRE) to help us build, operate, and continuously improve the reliability, scalability, and performance of our production systems. This role sits at the intersection of software engineering and systems operations – applying software engineering principles to infrastructure and operations problems and treating operational work (including incident response) as an engineering discipline rather than pure firefighting.
A core part of this role is Application Operations – continuously monitoring application health, triaging issues (including Operation tickets and alerts) as they come in, and actively participating in the incident resolution process when problems are confirmed, so issues are caught early and resolved effectively.

Job Description

Reliability & Platform Engineering

  • Design, build, and maintain scalable, resilient, and secure infrastructure (cloud, containers, orchestration platforms).
  • Define and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for key services.
  • Build and maintain observability: metrics, logging, distributed tracing, dashboards, and alerting.
  • Automate manual operational tasks (toil reduction) through tooling, scripting, and infrastructure-as-code.
  • Participate in architecture and design reviews, representing reliability, scalability, and operability concerns.
  • Conduct capacity planning and performance testing to anticipate and prevent failures before they occur.

Application Operations

This role is responsible for the day-to-day operational health of production applications -keeping a close eye on how systems are behaving, catching problems early, and playing an active role when things go wrong:

  • Application monitoring: Continuously monitor application health, performance, and behavior using dashboards, metrics, and alerting; proactively spot early warning signs before they become customer-facing problems.
  • Issue triage: Review incoming alerts and reported issues to assess severity, likely impact, and urgency, and route them appropriately – resolve directly, escalate, or flag for deeper investigation.
  • Involvement in incident resolution: Actively participate in the incident resolution process once an issue is confirmed as an incident – supporting the response effort with diagnosis, data gathering, and remediation steps under the direction of the incident lead .
  • Pattern recognition: Notice recurring issues or noisy alerts over time and flag them for follow-up (e.g., tuning alert thresholds, requesting a permanent fix) rather than repeatedly triaging the same symptoms.
  • Operational reporting: Contribute status and health updates during active issues, and provide input to postmortems and ticket-trend reviews for problems worked on.

Collaboration & Enablement

  • Partner with software engineering teams to review designs for reliability, failure modes, and operational readiness (production readiness reviews) .
  • Provide guidance and coaching to development teams on reliability best practices, observability instrumentation, and incident response .
  • Contribute to a healthy on-call culture: sustainable rotations, clear expectations, and support for on-call engineers.

Required Skills & Experience

  • 3+ years of experience in Site Reliability Engineering/DevOps/Systems/Platform Engineering, or a related infrastructure/operations role.
  • Solid experience operating production systems in a cloud environment (Azure preferred, AWS, GCP).
  • Hands-on experience with containerization and orchestration (Docker, Kubernetes).
  • Scripting/programming proficiency in at least one language (e.g., Python, Go, Bash).
  • Practical experience with monitoring/observability tooling (e.g., Prometheus, Grafana, Datadog, New Relic, ELK/EFK stack).
  • Direct experience participating in an on-call rotation and responding to live production incidents
  • Experience working with ticket-based support/service-management systems (e.g., Jira Service
  • Management, ServiceNow, or an internal ticketing platform) to triage and resolve operational requests within SLA.
  • Familiarity with incident management frameworks (e.g., ITIL incident management, Google SRE incident response model) and postmortem/RCA practices.
  • Working knowledge of networking, Linux systems administration, and distributed systems fundamentals.
  • Strong written and verbal communication skills, especially the ability to communicate clearly under pressure during an active incident.

Benefits and Perks

  • Attractive Salary & Benefits, full salary in probation.
  • Performance appraisal twice a year, 13th-month salary.
  • Healthcare and incident insurance.
  • 16 Paid Leave days / year (12 Annual Leave, 02 Summer Leave ,01 Birthday Leave & 01 Christmas Leave).
  • Company trip, annual year-end party, team building activities.
  • Diverse careers opportunities with Software Services, Software Product Development
  • Working and growing in a value driven, international working environment and standard Agile culture with passionate and talented teams.
  • Various training on hot-trend technologies, best practices, and soft skills.
  • Free in-house entertainment facilities: football, PS, badminton, coffee, and snacks (instant noodles, cookies, candies…).

Full name

Email address

Message

Upload CV
Upload your CV/resume or any other relevant file. Max. file size: 300 MB.

I accept the Terms and Conditions.
Are you human?

  • Category Software Development
  • Job Type
    Full time
  • Location Ho Chi Minh city
  • Date posted Posted 19 hours ago
  • Expiration date 4th November 2026