JOB DESCRIPTION
About the Role
We're hiring a Site Reliability Engineer to make Techcombank's data platform reliable. You'll write real code, own production systems end-to-end, and replace manual toil with automation — increasingly with AI in the loop.
Our platform runs on AWS: Databricks and Unity Catalog, EKS, Glue, Lambda, MSK/Kafka, Flink, Airflow, provisioned with Terraform and delivered through Jenkins, GitLab and ArgoCD. Data engineers, scientists and analysts across the bank depend on it.
We're deliberately open on background. Strong candidates come from SRE, backend/platform software engineering, or infrastructure and DevOps. What matters is that you build software to solve operational problems rather than manually absorbing them.
What You'll Do
- Own reliability for production data platform services: define SLOs, run incident response, drive blameless postmortems, and fix root causes.
- Build automation and internal tooling that removes toil — self-service provisioning, automated remediation, guardrails in CI/CD.
- Improve observability across ingestion, transformation and delivery so failures surface before users notice.
- Harden the delivery path: safer deployments, faster rollbacks, change risk caught before production.
- Apply AI where it pays off — incident summarisation, log and metric triage, RCA assistance, runbooks that execute instead of instruct. We'll support you in learning this if it's new.
- Partner with data engineering squads and mentor engineers on reliability practice.
What We're Looking For
- 8+ years building and operating production systems as an SRE, software engineer, or infrastructure/DevOps engineer.
- Strong coding ability in Python (or Go/Scala/Java) — you ship tools and services, not just scripts.
- Solid experience on AWS and with Infrastructure as Code (Terraform preferred).
- Hands-on with containers and Kubernetes, and with CI/CD pipelines.
- Practical observability experience: metrics, logs, tracing, alerting that people trust.
- Good at English
Nice to Have
Any of these will make you stand out; none are required.
- Data platform experience: Databricks, Spark, Kafka, Flink, Airflow, Glue.
- Using or building with GenAI — LLM APIs, prompt engineering, RAG, agents in production, AI-assisted coding at scale.
- Data quality, lineage, or governance tooling (Great Expectations, Deequ, Unity Catalog).
- Incident management practice in a regulated environment.
- Clear writing. Postmortems and runbooks are engineering artefacts here.
Why Join us?
At Techcombank, you will have the resources of a leading bank and the opportunity to directly shape data-driven strategies that impact tens of millions of customers.
We offer: