Build and operate batch and near-real-time data pipelines on our lakehouse platform, powering reporting, analytics, and data products across the company.
Own assigned data domains end to end — from source ingestion to the trusted, documented tables that business and technical teams rely on.
Design and maintain data models that turn raw operational data into reusable, well-structured datasets instead of one-off queries.
Improve pipeline performance, reliability, and infrastructure cost as data volume and business complexity grow.
Establish and maintain data quality standards, including validation, freshness monitoring, and alerting on critical datasets.
Keep metadata, lineage, and documentation up to date so data consumers can find and trust the data they use.
Support BI, Analytics, and Data Science teams by preparing serving datasets and resolving data issues they raise.
Partner with Backend and Product teams to define event and CDC data contracts for new and existing services.
Participate in the team's on-call rotation; investigate incidents, restore data SLAs, and drive follow-up improvements.
Contribute to code review, technical documentation, and the team's deployment and engineering practices.
Your skills and experience
2+ years of hands-on experience in Data Engineering, or an equivalent role with ownership of production data pipelines.
Strong SQL, including window functions, complex CTEs, and query optimization on tables with hundreds of millions of rows.
Production-level Python: structured, tested, and packaged code — not only notebooks or one-off scripts.
Hands-on Spark experience (PySpark or Scala): partitioning, shuffle, join strategies, data skew, and the ability to read and interpret execution plans.
Practical experience with a workflow orchestrator; Airflow strongly preferred.
Solid understanding of data warehouse and lakehouse fundamentals: dimensional modeling, incremental loads, CDC handling, snapshot vs delta processing.
Comfortable with Git and CI/CD workflows as part of daily development.
Working knowledge of Docker and Kubernetes: able to containerize a job, read pod logs and events, understand resource requests and limits, and debug a failing or evicted pod.
Strong debugging discipline: reads logs and metrics, isolates variables, and verifies hypotheses with data instead of guessing.
Able to read English technical documentation and open-source code independently.
Nice to have
Production experience with Apache Iceberg, Delta Lake, or Hudi, including snapshot, manifest, and metadata layer concepts.
Experience with Trino/Presto, or OLAP engines such as Apache Doris, ClickHouse, or StarRocks.
Deeper Kubernetes experience: Helm, ArgoCD/GitOps, and resource tuning for Spark executors.
Streaming experience with Spark Structured Streaming or Flink, including state, watermark, and exactly-once semantics.
Familiarity with data catalog and quality tooling such as DataHub, OpenMetadata, dbt, or Great Expectations.
Experience with Vault, External Secrets Operator, Terraform, or Terragrunt.
Exposure to logistics, marketplace, mobility, e-commerce, or on-demand platform data.
Experience with billing, reconciliation, or finance data pipelines.
Why you'll love working here
Physical Wellbeing Benefit: General Insurance, Medical check-up, Accident Insurance, Healthcare Insurance.
Emotional Wellbeing Benefit: Company Trip, Year End Party, Aha Hour Activities, Special Day Gifts, Aha Club (Badminton, Soccer).
Financial Wellbeing Benefit: Grab/Be For Work (Tech/Lead Level), Workplace Relocation, 13th Month Salary, PP Appreciate, Annual Leave Remain.