DataOps Engineer

LinkedIn|VinMotion|Hanoi, Hanoi, Vietnam|11 Aug 2026
Apply Now →

JOB DESCRIPTION

About the Role

We are looking for a DataOps Engineer to build and operate the data infrastructure that

powers VinMotion's next-generation humanoid robots. In this role, you will develop scalable

data pipelines, manage cloud-based data platforms, and support the delivery of high-quality

datasets for AI model training and deployment.


Key Responsibilities

DataOps Management

• Experience deploying, scaling, and managing Highly Available (HA) Kubernetes clusters on bare-metal or on-premises infrastructure (e.g., K3s, RKE2).

• Proficiency managing embedded datastores (e.g., etcd) and integrating on-premises storage using Kubernetes CSI drivers.

• Hands-on experience with declarative cluster management, leveraging tools like Helm, Helmfile, and Kubernetes Operators.

• Understanding of on-premises container networking, ingress controllers, and load balancing without reliance on managed cloud providers.


Data Storage

• Deep understanding of on-premises distributed object and file storage architectures

(e.g., MinIO, Ceph).

• Experience deploying and tuning databases for AI/ML and robotics workloads, including vector databases (e.g., Milvus), document stores (e.g., MongoDB), and relational databases (e.g., PostgreSQL).

• Familiarity with optimizing storage and access patterns for heavy multimodal datasets and specialized robotics formats (e.g., MCAP, Rosbag, Parquet, Delta Lake).

• Proven ability to architect storage infrastructure for fault tolerance, data replication, and rapid disaster recovery.


Data Pipeline Development

• Build and maintain scalable DataOps pipelines for collecting, processing, and distributing robotics datasets.

• Develop automated ETL/ELT workflows to ingest data from robots, simulation environments, and cloud systems.

• Ensure reliable data movement across edge devices, cloud infrastructure, and AI training

platforms.


DevOps & MLOps Collaboration

• Work with DevOps teams to improve CI/CD pipelines supporting data and AI workflows.

• Collaborate with AI and MLOps engineers to prepare datasets for model training and evaluation.

• Support model deployment workflows by ensuring reliable access to training and inference data.

• Contribute to automation and observability across the data platform.

• Monitoring data infrastructure, data pipelines, and loggings


Qualifications

Required

• Bachelor's or Master's degree in Computer Science, Data Engineering, Software Engineering, or a related field.

• 2 years of experience in DataOps, DevOps, Data Engineering, Platform Engineering, or Cloud Infrastructure.

• Hands-on experience with Docker, Kubernetes

• Strong programming skills in Python and SQL; familiarity with Bash scripting.

• Experience building data pipelines and ETL/ELT workflows.

• Familiarity with workflow orchestration tools (Airflow, Temporal, or similar).

• Experience with data storage technologies and formats such as Parquet, Delta Lake, or Iceberg.

• Knowledge of Linux, Git, CI/CD pipelines, and monitoring tools.

• Strong problem-solving skills and ability to work in a fast-paced, collaborative

environment.


Nice to Have

• Experience with AWS services such as S3, EC2, IAM, Glue, or EMR.

• Infrastructure-as-Code tools such as Terraform or Ansible.

• Familiar with robot data format (LeRobot, Zstd, Rosbag, Mcap,..)

• Have experiences with architecture designs

• Have experiences with Frontend, backend development

• Exposure to MLOps practices, including model packaging, ML pipelines, model

deployment, or model versioning (e.g., MLflow).

• Experience supporting AI/ML training infrastructure or GPU-based workloads.

• Familiarity with robotics, ROS 2, IoT, or edge computing environments.

• Experience with streaming platforms such as Apache Kafka or Amazon Kinesis.

• Knowledge of robotics simulation platforms such as NVIDIA Isaac Sim, Mujoco or Gazebo.