JOB DESCRIPTION
About the Role
We are looking for a DataOps Engineer to build and operate the data infrastructure that
powers VinMotion's next-generation humanoid robots. In this role, you will develop scalable
data pipelines, manage cloud-based data platforms, and support the delivery of high-quality
datasets for AI model training and deployment.
Key Responsibilities
DataOps Management
• Experience deploying, scaling, and managing Highly Available (HA) Kubernetes clusters on bare-metal or on-premises infrastructure (e.g., K3s, RKE2).
• Proficiency managing embedded datastores (e.g., etcd) and integrating on-premises storage using Kubernetes CSI drivers.
• Hands-on experience with declarative cluster management, leveraging tools like Helm, Helmfile, and Kubernetes Operators.
• Understanding of on-premises container networking, ingress controllers, and load balancing without reliance on managed cloud providers.
Data Storage
• Deep understanding of on-premises distributed object and file storage architectures
(e.g., MinIO, Ceph).
• Experience deploying and tuning databases for AI/ML and robotics workloads, including vector databases (e.g., Milvus), document stores (e.g., MongoDB), and relational databases (e.g., PostgreSQL).
• Familiarity with optimizing storage and access patterns for heavy multimodal datasets and specialized robotics formats (e.g., MCAP, Rosbag, Parquet, Delta Lake).
• Proven ability to architect storage infrastructure for fault tolerance, data replication, and rapid disaster recovery.
Data Pipeline Development
• Build and maintain scalable DataOps pipelines for collecting, processing, and distributing robotics datasets.
• Develop automated ETL/ELT workflows to ingest data from robots, simulation environments, and cloud systems.
• Ensure reliable data movement across edge devices, cloud infrastructure, and AI training
platforms.
DevOps & MLOps Collaboration
• Work with DevOps teams to improve CI/CD pipelines supporting data and AI workflows.
• Collaborate with AI and MLOps engineers to prepare datasets for model training and evaluation.
• Support model deployment workflows by ensuring reliable access to training and inference data.
• Contribute to automation and observability across the data platform.
• Monitoring data infrastructure, data pipelines, and loggings
Qualifications
Required
• Bachelor's or Master's degree in Computer Science, Data Engineering, Software Engineering, or a related field.
• 2 years of experience in DataOps, DevOps, Data Engineering, Platform Engineering, or Cloud Infrastructure.
• Hands-on experience with Docker, Kubernetes
• Strong programming skills in Python and SQL; familiarity with Bash scripting.
• Experience building data pipelines and ETL/ELT workflows.
• Familiarity with workflow orchestration tools (Airflow, Temporal, or similar).
• Experience with data storage technologies and formats such as Parquet, Delta Lake, or Iceberg.
• Knowledge of Linux, Git, CI/CD pipelines, and monitoring tools.
• Strong problem-solving skills and ability to work in a fast-paced, collaborative
environment.
Nice to Have
• Experience with AWS services such as S3, EC2, IAM, Glue, or EMR.
• Infrastructure-as-Code tools such as Terraform or Ansible.
• Familiar with robot data format (LeRobot, Zstd, Rosbag, Mcap,..)
• Have experiences with architecture designs
• Have experiences with Frontend, backend development
• Exposure to MLOps practices, including model packaging, ML pipelines, model
deployment, or model versioning (e.g., MLflow).
• Experience supporting AI/ML training infrastructure or GPU-based workloads.
• Familiarity with robotics, ROS 2, IoT, or edge computing environments.
• Experience with streaming platforms such as Apache Kafka or Amazon Kinesis.
• Knowledge of robotics simulation platforms such as NVIDIA Isaac Sim, Mujoco or Gazebo.