Computer Vision Engineer (Ha Noi)

LinkedIn|VinDynamics|Vietnam|6 Aug 2026
Apply Now →

JOB DESCRIPTION

Job summary: Architect and develop advanced Computer Vision pipelines to automate the ingestion, quality control (QA/QC), and semantic analysis of massive egocentric (first-person) video datasets. Your work will directly ensure the quality of training data for our embodied AI and humanoid robotics models.


JOB DESCRIPTION:

  • Automated Data QA/QC Pipeline: Design and implement an end-to-end automated pipeline to parse, validate, and process incoming video batches against strict programmatic standards (metadata, FPS, aspect ratio, duration, compression, and batch-list matching).
  • Egocentric Interaction Analysis: Develop robust CV modules to analyze first-person view (FPV) mechanics:
  • Ensure hands are continuously visible, working, and not occluded at the frame edges.
  • Calculate spatial approximations to guarantee manipulated objects remain centered in the frame.
  • Action & Task Verification: Implement video-language models or action classifiers to automatically verify that the recorded actions semantically match the assigned task descriptions.
  • Visual & Sensor Quality Control: Build heuristic and AI-based checks to filter out videos with excessive camera shake, poor lighting conditions (over/underexposed), or missing/corrupted head-mounted IMU data.
  • Automated Anonymization: Architect a zero-leakage PII redaction system to scan all frames for human faces, children, and sensitive text (IDs, licenses, logos), applying automated blurring before data enters the training lake.
  • Audio & Speech Validation: Integrate lightweight audio processing to verify the presence of audio tracks and extract/validate spoken voice commands from actors.
  • Performance Optimization: Optimize the entire video scanning pipeline for high-throughput execution on AWS infrastructure, minimizing processing time and compute costs per GB of video.


REQUIREMENTS:

Computer Vision & Deep Learning Expertise:

  • Egocentric Vision & Hand Tracking: Deep expertise in hand pose estimation, object tracking, and egocentric action recognition (e.g., MediaPipe, HaMeR, Ego4D ecosystem). Must be able to detect hand presence, occlusion, and interaction with objects in the center of the frame.
  • Action Recognition & Semantic Understanding: Strong experience with video understanding models (e.g., SlowFast, TimeSformer, VideoCLIP) to verify if the actor's actions match the assigned tasks.
  • Image Quality Assessment (IQA): Proficiency in classical and deep-learning-based IQA methods to automatically detect over/underexposure, motion blur, and excessive camera shake.
  • Privacy & Anonymization: Proven ability to build robust detection pipelines for faces, demographics (children), and PII (ID cards, passports, logos) using models like RetinaFace, YOLO, and OCR (Tesseract/EasyOCR) for automatic blurring/redaction.

Video Processing & Sensor Fusion:

  • Advanced Video Processing: Mastery of FFmpeg, OpenCV, and GStreamer for programmatic video manipulation, trimming, metadata extraction (FPS, bitrates, compression), and batch validation.
  • Sensor Data Handling (IMU): Experience working with time-series data and sensor fusion, specifically synchronizing head-mounted IMU telemetry (accelerometer/gyro) with video frames.
  • Multimodal Processing (Bonus): Familiarity with audio processing and Speech-to-Text (e.g., Whisper) to validate spoken commands in video audio tracks.

Engineering & Deployment:

  • Programming: 5+ years of software engineering experience with strong proficiency in Python and C++.
  • Frameworks: Expert in PyTorch; familiar with model optimization and inference acceleration using TensorRT, ONNX, or OpenVINO.
  • Cloud & MLOps: Experience deploying high-throughput CV pipelines on AWS (EC2 GPU instances, SageMaker, or EKS) and working with large-scale storage (S3).


BENEFITS:

  • Competitive compensation package based on experience and qualifications
  • Opportunity to build a strategic global data marketplace for robotics and AI training data from zero to one.
  • Work in a high-speed technology environment backed by Vingroup and VinDynamics leadership.
  • Competitive compensation package aligned with capability and business impact.
  • Clear ownership, measurable KPIs, and exposure to global partners, US platform models, and frontier robotics businesses.