GPU Engineer Team Leader

LinkedIn|Sunbytes|Hanoi, Hanoi, Vietnam|4 Aug 2026
Apply Now →

JOB DESCRIPTION

About The Opportunity

On behalf of our clients, we are hiring a GPU Engineer Team Leader to support their high-performance AI framework engineering team. Our client is an innovative AI infrastructure pioneer specializing in GPU/NPU optimizations and advanced LLM inference frameworks at the intersection of HPC and AI systems.

Joining the team as a GPU Engineer Team Leader, you will guide a dedicated team of 4-5 engineers while remaining hands-on with high-performance GPU software development. You will play a pivotal role in driving technical delivery, conducting code reviews, and optimizing single- and multi-GPU architectures for state-of-the-art AI workloads.

Key Responsibilities

  • Team Leadership & Delivery: Lead, mentor, and develop a small team of GPU/HPC engineers. Plan sprints, manage workload distribution, and ensure the timely delivery of high-quality software components.
  • Mentorship & Culture: Conduct regular 1:1s, support engineers' career development, and foster a culture of continuous improvement and technical excellence.
  • Technical Strategy: Translate high-level technical goals from senior management into actionable tasks, surface blockers early, and maintain clear communication with stakeholders.
  • Hands-on Development: Write production-ready, low-level GPU kernel code using CUDA, HIP, or OpenCL for AI training and inference workloads.
  • Quality Assurance: Lead code reviews, enforce coding standards, and perform deep performance profiling and memory hierarchy optimizations to solve critical architectural challenges.

Requirements

Technical skills:

  • Bachelor's degree in Computer Science, Computer Engineering, or a related technical field.
  • 2+ years of professional experience writing system software for GPUs.
  • Strong programming proficiency in C++ and Python.
  • Direct experience writing and optimizing GPU software using CUDA, HIP, or OpenCL.
  • Deep knowledge of GPU memory hierarchies, including shared memory utilization, registers, coalescing, and occupancy optimization.
  • Familiarity with deep learning frameworks (such as PyTorch or TensorFlow) and how they interact with underlying GPU hardware.

Nice-to-have (Preferred Qualifications)

  • Experience with distributed GPU computing, multi-GPU coordination, or parallel runtime systems.
  • Strong understanding of AI model architectures (e.g., attention mechanisms, matrix operations) and their impact on GPU workload design.
  • Hands-on experience with performance profiling tools such as Nsight Compute, Nsight Systems, or AMD ROCm profiler.
  • Active contributions to open-source GPU/HPC projects or publications at top-tier relevant conferences (PPoPP, HPDC, SC, MICRO, etc.).

Benefits

  • Competitive salary package with performance bonuses.
  • Premium healthcare insurance coverage.
  • Opportunity to work with cutting-edge HPC, multi-GPU systems, and generative AI infrastructure.
  • Clear career growth paths and continuous professional development support.
  • Dynamic, open, and technical engineering work environment.