About The Opportunity
On behalf of our clients, we are hiring a
GPU Engineer Team Leader
to support their high-performance AI framework engineering team. Our client is an innovative AI infrastructure pioneer specializing in GPU/NPU optimizations and advanced LLM inference frameworks at the intersection of HPC and AI systems.
Joining the team as a GPU Engineer Team Leader, you will guide a dedicated team of 4-5 engineers while remaining hands-on with high-performance GPU software development. You will play a pivotal role in driving technical delivery, conducting code reviews, and optimizing single- and multi-GPU architectures for state-of-the-art AI workloads.
Key Responsibilities
-
Team Leadership & Delivery: Lead, mentor, and develop a small team of GPU/HPC engineers. Plan sprints, manage workload distribution, and ensure the timely delivery of high-quality software components.
-
Mentorship & Culture: Conduct regular 1:1s, support engineers' career development, and foster a culture of continuous improvement and technical excellence.
-
Technical Strategy: Translate high-level technical goals from senior management into actionable tasks, surface blockers early, and maintain clear communication with stakeholders.
-
Hands-on Development: Write production-ready, low-level GPU kernel code using CUDA, HIP, or OpenCL for AI training and inference workloads.
-
Quality Assurance: Lead code reviews, enforce coding standards, and perform deep performance profiling and memory hierarchy optimizations to solve critical architectural challenges.
Requirements
Technical skills:
-
Bachelor's degree in Computer Science, Computer Engineering, or a related technical field.
-
2+ years of professional experience writing system software for GPUs.
-
Strong programming proficiency in C++ and Python.
-
Direct experience writing and optimizing GPU software using CUDA, HIP, or OpenCL.
-
Deep knowledge of GPU memory hierarchies, including shared memory utilization, registers, coalescing, and occupancy optimization.
-
Familiarity with deep learning frameworks (such as PyTorch or TensorFlow) and how they interact with underlying GPU hardware.
Nice-to-have (Preferred Qualifications)
-
Experience with distributed GPU computing, multi-GPU coordination, or parallel runtime systems.
-
Strong understanding of AI model architectures (e.g., attention mechanisms, matrix operations) and their impact on GPU workload design.
-
Hands-on experience with performance profiling tools such as Nsight Compute, Nsight Systems, or AMD ROCm profiler.
-
Active contributions to open-source GPU/HPC projects or publications at top-tier relevant conferences (PPoPP, HPDC, SC, MICRO, etc.).
Benefits
-
Competitive salary package with performance bonuses.
-
Premium healthcare insurance coverage.
-
Opportunity to work with cutting-edge HPC, multi-GPU systems, and generative AI infrastructure.
-
Clear career growth paths and continuous professional development support.
-
Dynamic, open, and technical engineering work environment.