Junior AI Engineer (Real-Time Translation Engine)

LinkedIn|KVY TECH CO LTD|Ho Chi Minh City, Vietnam|10 Aug 2026
Apply Now →

JOB DESCRIPTION

Company Description


KVY TECH CO LTD is a software engineering partner dedicated to helping startups and SMEs bring new software innovations to market quickly and with reduced risk. Since 2018, the company has provided full-cycle custom software development, acting as a long-term technology partner that manages end-to-end product development so clients can focus on business growth. KVY TECH builds scalable, high-quality solutions including MVPs, web and mobile applications, legacy system modernization, generative AI solutions, and custom ecommerce platforms. The team specializes in modern, in-demand technologies such as Ruby on Rails, Node.js, React, React Native, Nest.js, Typescript, and AWS. KVY TECH has a strong focus on software solutions for financial services, fintech, retail, edtech, and IoT, collaborating closely with clients to create impactful products.



Role Description

The Junior AI Engineer — Real-Time Translation Engine is a full-time, in-house position based in Ho Chi Minh City, Vietnam, in a hybrid arrangement with flexibility for some work from home.

We are building a real-time speech translation engine: speech in one language, translated speech out in another, live, under a strict latency budget. It is delivered to a client as an API and SDKs. The engine runs open-source speech recognition, machine translation, and text-to-speech models self-hosted on our own GPUs — this is applied ML systems engineering, not model research.

This role is designed as a structured apprenticeship into that domain. You will work directly alongside a senior speech engineer leading the current proof-of-concept, learning the pipeline hands-on and progressively taking ownership of it as an in-house engineer. We are hiring for fundamentals and trajectory rather than years of experience.

Day-to-day responsibilities include running and extending the benchmarking harness that measures per-stage and end-to-end latency across the pipeline; assembling and maintaining frozen audio and text test sets; evaluating model quality using the correct metric per stage (word error rate for speech recognition, chrF/COMET for translation, round-trip intelligibility for text-to-speech); profiling GPU memory and throughput; and reporting results as reproducible percentile figures with the full configuration that produced them.

As you grow into the role, responsibilities expand to model serving behind a swappable-model layer, streaming pipeline work (voice activity detection, partial versus final transcripts, sentence-boundary segmentation, and overlapping pipeline stages to cut end-to-end latency), the gRPC service boundary between the Python inference worker and the Go control plane, concurrent-stream load testing to establish capacity per GPU, and automated regression guards for latency and quality.

You will collaborate closely with the backend and platform engineers who own the API layer, and with the project owner on scope and priorities. Clear written communication in English, reproducible experiments, and well-documented results are core expectations of the role, not extras.



Qualifications

  • Strong programming fundamentals in Python , with the ability to read, debug, and extend an existing production codebase — not only write new scripts from scratch.
  • Solid computer science foundations — data structures, algorithms, and complexity — and the ability to reason about why a system is slow rather than guess.
  • Working knowledge of machine learning fundamentals , particularly model inference: what runs on a GPU, what numerical precision (fp16/int8) trades away, and roughly how transformer-based models work.
  • A measurement-first mindset. Comfort with basic statistics as applied to system evaluation — percentiles versus averages, sample size, and reproducibility. You should instinctively ask how a reported number was measured.
  • Comfort in a Linux environment — command line, SSH, Git, Docker, and log analysis.
  • Exposure to speech, audio, or NLP work through professional experience, academic projects, or personal projects — for example Whisper or another speech-to-text model, a text-to-speech model, machine translation, or general audio processing. Depth is not expected; genuine hands-on curiosity is.
  • Fluent reading comprehension of English technical material — model cards, documentation, licenses, research papers, and GitHub issues.
  • Bachelor's degree in Computer Science, Data Science, Mathematics, or a related technical field, or equivalent practical experience.
  • Clear communication and documentation habits , and the willingness to ask questions early rather than spend days blocked in silence.


Advantageous, but not required:

  • Experience with GPU tooling and profiling (CUDA, nvidia-smi, VRAM budgeting)
  • Asynchronous Python, WebSockets, or other streaming/real-time protocols
  • Familiarity with PyTorch, Hugging Face, or self-hosted model serving
  • Docker, Kubernetes, CI/CD pipelines, or observability tooling
  • Any exposure to Go or TypeScript, as the API layer and SDKs are built in those languages


Why KVY TECH:

  • Hybrid working environment (Remote/Onsite)
  • Competitive compensation
  • Laptop and equipment available on request
  • 15 days off annually
  • Birthday off with a company gift
  • Continuing education budget so you can keep learning outside of your day-to-day job
  • Work with an exciting and growing company with lots of interesting technical problems