JOB DESCRIPTION
Company Description
KVY TECH CO LTD is a software engineering partner dedicated to helping startups and SMEs bring new software innovations to market quickly and with reduced risk. Since 2018, the company has provided full-cycle custom software development, acting as a long-term technology partner that manages end-to-end product development so clients can focus on business growth. KVY TECH builds scalable, high-quality solutions including MVPs, web and mobile applications, legacy system modernization, generative AI solutions, and custom ecommerce platforms. The team specializes in modern, in-demand technologies such as Ruby on Rails, Node.js, React, React Native, Nest.js, Typescript, and AWS. KVY TECH has a strong focus on software solutions for financial services, fintech, retail, edtech, and IoT, collaborating closely with clients to create impactful products.
Role Description
The Junior AI Engineer — Real-Time Translation Engine is a full-time, in-house position based in Ho Chi Minh City, Vietnam, in a hybrid arrangement with flexibility for some work from home.
We are building a real-time speech translation engine: speech in one language, translated speech out in another, live, under a strict latency budget. It is delivered to a client as an API and SDKs. The engine runs open-source speech recognition, machine translation, and text-to-speech models self-hosted on our own GPUs — this is applied ML systems engineering, not model research.
This role is designed as a structured apprenticeship into that domain. You will work directly alongside a senior speech engineer leading the current proof-of-concept, learning the pipeline hands-on and progressively taking ownership of it as an in-house engineer. We are hiring for fundamentals and trajectory rather than years of experience.
Day-to-day responsibilities include running and extending the benchmarking harness that measures per-stage and end-to-end latency across the pipeline; assembling and maintaining frozen audio and text test sets; evaluating model quality using the correct metric per stage (word error rate for speech recognition, chrF/COMET for translation, round-trip intelligibility for text-to-speech); profiling GPU memory and throughput; and reporting results as reproducible percentile figures with the full configuration that produced them.
As you grow into the role, responsibilities expand to model serving behind a swappable-model layer, streaming pipeline work (voice activity detection, partial versus final transcripts, sentence-boundary segmentation, and overlapping pipeline stages to cut end-to-end latency), the gRPC service boundary between the Python inference worker and the Go control plane, concurrent-stream load testing to establish capacity per GPU, and automated regression guards for latency and quality.
You will collaborate closely with the backend and platform engineers who own the API layer, and with the project owner on scope and priorities. Clear written communication in English, reproducible experiments, and well-documented results are core expectations of the role, not extras.
Qualifications
Advantageous, but not required:
Why KVY TECH: