About the company
: Tensormesh is building the next generation of AI inference infrastructure.
Our mission is to make large language models faster, cheaper, and easier to deploy across any
environment — cloud, on-prem, or hybrid. We help enterprises and AI teams optimize GPU
utilization and scale inference workloads with up to 10× better performance.
1. What You'll Own
-
Pipeline architecture: GitHub Actions workflows + self-hosted GPU runner fleet; multi-stage pipeline from lint → unit → GPU integration → cross-framework compatible (vLLM/SGLang) → performance regression
-
Release engineering: semantic versioning, PyPI publishing, multi-arch container images, Helm charts, Sigstore/cosign signing, coordination with downstream integrators
-
Performance gates: Continuous benchmarking that blocks regressions in cache hit rate, TTFT, throughput, memory before merge
-
Contributor experience: fast PR feedback, eliminate flakiness, dev containers that don't require expensive GPUs
-
Security & IaC: SBOM/SLSA provenance, secret rotation, runner fleet via Terraform with cost-optimized autoscaling.
2. Required
-
4+ years MLOps/DevOps/SRE; 2+ years CI/CD for GPU or ML workloads
-
Deep GitHub Actions expertise (workflows, composite actions, self-hosted runners at scale)
-
Python packaging & PyPI release flow (incl. wheels with native extensions)
-
Docker multi-stage/multi-arch; NVIDIA Container Toolkit
-
Terraform/Ansible for cloud GPU infrastructure
-
Track record building CI that contributors trust — fast, non-flaky, clear failures
3. Strongly Preferred
-
Maintainer/contributor experience on a popular OSS project
-
Familiarity with vLLM, SGLang, NVIDIA Dynamo, KServe, or Triton
-
Kubernetes in CI (Kind/k3s, multi-node integration tests)
-
Continuous benchmarking tools + time-series perf tracking
-
Supply chain security (Sigstore, SLSA, syft/grype)
-
RDMA / high-perf networking / P2P system testing