Site Reliability Engineer, EDA Engineering Compute
Job Description
Synopsys is expanding Synopsys.ai alongside traditional EDA engineering compute for semiconductor customers in Vietnam. We are looking for an experienced Unix/Linux infrastructure professional to keep these platforms reliable, secure, and operable in Synopsys sites and at major customer environments (including customer-hosted, air-gapped, and high-security on-prem deployments).
You will join the regional EDA Engineering Compute / platform operations team working with global IT, AE engineering, and product teams. Your work directly enables R&D and customer engineering teams to run EDA and GenAI workloads—from HPC batch grids through containerized AI gateways, vector data services, and agent/MCP-based tool execution on the grid.
The successful candidate will support Synopsys Vietnam engineering compute and data center services, and partner on in-country deployment and sustainment of the related platform components for key semiconductor accounts.
How This Role Contributes
-
Improves availability, operability, and utilization of business-critical EDA and Synopsys.ai infrastructure in Vietnam.
-
Reduces operational risk for customer air-gapped and customer-hosted deployments through strong runbooks, monitoring, and incident practices.
-
Connects customer HPC schedulers, shared storage, GPU compute, and platform gateways so Synopsys tools and GenAI services run predictably at scale.
-
Extends APAC shared-support coverage with Japan, Korea, Taiwan, and China peers.
Responsibilities
Engineering compute and data center
-
Maintain Synopsys Vietnam engineering compute and data center environments per corporate IT, data center, and security standards.
-
Support server hardware lifecycle activities: installation, provisioning, maintenance, upgrade planning, retirement, and decommissioning.
-
perform capacity planning, performance troubleshooting, and vendor/customer coordination at customer sites.
-
Operate and troubleshoot Linux-based EDA compute: virtualization, engineering resource management, job schedulers, license connectivity, and remote access services.
-
Run HPC operations: cluster health, scheduler integration (LSF, Slurm, or similar), InfiniBand where deployed, and performance-related incidents.
-
Collaborate with corporate network and information security on data center connectivity, switching, firewalls, routing, and circuits.
Synopsys.ai platform operations (customer and Synopsys environments)
-
Support deployment and day-2 operations for Synopsys EDA and AI reference patterns in Vietnam: customer-hosted and air-gapped container environments, API/AI gateways, LLM gateway integration (cloud or on-prem inference endpoints per customer policy), and customer-controlled vector database/GPU compute where applicable.
-
Partner on install, upgrade, and configuration of platform components that connect user/agent workflows to the engineering grid—e.g., job services, unified MCP server patterns, and tool execution on LSF/Slurm-backed clusters (in coordination with products and global platform teams).
-
Implement and maintain observability, alerting, authentication/authorization integration (e.g., customer IdP, OIDC/SSO/OAuth patterns), and operational documentation for platform services.
-
Support agentic AI platform rollouts where components run locally, centrally, or on the grid; escalate design questions to global architecture and engineering teams.
Reliability, improvement, and customer engagement
-
Own or co-own runbooks, monitoring/alerting policies, and automation (scripts, IaC, or equivalent) to reduce toil across the APAC team.
-
Apply AI-assisted workflows to improve troubleshooting, knowledge management, and routine operations.
-
On-call and incident response: shared rotation (weekday nights, weekends, and holidays); lead or support incident bridges; produce post-incident reports in English.
-
Customer air-gapped and on-site support: planned upgrades, break-fix, and customer training for isolated HPC and on-prem AI sites. Travel within Vietnam and APAC as needed (typically 10–20% of time).
-
Cross-regional collaboration: handoffs, documentation, and mentor junior staff where applicable.
Required Skills And Experience
Technical
-
5+ years in Linux system administration, infrastructure engineering, or production network/storage support.
-
Enterprise data center experience: servers, storage, networking, monitoring, and operational processes.
-
HPC or EDA engineering compute: Linux clusters, workload schedulers (LSF, Slurm, or similar), and production support.
-
Ability to support GPU-enabled compute and containerized platform services in customer-controlled environments.
-
Familiarity with virtualization, remote access, observability stacks, and infrastructure automation.
-
Understanding of secure multi-tier deployments: gateways, service-to-service auth, and operational telemetry (logs, metrics, tracing) for distributed platform components.