Active listing
Senior Platform and EngOps Engineer
About the role
NVIDIA is seeking a Senior Platform and EngOps Engineer to develop and maintain GPU clusters using NVLink and InfiniBand. The role involves automating deployment, provisioning, and maintenance of large GPU clusters, implementing DevOps tools for software updates, monitoring cluster availability, and troubleshooting failures. Responsibilities include managing rollout and rollback of cluster software and firmware, collaborating with engineering and product teams across time zones, and ensuring optimal cluster performance. Candidates should hold a BS or MS in Computer Science, Engineering, or related field, with 5+ years of experience in cluster administration, automation using Ansible, Python, and Shell Scripting, and deep knowledge of operating systems, networking, and high-performance applications. Preferred experience includes resource scheduling managers like Slurm, alerting tools, GPU-focused hardware such as DGX systems, metrics collection, and large-scale networking design.
What you’ll do
- Develop automated tools to efficiently deploy, provision, and maintain extensive GPU clusters interconnected via NVLink and InfiniBand.
- Implement modern DevOps tools to automate software updates, perform maintenance tasks, and monitor cluster availability, ensuring seamless operations.
- Take ownership of daily cluster failures and issues, troubleshooting them promptly to maintain optimal cluster availability and performance.
- Manage the rollout and rollback of cluster software and firmware updates, ensuring smooth transitions and minimal disruptions.
- Collaborate effectively with dynamic Engineering and Product Teams across multiple time zones to align cluster operations with evolving project requirements.
What you’ll bring
- BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of hands-on experience in deploying and administrating clusters, servers, switches, and related infrastructure.
- Automation expert with hands on skills in Ansible, Python and Shell Scripting.
- Deep understanding of operating systems, computer networks, and high-performance applications.
- Proven ability to work effectively with developers and test engineers across different teams and time zones.
- Proficient with Linux fundamentals.
Nice to have
- Familiarity with resource scheduling managers, preferably Slurm.
- Direct experience with industry standard alerting tools and emergency response practices.
- Hands-on experience with GPU-focused hardware and software, such as DGX systems and Compute Clusters.
- Proficiency in crafting and implementing a robust metrics collection and alerting infrastructure.
- Proficiency in designing large scale networking technologies and the associated challenges.
Skills
Education
Bachelor's or Master's