Back to jobs
M
Hiring companyMicrosoft

Verified 2 days ago

Senior Site Reliability Engineer

Hyderabad, Telangana, India Full-time On-site

About the role

Azure AI Infrastructure team at Microsoft ensures the reliability, scalability, and security of AI workloads across cloud and HPC environments. The Senior Site Reliability Engineer leads incident response, performance optimization, and automation of containerized services, collaborating with cross‑functional engineers to drive innovation. Hyderabad, Bengaluru, Noida, India, on-site.

What you’ll do

  • Lead incident response and root cause analysis for AI workloads
  • Identify and resolve performance bottlenecks in compute, storage, networking, and GPUs
  • Develop and maintain automation tools for deployment and monitoring
  • Provide technical guidance on cloud and AI infrastructure technologies
  • Advise customers on service excellence and reliability

What you’ll bring

  • Master's Degree in Computer Science or related field
  • Bachelor's Degree in Computer Science or related field
  • 12+ years professional software engineering experience
  • 8+ years service operations and reliability
  • 5+ years AI or cloud platform infrastructure experience
  • 1+ years incident management in cloud/AI environments

Skills

KubernetesDockerAzureGPUInfiniBandAutomationIncident ManagementPerformance Optimization