Verified today
ML Infrastructure Service Reliability Engineer
About the role
ML Infrastructure team at Apple Services Engineering manages the company’s largest ML compute platform and multi‑cloud storage abstraction, enabling critical training workloads for user‑facing features. As a Site Reliability Engineer you will ensure high availability and efficiency of the stack—from nodes to network—by automating operations, troubleshooting complex cloud networking, and advancing Kubernetes‑based services. Based in India, on‑site work model.
What you’ll do
- Design, implement and maintain ML infrastructure services ensuring high availability
- Automate operational tasks and improve efficiency through tooling
- Troubleshoot and resolve complex networking and storage issues
- Develop code in Python/Go/Rust for platform enhancements
- Collaborate with cross‑functional teams to align technical direction
- Monitor and support services using observability tools
- Contribute to architecture decisions for scalable distributed systems
What you’ll bring
- 5+ years building, operating and scaling large applications in private, public or hybrid cloud
- Deep expertise in Kubernetes (GKE or EKS)
- Proficiency in Python, Go or Rust development
- Experience with object storage (Amazon S3 or GCS)
- Strong background in complex networking in cloud environments
- Solid understanding of Linux internals and distributed systems
- Hands‑on experience with configuration management tools (Spinnaker, Helm, Flux)