Verified today
Senior Site Reliability Engineer
About the role
Employ is transforming talent acquisition with its ATS solutions (Jobvite, Lever, JazzHR) and AI Companions, helping over 26,000 global customers hire smarter and at scale. The Senior Site Reliability Engineer will build highly reliable, scalable systems by blending software engineering and infrastructure expertise, maintaining production health, implementing automation, and working cross-functionally with developers and security teams. Based in Bangalore, India, this full-time role operates on a hybrid work model.
What you’ll do
- Use AI as a core tool in daily SRE work, embedding architectural context and change history directly into code and systems so they remain legible to both engineers and AI agents over time
- Deep dive into application codebases and directly contribute to improvements that drive the SRE mission, leaving code better than you found with each issue you tackle
- Participate in on-call rotations and command incident response with effective practices that improve time to recovery, preservation of evidence, and root cause analysis that prevents future occurrence
- Own the design and effectiveness of observability systems that empower all engineers with alerting and visibility into the applications we are supporting
- Develop and manage sustainable Infrastructure as Code (IaC) automation using tools such as Terraform, Ansible, or similar along with CI/CD and orchestrators such as Kubernetes and ArgoCD to bake SRE i
- Partner with roadmap delivery teams to implement and promote Site Reliability Engineering best practices within their workflows such as the systematic implementation of SLIs/SLOs, production readiness
- Collaborate with Security teams to ensure systems align with ISO 27001, SOC 2, and other compliance standards
- Stay current with industry trends and emerging technologies to continually improve our SRE capabilities
What you’ll bring
- 5+ years of experience in Site Reliability Engineering, Software Engineer, or a similar role
- Proficiency in one or more programming/scripting languages such as Typescript, Java, Python, Go, PHP, or Ruby
- Strong experience with Unix/Linux systems administration and internals
- Solid understanding of system design, distributed computing, and SRE principles
- Expertise with containerization and orchestration technologies such as Docker and Kubernetes
- Experience with one or more cloud platforms: AWS, Azure, or Google Cloud Platform (GCP)
- Hands-on experience with CI/CD tools such as GitHub Actions CI/CD, Argo CD, etc.
- Experience with relational databases like PostgreSQL, MySQL, or SQL Server
Nice to have
- Proficiency in scripting and automation using tools such as NodeJS
- Familiarity with NoSQL solutions such as MongoDB, Redis, DynamoDB
- Ability to monitor, optimize, and troubleshoot performance in large scale high-availability environments
- Experience with backup, replication, and data recovery strategies
- Excellent communication and collaboration abilities