Verified today
Site Reliability Engineer III
About the role
Innovaccer is seeking a Site Reliability Engineer III to build secured, modern healthcare cloud infrastructure and a massive data stack, with an emphasis on writing everything as code. Day-to-day work includes automating Infrastructure as Code across cost, reliability, scalability, and performance pillars; designing SRE domains; building CI/CD stacks with Dev and QA; optimizing production metrics and cost; enforcing CISO security guidelines; leading least-privilege RBAC; and driving disaster recovery and incident response. This is an on-site, full-time role based in Noida, Uttar Pradesh, India.
What you’ll do
- Build and automate secure cloud infrastructure (Infrastructure as Code) across pillars such as cost, reliability, scalability, and performance
- Design and architect various domains of SRE
- Build CI/CD stack collaborating across Dev and QA/Automation teams and drive the organization toward continuous delivery and deployment
- Collaborate with Dev and QA teams to increase adoption of DevOps practices and toolchains
- Apply analytical skills to understand production system metrics, drive change, optimize system utilization, and drive cost efficiency
- Ensure the platform is secured per CISO guidelines, including DDoS protection via WAF, vulnerability and patch management, and required security agents
- Lead least-privilege RBAC for various production services and toolchains
- Build and execute disaster recovery plans and serve as a key stakeholder in incident response
What you’ll bring
- 6–9 years in production engineering, site reliability, or related roles
- Solid hands-on experience with at least one cloud provider (AWS, Azure, or GCP) with an automation focus; certifications preferred
- Strong expertise in Kubernetes and Linux
- Proficiency in scripting/programming with Python required
- Knowledge of CI/CD pipelines and toolchains (Jenkins, ArgoCD, GitOps)
- Experience in production reliability, scalability, and performance systems
- Experience in 24x7 production environments with a process focus
- Security-first mindset with knowledge of vulnerability management and compliance
Nice to have
- Certifications preferred for cloud provider automation experience
- Familiarity with persistence stores (Postgres, MongoDB), data warehousing (Snowflake, Databricks), and messaging (Kafka)
- Exposure to monitoring/observability tools such as ElasticSearch, Prometheus, Jaeger, NewRelic, etc.
- Hands-on experience with Kafka, Postgres, and Snowflake is advantageous