Back to jobs
I
Hiring companyInnovaccer

Verified today

Site Reliability Engineer III

Noida, Uttar Pradesh, India Full-time On-site

About the role

Innovaccer is seeking a Site Reliability Engineer III to build secured, modern healthcare cloud infrastructure and a massive data stack, with an emphasis on writing everything as code. Day-to-day work includes automating Infrastructure as Code across cost, reliability, scalability, and performance pillars; designing SRE domains; building CI/CD stacks with Dev and QA; optimizing production metrics and cost; enforcing CISO security guidelines; leading least-privilege RBAC; and driving disaster recovery and incident response. This is an on-site, full-time role based in Noida, Uttar Pradesh, India.

What you’ll do

  • Build and automate secure cloud infrastructure (Infrastructure as Code) across pillars such as cost, reliability, scalability, and performance
  • Design and architect various domains of SRE
  • Build CI/CD stack collaborating across Dev and QA/Automation teams and drive the organization toward continuous delivery and deployment
  • Collaborate with Dev and QA teams to increase adoption of DevOps practices and toolchains
  • Apply analytical skills to understand production system metrics, drive change, optimize system utilization, and drive cost efficiency
  • Ensure the platform is secured per CISO guidelines, including DDoS protection via WAF, vulnerability and patch management, and required security agents
  • Lead least-privilege RBAC for various production services and toolchains
  • Build and execute disaster recovery plans and serve as a key stakeholder in incident response

What you’ll bring

  • 6–9 years in production engineering, site reliability, or related roles
  • Solid hands-on experience with at least one cloud provider (AWS, Azure, or GCP) with an automation focus; certifications preferred
  • Strong expertise in Kubernetes and Linux
  • Proficiency in scripting/programming with Python required
  • Knowledge of CI/CD pipelines and toolchains (Jenkins, ArgoCD, GitOps)
  • Experience in production reliability, scalability, and performance systems
  • Experience in 24x7 production environments with a process focus
  • Security-first mindset with knowledge of vulnerability management and compliance

Nice to have

  • Certifications preferred for cloud provider automation experience
  • Familiarity with persistence stores (Postgres, MongoDB), data warehousing (Snowflake, Databricks), and messaging (Kafka)
  • Exposure to monitoring/observability tools such as ElasticSearch, Prometheus, Jaeger, NewRelic, etc.
  • Hands-on experience with Kafka, Postgres, and Snowflake is advantageous

Skills

PythonKubernetesLinuxAWSAzureGCPKafkaPostgres