Verified today
Senior Lead Systems Operations Engineer
About the role
Wells Fargo is seeking a Senior Lead Systems Operations Engineer to provide deep technical leadership in Platform Reliability Engineering (PRE) within the Bengaluru, India location. The role involves acting as an advisor to senior leadership, leading complex initiatives for platform support solutions, and translating business objectives into technical engineering solutions. Day-to-day work includes defining SRE practices, driving observability standards, designing automation-first solutions, and mentoring team members on reliability and operations.
What you’ll do
- Act as a Platform Reliability Engineering (PRE) subject matter expert, providing deep technical leadership in one core domain
- Lead analysis and resolution of complex, systemic production reliability issues, translating recurring incidents into long-term engineering solutions
- Apply SRE principles including SLIs, SLOs, error budgets, and incident-driven engineering improvements to both new and legacy platforms
- Define and drive enterprise observability standards, including metrics, logs, traces, alerting, and service health dashboards
- Design and implement automation-first solutions to reduce operational toil, improve MTTR, and enable self-healing and self-service
- Partner with application, infrastructure, cloud, and support teams to improve availability, performance, capacity, and resiliency
- Lead or contribute to blameless post-mortems, ensuring measurable and sustained reduction of repeat incidents
- Translate complex technical and operational risks into clear, data-driven guidance for senior leadership
What you’ll bring
- 7+ years of Systems Engineering, Technology Architecture experience, or equivalent
- 7+ years of experience in Systems Operations, SRE, Platform Engineering, or Production Support with deep expertise in at least one platform domain
- Strong hands-on experience applying SRE practices, including SLI/SLO definition, error budgets, and reliability metrics
- Proven experience troubleshooting and resolving large-scale, distributed production systems
- Hands-on experience with observability and monitoring tools such as Grafana, Splunk, Prometheus, Cribl, ThousandEyes, AppDynamics, or equivalent
- Strong scripting and automation skills using Python, Bash, and/or PowerShell
- Solid understanding of incident, problem, and change management in enterprise production environments
- Strong communication and influencing skills across engineering teams and senior leadership
Nice to have
- Experience with capacity management, performance engineering, and resiliency design (HA, fault tolerance, RTO/RPO)
- Experience operating in hybrid environments (on-prem + cloud) with complex enterprise dependencies
- Familiarity with infrastructure automation / IaC tools such as Ansible or Terraform
- Ability to drive technical debt remediation for critical legacy platforms using structured backlogs
- Experience mentoring or leading senior engineers in reliability, operations, or SRE-focused roles
- Prior project or initiative leadership experience is highly desirable
Skills
Education
7+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education