Verified 3 days ago
Lead Systems Operations Engineer
About the role
Lead Systems Operations Engineer on the Platform Reliability Engineering (PRE) team at Wells Fargo, driving reliability, resiliency, and observability across critical platform services. Day‑to‑day responsibilities include leading complex initiatives, consulting on technical changes, and mentoring senior engineers to improve platform stability and scalability. Bengaluru, India, on‑site with hybrid support for on‑prem and cloud environments.
What you’ll do
- Lead complex, high‑impact initiatives and provide systems consultation to technology teams
- Plan and design large‑scale computer systems and network infrastructure
- Analyze and resolve escalated support issues for core business solutions
- Make decisions on technical changes and enhancements
- Consult on change design with engineering teams
- Collaborate with peers and managers to resolve systems support issues
- Drive technical debt remediation for critical legacy platforms
- Mentor and lead senior engineers in reliability and operations
What you’ll bring
- 5+ years Systems Engineering or Architecture experience
- Deep expertise in at least one platform domain (Database, Cloud, Network, Compute/Storage, Middleware, or Enterprise Application Support)
Nice to have
- Strong hands‑on SRE practices (SLI/SLO, error budgets, reliability metrics)
- Proven troubleshooting of large‑scale distributed production systems
- Experience with observability tools (Grafana, Splunk, Prometheus, Cribl, ThousandEyes, AppDynamics)
- Scripting and automation skills (Python, Bash, PowerShell)
- Infrastructure automation/IaC tools (Ansible, Terraform)
- Incident, problem, and change management in enterprise environments
- Capacity and performance engineering (HA, fault tolerance, RTO/RPO)
- Mentoring senior engineers in reliability and SRE roles