Verified today
Manager, Site Reliability Engineering
About the role
Manager, Site Reliability Engineering at LSEG leads a team focused on maintaining and improving system reliability across financial market infrastructure. Responsibilities include defining and monitoring Service Level Objectives, measuring availability, latency, and overall health, and automating scaling and recovery processes. The role partners with development teams to enhance reliability, observability, and release velocity, and participates in on-call rotations, incident response, postmortems, and root cause analysis. Advocating strong engineering practices, the manager drives scalable, reliable, and performant services. Key enablers for cloud migration involve architectural reviews, operational testing, and configuring Datadog dashboards and metrics. The position requires a Bachelor's degree in computer science, experience with Java, C#, Python, Go, Unix/Linux, Windows, and cloud platforms Azure, AWS, or GCP. Preferred qualifications include 12+ years in the industry, DevOps experience, algorithms, data structures, observability practices, Infrastructure as Code, IAM, and application security. Tools such as Datadog, BigPanda, Terraform, and EntraID are used, with openness to other tools. The manager works onsite in Bengaluru, India, and benefits include healthcare, retirement planning, paid volunteering days, and wellbeing initiatives.
What you’ll do
- Maintain Service Level Objectives for the systems they own.
- Constantly measuring and improving availability, latency, and overall system health.
- Write automation to scale systems sustainably, prevent service issues, or when they occur, quickly recover service.
- Partner with development teams to improve system reliability, observability, and release velocity.
- Participate in on-call rotations, incident response, postmortems, and root cause analysis and resolution.
- Be a vocal advocate of strong/sound engineering practices that allow us to build, deploy, and run scalable, reliable, and performant services.
- Enable cloud migration by working with foundation and migration teams from inception of the projects, performing architectural reviews, operational acceptable testing and configuring Datadog dashboards and metrics.
- Be part of continuous learning and development culture.
What you’ll bring
- Bachelor's degree in computer science, a related technical field involving software/systems engineering, or equivalent practical experience.
- Experience with Object Oriented programming languages such as: Java, C#, Python, or Go.
- Experience with Unix/Linux and Windows operating systems.
- Hands on Experience with one of the following cloud platforms: Azure, AWS, or GCP.
Nice to have
- Minimum 12 years in the industry.
- Experience on DevOps concepts and way of working.
- Experience with algorithms and data structures.
- Experience in Observability practices with logging, metrics, tracing, and alerting.
- Experience with Infrastructure as Code.
- Understanding of identity and access management, and application security.
- We use Datadog and BigPanda for our observability stack, Terraform for our cloud infrastructure, and EntraID as our IAM solutions but we’re very open to incorporating your experience with any other tools.
Skills
Benefits
- Healthcare
- Retirement planning
- Paid volunteering days
- Wellbeing initiatives
Education
Bachelor's