Verified 5 days ago
Site Reliability Engineer (SRE)
About the role
The Site Reliability Engineer (SRE) role at Five9 combines software engineering and operations to ensure highly reliable, scalable systems. This position involves 50% software development and 50% operational work, focusing on automation, monitoring, and system reliability. The SRE collaborates with platform, application, and database teams to maintain service reliability and availability. Key responsibilities include designing dashboards, establishing SLIs and SLOs, building alerting systems, managing CI/CD pipelines, and ensuring security and compliance. The role also involves cost optimization, managing shared infrastructure, and participating in incident response. Five9 values a blameless culture, automation, data-driven decisions, and knowledge sharing.
What you’ll do
- Design and implement comprehensive dashboards for OS and application monitoring
- Establish and maintain SLIs, SLOs, and error budgets for the service
- Build alerting systems and performance monitoring to proactively resolve issues
- Participate in on-call rotations and lead incident response efforts
- Build and optimize CI/CD pipelines for speed and resilience
- Develop and maintain infrastructure using tools like Terraform and Ansible
- Automate system configuration and ensure consistency across environments
- Ensure security scanning systems are in place and review vulnerabilities
- Maintain proper authentication, authorization, and audit logging systems
- Ensure systems meet regulatory and industry standards
- Participate in security incident response and remediation efforts
- Monitor and optimize cloud resource usage and costs
- Analyze usage patterns and plan for future capacity needs
- Provide recommendations for cost-effective architecture and resource allocation
- Implement automated scaling and resource optimization strategies
- Build and maintain common services like notification systems and caching layers
- Manage database reliability, performance, and scaling
- Implement and maintain service discovery, load balancing, and network policies
- Create and maintain tools that improve developer productivity and reliability
What you’ll bring
- Proficient in Python, Shell, Java, NodeJS, or similar languages
- Experience with AWS, GCP, or Azure cloud platforms
- Hands-on experience with Docker, Kubernetes, and container orchestration
- Knowledge of Prometheus, Grafana, ELK stack, or similar monitoring tools
- Proficiency with Ansible, Terraform, Helm, or similar infrastructure tools
- Expert-level Git usage and collaborative development practices
- Experience with GitLab CI/CD, GitHub Actions, or similar CI/CD pipelines
- Experience defining and maintaining SLOs and SLIs
- Understanding and implementation of error budget policies
- Proven track record in toil reduction and automation
- Experience with capacity planning and performance testing
Nice to have
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience
- Experience with microservices and distributed systems
- Knowledge of security best practices and compliance frameworks
- Experience with chaos engineering and reliability testing
- Prior experience in an SRE or DevOps role at a tech company
- Contributions to open-source projects or technical communities