Verified today
Site Reliability Engineer II
About the role
Backblaze is the object storage leader in the open cloud movement, and its Site Reliability Engineer II (SRE II) role helps ensure the stability, scalability, and reliability of customer-facing services and infrastructure. Day-to-day work centers on building automation, maintaining observability, supporting incident response and on-call rotations, and collaborating with engineering, product, and operations teams to embed reliability practices. This is a Remote - Bangalore position.
What you’ll do
- Support the availability and durability of critical services across production environments
- Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk
- Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements
- Follow established ITIL/OSS processes (incident, change, problem, and capacity management)
- Develop automation for common operational tasks to reduce manual intervention and toil
- Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK)
- Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins)
- Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency
What you’ll bring
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience)
- 2–4 years of experience in site reliability, systems engineering, or operations
- Solid Linux systems administration and troubleshooting skills
- Familiarity with monitoring, alerting, incident response, and root cause analysis
- Proficiency in at least one scripting language (Python, Bash, or Go)
- Understanding of containers (Kubernetes, Docker) and microservices concepts
- Knowledge of incident response and operational best practices
- Exposure to large-scale, production-grade systems
Nice to have
- Experience in a SaaS, service provider, or distributed systems environment
- Familiarity with ITIL/OSS practices and SLO/SLA’s
- Strong problem-solving skills and willingness to learn new technologies
- Experience with cloud platforms (AWS, GCP, or Azure)
- Ability to work independently, take ownership, and drive projects from problem discovery through resolution
Skills
Education
Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience).