Verified today
Staff Site Reliability Engineer
About the role
Site Reliability Engineering team at AlphaSense builds and operates reliability platforms to ensure 99.99% uptime for its AI-driven market intelligence services. The Staff Site Reliability Engineer architects core reliability frameworks, leads incident response, and drives AIOps automation across global engineering. Pune, hybrid
What you’ll do
- Architect reliability platforms and self‑service tooling
- Lead AI‑driven AIOps initiatives for automated diagnostics
- Mentor engineers and champion SRE culture
- Act as Incident Commander for critical events
- Design and implement end‑to‑end observability solutions
- Drive production readiness reviews and operational standards
- Collaborate across engineering to embed reliability best practices
- Influence architectural decisions for scalability
What you’ll bring
- 8+ years SRE/DevOps experience
- 3+ years senior SRE role
- Proficiency in Python or Go
- Hands‑on AWS/GCP/Azure
- Kubernetes expertise
- Monitoring with Prometheus/Grafana/Datadog
- Incident management leadership
- Strong networking fundamentals
Nice to have
- 8+ years SRE/DevOps experience
- 3+ years senior SRE role
- Python or Go programming
- AWS/GCP/Azure cloud platforms
- Kubernetes orchestration
- Observability tools (Prometheus, Grafana, OTEL)
- Incident response and postmortem process
- Networking (TCP/IP, DNS, HTTP)