Verified today
Staff Site Reliability Engineer
About the role
AlphaSense seeks a Staff Site Reliability Engineer to architect reliability platforms, lead AI‑driven initiatives, embed SRE practices, command incidents, deliver observability, and mentor teams.
What you’ll do
- Architect core reliability platforms and self‑service tooling
- Lead AI‑driven reliability initiatives and AIOps strategy
- Embed SRE practices across engineering via design reviews and production readiness
- Act as Incident Commander during critical events, ensuring blameless postmortems
- Deliver end‑to‑end monitoring, tracing, and profiling to optimize performance
- Mentor and multiply engineering talent across SRE and product teams
What you’ll bring
- 8+ years in SRE/DevOps with senior-level experience
- Proven track record running production SaaS systems at scale
- Strong programming skills in Python, Go, or similar
- Hands‑on experience with AWS, GCP, Azure and Kubernetes
- Deep knowledge of networking fundamentals (TCP/IP, DNS, HTTP/S)
- Expertise in monitoring, alerting, and observability tools (Prometheus, Grafana, Datadog, ELK, OTEL)
- Incident management experience leading high‑severity incidents and post‑mortems
- Excellent troubleshooting, communication, and collaboration skills