Active listing
Staff Engineer, Reliability
About the role
AI Operations & Site Reliability team at The Hartford builds and maintains enterprise AI platforms, ensuring LLM services, RAG pipelines, and inference endpoints run reliably. The Senior Staff Engineer designs high‑availability architectures, defines SLOs/SLIs, and creates observability stacks to monitor AI‑specific metrics like hallucination rates and token throughput. Works from Hyderabad, India on an on‑site model, collaborating with AI engineers to automate deployments and drive continuous reliability improvements.
What you’ll do
- Own end‑to‑end reliability of production AI systems
- Define and maintain SLOs/SLIs/SLAs for AI services
- Design high‑availability architectures and disaster recovery procedures
- Build observability stacks and real‑time dashboards for AI metrics
- Lead incident response and maintain runbooks for AI failures
- Automate deployment pipelines and infrastructure as code
What you’ll bring
- 8+ years software engineering, DevOps, or SRE experience
- 1+ year operating AI/ML systems in production
- Strong proficiency in Python and a systems language
- Advanced experience with GCP, AWS, Kubernetes, Terraform
- Experience with observability tools like Prometheus, Grafana, Datadog
- Knowledge of AI failure modes and LLM provider platforms
Skills
Education
Bachelor's degree in Computer Science or related field