Back to jobs
TH
Hiring companyThe Hartford

Active listing

Staff Engineer, Reliability

Hyderabad, India Full-time On-site

About the role

AI Operations & Site Reliability team at The Hartford builds and maintains enterprise AI platforms, ensuring LLM services, RAG pipelines, and inference endpoints run reliably. The Senior Staff Engineer designs high‑availability architectures, defines SLOs/SLIs, and creates observability stacks to monitor AI‑specific metrics like hallucination rates and token throughput. Works from Hyderabad, India on an on‑site model, collaborating with AI engineers to automate deployments and drive continuous reliability improvements.

What you’ll do

  • Own end‑to‑end reliability of production AI systems
  • Define and maintain SLOs/SLIs/SLAs for AI services
  • Design high‑availability architectures and disaster recovery procedures
  • Build observability stacks and real‑time dashboards for AI metrics
  • Lead incident response and maintain runbooks for AI failures
  • Automate deployment pipelines and infrastructure as code

What you’ll bring

  • 8+ years software engineering, DevOps, or SRE experience
  • 1+ year operating AI/ML systems in production
  • Strong proficiency in Python and a systems language
  • Advanced experience with GCP, AWS, Kubernetes, Terraform
  • Experience with observability tools like Prometheus, Grafana, Datadog
  • Knowledge of AI failure modes and LLM provider platforms

Skills

Site Reliability EngineeringAI Platform OperationsObservabilityKubernetesTerraformPythonCloud (GCP/AWS)Incident Management

Education

Bachelor's degree in Computer Science or related field