Verified today
Applied AI Engineer
About the role
Dscout is building the most flexible and powerful UX research platform, trusted by top brands across finance, healthcare, consumer goods, and tech. As an Applied AI Engineer, you will own the production improvement loop for LLM-based agent systems—turning powerful models into reliable product features through evaluation, instrumentation, experimentation, and staged rollouts. You will design quality standards, investigate underperformance, ship behavior improvements, and build backend services that make agent behavior observable and steerable. This role is based in India and is remote.
What you’ll do
- Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes
- Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context
- Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring
- Investigate why agents underperform across context, knowledge, instructions, tools, routing, guardrails, or workflow design
- Design and ship targeted behavior improvements, including changes to prompting, context construction, decision logic, tool use, and human-review paths
- Build backend services, APIs, data models, and feedback pipelines that make agent behavior observable, steerable, and reproducible
- Run controlled experiments, production replays, or staged rollouts to measure whether changes improve quality and downstream business results
- Partner with Product, Data Science, and Sales to prioritize high-value problems and define customer and business success
What you’ll bring
- 2-5 years of software engineering experience with hands-on experience building or operating LLM-powered features or agents in production (not just prototypes or demos)
- Fluency with prompting and context engineering as an engineering discipline, iterating on prompts, context construction, and tool definitions like code
- Experience building or maintaining evaluation harnesses for AI systems, including offline eval sets, LLM-as-judge or human-in-the-loop scoring, and regression detection
- Comfort with non-determinism, reasoning about agent behavior across a distribution of production traffic rather than fixed test cases
- Experience running experiments (A/B, staged rollouts, production replay) to validate whether changes actually improved outcomes
- A track record of shipping features real users depended on and owning what happened after launch
- High-agency mindset comfortable investigating ambiguous underperformance problems across context, tools, routing, and workflow design without a fully-scoped ticket
- Comfort using AI coding tools (Cursor, Claude Code, Copilot, or similar) as a real part of the workflow
Nice to have
- Experience with voice or real-time conversational AI systems
- Familiarity with LLM observability/tracing tools (e.g., Braintrust, LangSmith, Datadog LLM Observability)
- Experience with agentic orchestration frameworks (LangChain/LangGraph or similar)
- Exposure to MCP-based tooling or agentic data workflows
Skills
Benefits
- Strong and competitive compensation package with a built-in bonus and equity program
- Progressive benefits package for employees and dependents, including flexible PTO, 15 company holidays, 12 weeks of paid parental leave, and 401k match
- Education stipend to support growth and development, and a remote work stipend