Back to jobs
EL
Hiring companyEmergent Labs

Verified today

Data Scientist

Bengaluru

About the role

Emergent builds autonomous coding agents that generate, test, and deploy production applications from plain-language intent, with systems running at global scale. As a Data Scientist in Bengaluru, you own the end-to-end analytics loop across product, growth, and user behavior—turning agent trajectories, support tickets, logs, and user prompts into structured signal, building predictive and margin models, and driving rigorous experimentation. The role blends classical ML, LLM-native insight extraction, attribution, and strategic recommendations to inform decisions at scale.

What you’ll do

  • Turn agent trajectories, support tickets, logs, and user prompts into structured, queryable signal—summarize-then-embed-then-cluster pipelines (à la Anthropic's Clio / Braintrust Topics): distill each
  • Surface early indicators—of confusion, of a coming bug wave, of churn risk, of fraud—that no dashboard would ever surface on its own, and route them to the right team
  • Build predictive models that forecast conversion, retention, expansion, and churn, and embed those signals directly into product and growth workflows
  • Own marketing attribution and MMM: build the media-mix and incrementality models that tell us what's actually driving signups and paid conversions when per-user attribution is partial and, on mobile,
  • Own product and growth analytics across the self-serve funnel, web and mobile—activation, engagement, retention, conversion—and design and analyze A/B and growth tests with real rigor around power, no
  • Run clustering pipelines over hundreds of thousands of agent trajectories to discover the recurring kinds of things users try to build and the recurring ways builds fail, then hand product a taxonomy
  • Model the 'aha moment' for new users, including text-derived features from their first prompts and first agent interactions, and rebuild onboarding around the earliest signals of long-term retention
  • Build a gross-margin model that attributes LLM and compute cost down to the individual app and cohort, and tell product which segments are net-positive

What you’ll bring

  • 2 to 5 years in data science or applied ML with a focus on product analytics, growth, or user behavior
  • Strong SQL and real comfort working with large, event-level behavioral data at scale
  • Solid classical ML foundations—clustering (k-means, HDBSCAN, hierarchical), embeddings and vector similarity, dimensionality reduction (UMAP/PCA), classification—and knowing when each is the right too
  • Genuine skill at deriving insight from unstructured natural-language data—LLM traces, logs, tickets, free text—and turning it into predictive, queryable signal
  • Proficient in Python and the standard data-science stack (pandas, scikit-learn, statsmodels, numpy)
  • Data engineering competence—can design and ship ETL and data models (dbt or equivalent), not just query what already exists
  • Experienced designing and analyzing experiments—sample sizing, power, significance, novelty effects, interference between tests, and causal methods
  • Able to move fluidly between exploratory analysis, ML modeling, hypothesis testing, and crisp strategic recommendations

Nice to have

  • Marketing Mix Modeling (MMM), media attribution, or incrementality and geo-testing experience, especially in low-tracking or post-cookie environments
  • Experience at a PLG company with a self-serve funnel and freemium or usage-based / credit-based pricing
  • Modern data stack (BigQuery, dbt) and product analytics platforms (PostHog, Amplitude, Mixpanel, Segment)
  • Causal inference methods (difference-in-differences, synthetic control, propensity score matching)
  • Fraud, trust and safety, or abuse analytics
  • Working knowledge of embedding models and vector search, and the practical tradeoffs of running them at scale
  • Familiarity with the economics of AI/LLM products, including COGS modeling where compute is the dominant variable cost
  • You've built or contributed to AI-powered analytical tooling or novel measurement approaches

Skills

SQLPythonpandasscikit-learnstatsmodelsnumpyclusteringembeddings

Benefits

  • Daily Meals: Lunch and Dinner provided
  • Family Insurance: 3 Lakhs worth of coverage for you and your family
  • Unlimited Paid Time Off: Take the time you need to recharge and come back refreshed
  • Flexible Working Hours: Work arrangements that fit your life and commitments