SVS recruiting recruiting
← Back to jobs

AI Research Engineer

⌂ Emergent ▤ 5.0–8 years ◉ Bangalore ◎ ml-ds, ai-engineer

Research Engineer to define, measure and advance coding-agent quality — evals, benchmarks (SWE-bench Pro, Terminal-Bench), and post-training experiments (SFT, RLHF/DPO, reward modeling). 5–8 yrs AI experience.

०१ At a glance
Location
Bangalore
Experience
5.0–8 years
Published
Aug 30, 2026
०२ Description

About Emergent

Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Since public launch, Emergent has crossed $100M ARR with 10M+ users across 190+ countries who have built 12M+ applications. Backed by Creaegis, Claypond, Sentinel Global, Khosla Ventures, SoftBank, Google, Lightspeed, Prosus, Together, and Y Combinator.

The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox.

The Role

We’re looking for a Research Engineer to characterize, measure, and advance the capabilities of our coding agents. You will turn ambiguous notions of “agent quality” into clear, defensible metrics that the team, leadership, and the field can rely on, and use those metrics to drive both incremental wins and moonshots in agent performance.

This is a deep-work role at the intersection of agent behavior, evaluation research, and applied training. You will define what good looks like for long-horizon coding agents, build the evaluation datasets and methodology that produce those signals, mine production data for failure modes most teams never see, and run targeted training, fine-tuning, RL, memory, and prompt-optimization experiments that translate research advances into shipped improvements.

What You’ll Do

  • Architect the next version of the Emergent agent: shape the core architecture and make the foundational design choices that define how the agent thinks, learns, and improves over time
  • Characterize agent behavior at depth: develop an evidence-grounded understanding of how the agent succeeds and fails across real-world usage, and convert that into rigorous, quantitative measurement
  • Design and ship evaluations across reasoning, planning, tool use, code correctness, long-horizon execution, security, and reliability: define the metric, build the dataset, validate against known signals, ship dashboards that make regressions impossible to miss
  • Drive step-function gains: take on the ambitious bets — 10-point leaps on hard capabilities, not incremental polish
  • Climb public benchmarks: move the needle on SWE-bench Pro, Terminal-Bench, and other industry-standard coding-agent benchmarks
  • Run training and post-training experiments: supervised fine-tuning, RLHF/RLAIF, DPO, distillation, reward modeling, prompt optimization, judge-model calibration — against production-grounded objectives
  • Own end-to-end: hypothesis → experiment design → execution → analysis → decision → rollout → post-launch measurement
  • Make hard calls in subjective systems: when a regression is real, when a win is noise, when a benchmark is overfit, when to ship despite mixed signals, when to kill a promising direction

Who You Are

  • 5 to 8 years of AI experience, with meaningful time spent either training and fine-tuning models, or designing rigorous evaluations and measurement systems for them. Both paths are equally valued
  • Hands-on with the modern AI stack and fluent in Python (Go a plus): training pipelines, eval harnesses, data processing, statistical analysis. Comfortable with transformers, RLHF/DPO/RL for agents, eval frameworks (Inspect, lm-eval-harness, or equivalent), prompt optimization, judge models, and agent frameworks
  • Take pride in numbers that move. You measure first, opine second. You can defend why a benchmark is the right benchmark, why a metric isn’t gameable, and why a result is statistically real
  • Comfortable in subjective, probabilistic systems: noise floors, confounds, distribution shift, judge bias, selection effects
  • Enjoy the long tail — sifting through large volumes of agent behavior to find the rare, hidden failure mode
  • Understand models like friends: intuitions about how a model will behave on a new task before running it, and you update when reality disagrees. You know what came out last week and which two-year-old paper is suddenly relevant again
  • Independent operator with leadership presence: scope your own work, push back on weak ideas (including your manager’s)
  • Ship fast without compromising rigor

Benefits and Perks

  • Daily meals: lunch and dinner provided
  • Family insurance: ₹5L coverage for you and your family
  • Unlimited paid time off
  • Flexible working hours
Ready when you are. Apply