Research Engineer
About Zstate
Zstate builds domain intelligence and agentic systems with credentialed expert networks. We package training data and production agents for software engineering, healthcare, finance, and related verticals where judgment is the scarce resource.
About the role
Founding role. You build the benchmarks, evals, and RL environments Zstate ships to AI lab clients, covering tool use, agentic behaviour, and domain reasoning, plus the methodology that makes them defensible: capability taxonomies, rubric architecture, reward signals, and inter-annotator reliability studies. You take a benchmark from task design through expert authorship, scoring, and failure analysis, and you run models against it to find where they break. You work closely with domain experts authoring tasks, building tooling, dashboards, and pipelines that turn their work into a benchmark a lab can trust. Reports to the Head of Engineering (currently the Founder).
What you will do
- Benchmarking: design and build benchmarks task by task (source artifacts, capability buckets, weighting, and negative criteria) following Zstate's Gold Standard methodology, and keep taxonomy and scoring aligned with what each AI lab client is trying to measure.
- Evaluation runs: build the pipeline that runs models against a benchmark, scores rollouts, and turns results into a report a lab can act on: pass@k, mean reward, variance, and bucket-level breakdown.
- Failure analysis: read model rollouts closely to find where and how they fail, and turn recurring failure patterns into new negative criteria or rubric refinements.
- Rubrics and IRR: write and refine task rubrics with domain experts, build the LLM-judge grading pipeline against them, and run inter-annotator reliability and calibration sessions that keep scores defensible.
- RL environments and reward signals: design RL environments and reward signals for post-training, and build reward verification checks that catch reward hijacking and misaligned reward functions before a client sees them.
- Authoring funnel: track the expert authoring funnel and data quality end to end (applications, work samples, accepted tasks, IRR scores) and use it to decide where the annotation pipeline needs to change.
What we look for
- 3-8 years of experience in software or ML engineering, with a track record of shipping production systems.
- Applied experience in model evaluation, benchmarking, or post-training, with strong software or ML engineering fundamentals.
- Strong coding skills, comfortable working with ML models and writing evaluation code.
- Solid backend fundamentals: APIs, databases, and cloud infrastructure for running and storing evaluation results at scale.
- Judgment to read evaluation results and model behaviour and draw the right conclusions about what they mean.
- Comfort designing rubrics and scoring criteria that hold up when an AI lab pushes back on them.
- High ownership, strong execution, and comfort with ambiguity in a high-iteration environment.
Nice to have
- Experience on a team that builds evals, benchmarks, or post-training data for AI labs.
- Experience designing synthetic scenarios or RL environments used for reward shaping.
- A research background in evaluation or benchmarking, including published work.
- Public work, eval frameworks, benchmark suites, or write-ups that show how you think about evaluation.