The full agent trajectory. Not just the code.

58k real engineering tasks captured end-to-end: reasoning traces, tool calls, code edits, and explicit human acceptance signals. Built for labs training the next generation of coding agents.

Trajectories · sample
TaskRepoTypeStepsTool callsRewardStatus
T-58021api-gatewaybugfix2360.91accepted
T-58018auth-servicefeature3190.87accepted
T-58014billing-corerefactor1850.74Scoring...
T-58011ui-kitbugfix1240.95accepted
T-58007data-pipelinefeature27110.82accepted
T-58003ml-servingbugfix41140.68Scoring...
T-57998search-indexperf1970.89accepted
T-57994api-gatewaybugfix1550.93accepted
T-57990worker-poolrefactor2280.79accepted
T-57986auth-servicefeature34120.71Scoring...
T-57981cli-toolsbugfix930.96accepted

How this goes beyond the benchmarks

CapabilitySWE-benchHumanEvalCodeSearchNetZstate SWE
Full reasoning traces
Tool use captured
Human acceptance signals
Real production tasks~
Action-level reward signal
Scale (tasks)300164~100k58k

One corpus. Three lenses.

58k
Tasks

Task dataset

Real engineering problems with cleaned prompts and execution summaries. The foundation for SFT on problem comprehension and solution planning.

833k
Steps

Trajectory dataset

Full step-by-step agent traces: reasoning, tool usage, and code generation at every decision point. The complete picture of how an expert agent solves problems.

29k
Signals

Reward dataset

Explicit user acceptance signals at the action level. 50% acceptance rate. Supports iterative multi-accept workflows and action-level reward modelling.

Every tool interaction logged with context.

Semantic search
Call graph analysis
File edits
CLI execution
Directory traversal
Test runner
Symbol lookup
Dependency graph
Code commenting
Diff generation
Static analysis
more tools across the corpus

Ready to see the SWE dataset schema?

We're packaging a sample and schema for AI lab outreach now. Get in touch to be first in line, or to discuss a curated subset built for your training pipeline.