Post-training adapts an existing model with additional training. For coding agents, useful training examples need more than a problem description: the code must run in a working environment, and the tests must reliably judge the result.
I designed and built a ScaleSWE-based post-training data pipeline, covering containerised task environments, task and test generation, independent validation, rollout collection with AweAgent, and export for supervised fine-tuning (SFT).
Building tasks that can be checked
The pipeline prepares an executable environment, generates a task and its tests, and checks whether those pieces work together before collecting an agent’s attempt. Independent validation checks the environment and the tests around a reference fix, while keeping reference answers and private acceptance tests separate from the solving agent.
The team built the data-production and acceptance pipeline based on the ScaleSWE paper and integrated AweAgent for solver execution and trajectory collection. A trajectory records the steps taken by an agent while working through a task.
I validated the concurrent launch of 256 task instances.
Fine-tuning and evaluation
I fine-tuned Qwen3-30B-A3B-Instruct on 2,954 agent trajectories generated by DeepSeek-v3.2. Under the same AweAgent evaluation setup, SWE-bench Verified improved from 18.0% to 54.6%.
Current focus: human–agent auditing
I am now leading validation of training tasks and environments reconstructed from real user–agent traces. I am developing methods for people and agents to audit environments, task specifications, and verifiers together, so validation can scale without treating a passing test as sufficient evidence by itself.
The implemented pipeline and initial evaluation are completed work; this auditing effort is ongoing. A public release of the full project code is not available yet.