Projects

Engineering · Huawei · Aug 2026–present

Coding-Agent Post-Training Data & Evaluation

Building executable coding tasks, validating their environments and tests, and turning agent trajectories into supervised fine-tuning data.

Role
Pipeline design and development; current validation lead
Focus
Engineering · Huawei · Aug 2026–present

Post-training adapts an existing model with additional training. For coding agents, useful training examples need more than a problem description: the code must run in a working environment, and the tests must reliably judge the result.

I designed and built a ScaleSWE-based post-training data pipeline, covering containerised task environments, task and test generation, independent validation, rollout collection with AweAgent, and export for supervised fine-tuning (SFT).

Building tasks that can be checked

The pipeline prepares an executable environment, generates a task and its tests, and checks whether those pieces work together before collecting an agent’s attempt. Independent validation checks the environment and the tests around a reference fix, while keeping reference answers and private acceptance tests separate from the solving agent.

The team built the data-production and acceptance pipeline based on the ScaleSWE paper and integrated AweAgent for solver execution and trajectory collection. A trajectory records the steps taken by an agent while working through a task.

I validated the concurrent launch of 256 task instances.

Fine-tuning and evaluation

I fine-tuned Qwen3-30B-A3B-Instruct on 2,954 agent trajectories generated by DeepSeek-v3.2. Under the same AweAgent evaluation setup, SWE-bench Verified improved from 18.0% to 54.6%.

Current focus: human–agent auditing

I am now leading validation of training tasks and environments reconstructed from real user–agent traces. I am developing methods for people and agents to audit environments, task specifications, and verifiers together, so validation can scale without treating a passing test as sufficient evidence by itself.

The implemented pipeline and initial evaluation are completed work; this auditing effort is ongoing. A public release of the full project code is not available yet.

All projects