Engineering
Post-Training Data & Evaluation
Building executable coding tasks, validating their environments and tests, and turning agent trajectories into supervised fine-tuning data.
18.0% to 54.6% on SWE-bench Verified

Senior Engineer at Huawei, Shanghai, China.
PhD in Computer Science, Fudan University.
Hi, I’m Junming Cao (曹峻铭), and I also go by James. I build execution and verification infrastructure for coding agents. My current focus is human–agent auditing of training environments, task specifications, and verifiers for post-training.
I’ll be attending NeurIPS 2026 in Sydney this December, and would love to meet fellow students, researchers, and engineers for a coffee chat.
I became a committer in the Cannbot Community. The Cannbot community builds open-source coding-agent tools for Ascend NPU development in the CANN ecosystem. It brings together reusable skills, domain knowledge, and workflows for kernel development, model migration, and inference optimization.
Our paper, Lifting the Veil on Composition, Risks, and Mitigations of the Large Language Model Supply Chain, was accepted to ACM TOSEM.
Our CodeMap paper, Understanding Codebase like a Professional! Human–AI Collaboration for Code Comprehension, received an ACM SIGSOFT Distinguished Paper Award at ICPC 2026.
Our paper, Understanding Performance Problems in CUDA Programs, was accepted to FSE 2026.
Our CodeMap paper, Understanding Codebase like a Professional! Human–AI Collaboration for Code Comprehension, was accepted to ICPC 2026.
I attended EMNLP 2025 in Suzhou, China, and enjoyed many inspiring conversations with fellow researchers and friends.
I joined Huawei’s Applied Software Engineering Lab, 2012 Lab, as a Senior Engineer in Shanghai.
Our paper, RegTrieve: Reducing System-Level Regression Errors for Machine Learning Systems via Retrieval-Enhanced Ensemble, was accepted to FSE 2025.
Engineering
Building executable coding tasks, validating their environments and tests, and turning agent trajectories into supervised fine-tuning data.
18.0% to 54.6% on SWE-bench Verified
Engineering
Shared execution and verification infrastructure for coding agents that generate, debug, and migrate accelerator kernels.
#1 on CannBench 910Cannbot Lingxi-Evo · Aug–Sep 2026
Research
Studying real performance problems, building a reproducible benchmark, and developing a static checker that found 488 new issues in open-source repositories.
27 issues were fixed
How milestones, architecture decision records, and task specifications helped me keep fast-moving coding agents focused on the next release.
A compute-kernel example of moving state, retries, and acceptance checks from prompt instructions into the software around a coding agent.
Research agents benefit from stable goals and fast, automatic feedback. Application development needs the same foundations, plus an explicit boundary around what tests cannot decide.
I contribute execution, verification, and diagnostic tools for Ascend kernel engineering and model training.
Committer, cannbot community, since July 2026. The Cannbot community builds open-source coding-agent tools for Ascend NPU development in the CANN ecosystem. It brings together reusable skills, domain knowledge, and workflows for kernel development, model migration, and inference optimization. Contributed the Ascend kernel-migration orchestration engine and layered knowledge management.
Contributed a GPU/NPU training-diagnostics skill using msProbe to compare tensors layer by layer and investigate numerical differences.
Co-authored a study of how professionals understand unfamiliar codebases, based on interviews with eight code auditors. The findings informed CodeMap, an LLM-powered system with hierarchical codebase visualisations and interactive navigation. Among experienced developers, CodeMap reduced time spent reading LLM responses by 79%.
ACM SIGSOFT Distinguished Paper Award
Developed RegTrieve, a retrieval-enhanced ensemble that dynamically combines old and updated models to reduce system-level regression errors in multi-model machine-learning systems.
Built CodeReview-New, a 14,568-example benchmark for code refinement from review comments. Compared ChatGPT with CodeReviewer and analysed 400 cases, identifying missing domain knowledge and ambiguous edit locations or review intent as major sources of failure.
Studied 224 performance problems and built a 58-case reproducible benchmark. Developed DeepPerf, a rule-based static checker; evaluation on 1,108 open-source repositories identified 488 new issues, of which 27 were fixed.