He Dong 董赫
I build AI systems that remain observable, reviewable, and safe to operate.
My work sits at the intersection of LLM agents, AI for code, developer tooling, and distributed reinforcement learning. I turn uncertain model behavior into inspectable engineering evidence—from repository-scale agents and evaluation pipelines to AST transformations and parallel training systems.
How I work
observe → reproduce → locate → patch → verify
Evidence at every boundary. Humans retain the final decision.
Now
Recent, verifiable signals from the work.
-
Six production pull requests merged into Everything Claude Code. The work spans metrics performance, cross-process dashboards, installer ownership, hook contracts, failure containment, and Unicode-safe CI.
-
Current top-10 ECC contributor by commit count. Seventeen upstream commits have landed in a project with more than 260k GitHub stars.
-
Earned GitHub's Galaxy Brain achievement through accepted, source-traced answers in open technical discussions.
-
Exploring reliable agent infrastructure and scalable RL. Current systems work covers auditable repository maintenance, structural code transformation, and distributed Actor–Learner training.
Open-source impact
Production changes, not drive-by patches.
Everything Claude Code
I contribute regression-driven changes across ECC's runtime and control surfaces. Bounded resource use, atomic transitions, explicit failure semantics, and focused tests ship with the feature.
- 6
- merged PRs
- 17
- upstream commits
- Top 10
- by commit count
- 260k+
- project stars
Review all merged PRs Browse upstream commits Contributor graph
Selected systems
Tools built around explicit contracts, evidence, and rollback paths.
RepoPilot
Agent infrastructureAn auditable AgentTeam for repository maintenance. It moves an issue or failed CI run toward a verified pull request while preserving decisions, approvals, tool calls, evidence, and rollback points.
codemod-pilot
Code intelligenceA safe, example-driven codemod engine. Give it before-and-after snippets and it infers a structural transformation, scans a repository, previews the diff, and prepares a rollback path before writing.
DRL MuJoCo
Distributed reinforcement learningA distributed Actor–Learner system for MuJoCo with parallel rollout collection, PPO optimization, multi-GPU experiments, a Rust replay buffer, and real-time experiment observability.
conda-helper
Cross-platform developer toolingA local-first CLI for safer Conda backup, restore, clone, offline packaging, cleanup, and diagnosis across Windows, macOS, and Linux—with actionable explanations for raw Conda failures.
Experience & approach
Research discipline translated into production engineering.
Beijing Institute of Technology
M.S. student researching deep reinforcement learning and agentic systems.
AI Evaluation Pipeline Automation Engineer
Led the 0→1 build of EvalHub, unifying performance benchmarks, skill evaluation, baselines, Meego releases, and Bits governance in one React + Fastify platform. Extended the shared evaluation pipeline with multi-turn skills, non-text artifacts, release gates, and a 70-case / 72-turn regression baseline.
- 58 MRs merged
- 15 days delivery window
- 70 / 72 cases / turns
AI Engineering Automation Intern
Built an Aime Workflow → PE → Skill system for MR defect tracing and automated code review. The production workflow reduced one review cycle from 20–30 minutes to 2.5–3 minutes while maintaining high-confidence localization and low false-positive rates.
- 217 defects traced
- 98% localization accuracy
- <3% false positives
Independent open-source engineering
Building agent infrastructure, repository-scale code transformations, cross-platform CLI tools, and distributed reinforcement-learning systems.
Engineering principles
- Evidence before confidence
- Reproduce the failure and make every conclusion inspectable.
- Safe autonomy
- Useful tools, explicit boundaries, approval gates, and reversible actions.
- Systems over demos
- Typed contracts, tests, observability, and deployment paths around the model.
- Performance with a baseline
- Measure against a frozen reference and understand regressions before promotion.
Working toolkit
LanguagesPython · TypeScript · Rust · Go · SQL · Shell
Agent systemsMCP · AgentTeams · evaluation harnesses · auditable execution
Code intelligencetree-sitter · AST transformations · CLI design · GitHub automation
ML & RLPyTorch · Ray · PPO · MuJoCo · parallel rollout collection
Backend & dataFastify · FastAPI · PostgreSQL · SQLite · WebSockets
ReliabilityOpenTelemetry · Docker · GitHub Actions · Vitest · pytest
Honors & recognition
Academic, competition, platform, and engineering recognition.
National Scholarship
GPA 3.91 / 4.0 and ranked 1st of 139 in Computer Science.
Outstanding Student of Hebei Province
Provincial recognition for academic performance and leadership, alongside Outstanding Student Leader and Outstanding Communist Youth League Member honors.
CUMCM 2024 · National Second Prize
Served as team leader in the national undergraduate mathematical modeling competition, coordinating modeling, collaboration, and final delivery.
Kaggle Expert · Featured Dataset
A dataset selected as Kaggle Featured with more than 3,000 downloads, alongside continued practical work across AI competitions.
ByteTech · Homepage Weekly No. 1
A technical article published on ByteDance's internal ByteTech forum ranked first on the homepage weekly chart. This is internal recognition and has no public link.
Notes
The archive remains available behind the new homepage.