AI Agent Evaluation · LLM-Enabled Software Systems · AI for Software Engineering

Shuhan Sun

Senior software engineer and PhD student in Artificial Intelligence working at the intersection of large-scale production systems and applied AI.

My research focuses on how AI agents and language models can be evaluated and integrated into complex software systems. I am particularly interested in realistic agent evaluation, AI-assisted software evolution, and runtime repair supported by testing and verification.

Portrait of Shuhan Sun

Background

Production systems shaping applied AI research

I build backend platforms and decision systems that operate under real production constraints. This experience shapes my research perspective: the value of AI in software engineering depends not only on model outputs, but also on how those outputs interact with APIs, workflows, system state, testing, observability, and runtime feedback.

I am pursuing a PhD in Artificial Intelligence and developing research at the intersection of agent evaluation, language-model applications, and software system evolution.

Research

Research interests

I study intelligent software systems that can be evaluated, adapted, and improved in realistic environments.

01

AI Agent Evaluation

Realistic, reproducible methods for measuring task success, execution quality, tool use, error recovery, efficiency, and generalization.

02

LLM-Enabled Software Systems

Architectures and methods for integrating language models with APIs, workflows, data, and production infrastructure.

03

AI-Assisted Software Evolution

Using language models to understand system changes, diagnose incompatibilities, and support the maintenance of evolving software.

04

Runtime Repair & Verification

Feedback-driven techniques that combine model reasoning with testing, observability, and verification to repair failures during execution.

Questions I am exploring

  • How can agent evaluations reflect realistic, multi-step tool workflows and explain why systems succeed or fail?
  • How can language models reason over code, API contracts, logs, and system state to support software evolution?
  • How can testing, verification, and runtime feedback determine whether AI-generated changes actually work?

Scholarship

Publications & research artifacts

Research in progress. Manuscripts, preprints, code, datasets, and reproducibility artifacts will be listed here as they become public.

Persistent researcher identifier: ORCID 0009-0004-0657-4600.

Industry

Experience

Senior Software Engineer · eBay

Large-scale backend and decision systems

Engineering work spanning decision orchestration, experimentation, recommendation systems, guardrails, platform services, and production reliability.

Academic

Education

PhD in Artificial Intelligence

University of the Cumberlands · In progress

MS in Computer Science

Maharishi International University

BEng in Electrical Engineering & Automation

Hangzhou Dianzi University

Identity

Research profiles