AI Agent Evaluation
Realistic, reproducible methods for measuring task success, execution quality, tool use, error recovery, efficiency, and generalization.
AI Agent Evaluation · LLM-Enabled Software Systems · AI for Software Engineering
Senior software engineer and PhD student in Artificial Intelligence working at the intersection of large-scale production systems and applied AI.
My research focuses on how AI agents and language models can be evaluated and integrated into complex software systems. I am particularly interested in realistic agent evaluation, AI-assisted software evolution, and runtime repair supported by testing and verification.
Background
I build backend platforms and decision systems that operate under real production constraints. This experience shapes my research perspective: the value of AI in software engineering depends not only on model outputs, but also on how those outputs interact with APIs, workflows, system state, testing, observability, and runtime feedback.
I am pursuing a PhD in Artificial Intelligence and developing research at the intersection of agent evaluation, language-model applications, and software system evolution.
Research
I study intelligent software systems that can be evaluated, adapted, and improved in realistic environments.
Realistic, reproducible methods for measuring task success, execution quality, tool use, error recovery, efficiency, and generalization.
Architectures and methods for integrating language models with APIs, workflows, data, and production infrastructure.
Using language models to understand system changes, diagnose incompatibilities, and support the maintenance of evolving software.
Feedback-driven techniques that combine model reasoning with testing, observability, and verification to repair failures during execution.
Scholarship
Research in progress. Manuscripts, preprints, code, datasets, and reproducibility artifacts will be listed here as they become public.
Persistent researcher identifier: ORCID 0009-0004-0657-4600.
Industry
Engineering work spanning decision orchestration, experimentation, recommendation systems, guardrails, platform services, and production reliability.
Academic
University of the Cumberlands · In progress
Maharishi International University
Hangzhou Dianzi University
Identity