Track 1: Testing and Evaluation of LLMs and SWE Agents

Track leader at TU Delft: Annibale Panichella
Track leader at JetBrains: Pouria Derakhshanfar
Phd researcher(s): Ali Asgari (TU Delft)

Autonomous AI agents, leveraging the reasoning capabilities of LLMs and multi-agent orchestration, assist developers in coding, testing, and automating workflows. However, these agents introduce new operational challenges. Ensuring the safety, robustness, reliability, and long-term maintainability of LLM-powered agents designed to solve software engineering tasks remains difficult. These agents rely on both probabilistic and continuously evolving models, making their behavior hard to predict and control. Hence, cContinuous testing and the (offline or online) evaluation of agent behavior in real-world environments are crucial.

These challenges motivate research into more rigorous, scalable, and production-ready assessment methods for LLM-driven software engineering agents, and thereby, this track targets this research direction from multiple aspects:

  1. User Experience and Operational Challenges Investigating and understanding the practical issues and friction points users encounter when integrating and operating Software Engineering (SWE) agents in real-world development workflows.
  2. Metamorphic Testing for Robustness Developing and applying metamorphic testing techniques to rigorously assess the robustness, reliability, and security of LLMs and multi-agent SWE systems, specifically targeting their probabilistic nature and unpredictable behavior.
  3. Efficient SWE Task Collection for Evaluation Focusing on methods and tools for the more efficient and scalable collection, curation, and generation of representative SWE tasks and benchmarks necessary for comprehensive and continuous evaluation of agent performance.
  4. Task Prioritization in Agent Development Cycles Researching strategies and frameworks for prioritizing which SWE tasks should be used in evaluation at different stages of the agent’s development lifecycle (e.g., pre-training, fine-tuning, production deployment) to ensure high-impact assessment and resource efficiency.

MSc Students:

Track News

06 August 2026
19 March 2026
04 April 2025
13 January 2025

Publications

  1. Ali Asgari, Mitchell Olsthoorn, and Annibale Panichella. Test Case Selection for Deep Neural Networks: A Replication Study on LLMs for Code. Accepted at the 35th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), 2026
  2. Milan de Koning, Ali Asgari, Pouria Derakhshanfar, and Annibale Panichella. A Metamorphic Testing Approach to Diagnosing Memorization in LLM-Based Program Repair. The 26th International Conference on Software Quality, Reliability, and Security (QRS), 2026
  3. Ali Asgari, Annibale Panichella, Pouria Derakhshanfar, and Mitchell Olsthoorn. What Challenges Do Developers Face in AI Agent Systems? An Empirical Study on Stack Overflow & GitHub Issues. Accepted in the Journal of Systems and Software (JSS), 2026
  4. Ali Asgari, Milan de Koning, Pouria Derakhshanfar, and Annibale Panichella. Metamorphic Testing of Deep Code Models: A Systematic Literature Review. ACM Transactions on Software Engineering Methodology (TOSEM), 2025
  5. Mohammad Mahdi Sayyadnejad, Ali Asgari, Ashkan Sami, and Hooman Tahayori. Exploring the black box: analysing explainable AI challenges and best practices through stack exchange discussions. Empirical Software Engineering (EMSE), 30, Article 176; journal-first presentation at ICSE 2026, 2025
  6. Mitchell Olsthoorn. Improving the Comprehensibility of Generated Test Suites Using Test Case Clustering. 18th IEEE International Conference on Software Testing, Verification and Validation (ICST — Short Papers, Vision and Emerging Results Track), 2025
  7. Azat Abdullin, Pouria Derakhshanfar, and Annibale Panichella. Test Wars: A Comparative Study of SBST, Symbolic Execution, and LLM-Based Approaches to Unit Test Generation. 18th IEEE International Conference on Software Testing, Verification and Validation (ICST — Research Papers Track), 2025
  8. Annibale Panichella. Metamorphic-Based Many-Objective Distillation of LLMs for Code-related Tasks. 47th IEEE/ACM International Conference on Software Engineering (ICSE — Research Track), 2025
  9. Arkadii Sapozhnikov, Mitchell Olsthoorn, Annibale Panichella, V. V. Kovalenko, and Pouria Derakhshanfar. TestSpark: IntelliJ IDEA's Ultimate Test Generation Companion. 46th IEEE/ACM International Conference on Software Engineering, Companion Proceedings (ICSE Companion — Tool Demonstration), 2024
  10. Calin Georgescu, Mitchell Olsthoorn, Pouria Derakhshanfar, Marat Akhin, and Annibale Panichella. Evolutionary Generative Fuzzing for Differential Testing of the Kotlin Compiler. ACM International Conference on the Foundations of Software Engineering (FSE — Industry Track), 2024