benchmarks
Benchmark Engineering
Benchmarking is among the most consequential activities in computer science, yet it has never matured into a discipline of its own. For over two decades, my research has examined how benchmarks are designed, what makes them fail, and what principled foundations would look like across domains including security, dependability, and AI. This page brings that work together under a single agenda: Benchmark Engineering. I also write opinion and perspective pieces on these challenges on my blog, Echoes of Saudade.
Opinion & Perspectives
Provocation, argument, and reflection on why computer science needs benchmarking as a first-class discipline. You can find more of my thoughts on Echoes of Saudade.
From Performance to Dependability Benchmarking: A Mandatory Path
Read paperKeynotes & Tutorials
Invited keynotes and tutorials on benchmark engineering across security, dependability, and AI.
Benchmarking GenAI for Software Engineering: Challenges and Insights
View details & presentationPerspectives on Dependability and Security Benchmarking: TO BEnchmark OR NOT TO BEnchmark
View details & presentationTrustworthiness Benchmarking of (Safety) Critical Systems
View details & presentationBenchmarking the Security of Software Systems OR TO BEnchmark or NOT TO Benchmark
View details & presentationBenchmarking Machine Learning-based Online Failure Prediction Models
View details & presentationOn the Metrics for Benchmarking Vulnerability Detection Tools
View details & presentationFoundational Research
Over two decades of publications forming the empirical foundation of the Benchmark Engineering agenda, spanning AI evaluation, security, and dependability. Full list available on the publications page.