publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- Repairing LLM Executions for Secure Automatic ProgrammingAli El Husseini, Yacine Izza, Blaise Genest, and Abhik RoychoudhuryIn Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering, 2026
While automatic code generation using Large Language Models (LLMs) has advanced significantly, these models frequently produce code containing security vulnerabilities. Existing approaches to improve the security of automatically generated code, such as fine-tuning or prompt engineering, have shown limited success and provide minimal insight into the underlying mechanisms causing these vulnerabilities. We propose an approach grounded in mechanistic interpretability to analyze and mitigate vulnerable code generation in LLMs. We begin by examining the knowledge stored inside LLMs, identifying and disentangling knowledge representations that contribute to generating vulnerable code. Next, we leverage these insights to repair model execution in real time: when the model attempts to access vulnerability-inducing representations during inference, our method intercepts and modifies this access, improving the security of the generated code. We implement our methodology in a tool called Thea and evaluate it on the CyberSecEval benchmark using Llama 3.1. Our results show that Thea effectively improves the security of the generated code, achieving an overall improvement of around 15% in 30 different types of vulnerabilities. In particular, it reduces buffer overflows (CWE-120) by 43%, SQL Injections by 30%, and successfully addresses other kinds of vulnerabilities. Our analysis further reveals that in cases where vulnerability reduction is less substantial (e.g. an 11% reduction for CWE-338), the insights behind Thea can be leveraged to reliably detect the occurrence of a vulnerability, enabling us to provide appropriate warnings to users when complete remediation is not possible. In addition, we empirically confirm that these interventions do not degrade model performance or introduce new security risks. Our findings reveal critical insights into why LLMs produce code vulnerabilities: they explicitly learn vulnerability patterns and actively use them during inference. We demonstrate how this can be leveraged to repair LLM executions, allowing us to avoid such vulnerability patterns.
@inproceedings{elhusseini2026repairing, title = {Repairing LLM Executions for Secure Automatic Programming}, author = {El Husseini, Ali and Izza, Yacine and Genest, Blaise and Roychoudhury, Abhik}, booktitle = {Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering}, series = {ICSE '26}, year = {2026}, pages = {818--829}, numpages = {12}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, isbn = {9798400720253}, doi = {10.1145/3744916.3773210}, url = {https://doi.org/10.1145/3744916.3773210}, keywords = {Large Language Models; Secure Code Generation; Software Security; Mechanistic Interpretability; Vulnerability Prevention; Automatic Programming}, }
2025
- Detecting Python Malware in the Software Supply Chain with Program AnalysisRidwan Shariffdeen, Behnaz Hassanshahi, Martin Mirchev, Ali El Husseini, and Abhik RoychoudhuryIn International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2025
@inproceedings{shariffdeen2025detecting, title = {Detecting Python Malware in the Software Supply Chain with Program Analysis}, author = {Shariffdeen, Ridwan and Hassanshahi, Behnaz and Mirchev, Martin and El Husseini, Ali and Roychoudhury, Abhik}, booktitle = {International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP)}, year = {2025}, }