When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generation
Yue Liu, Zhenchang Xing, Shidong Pan, Chakkrit Tantithamthavorn
摘要
In recent years, the AI wave has grown rapidly in software development. Even novice developers can now design and generate complex framework-constrained software systems based on their high-level requirements with the help of Large Language Models (LLMs). However, when LLMs gradually "take the wheel" of software development, developers may only check whether the program works. They often miss security problems hidden in how the generated programs are implemented.
In this work, we investigate the security properties of frameworkconstrained programs generated by state-of-the-art LLMs. We focus specifically on Chrome extensions due to their complex security model involving multiple privilege boundaries and isolated components. To achieve this, we built ChromeSecBench, a dataset with 140 prompts based on known vulnerable extensions. We used these prompts to instruct nine state-of-the-art LLMs to generate complete Chrome extensions, and then analyzed them for vulnerabilities across three dimensions: scenario types, model differences, and vulnerability categories. Our results show that LLMs produced vulnerable programs at alarmingly high rates (18%-50%), particularly in Authentication & Identity and Cookie Management scenarios (up to 83% and 78% respectively). Most vulnerabilities exposed sensitive browser data like cookies, history, or bookmarks to untrusted code. Interestingly, we found that advanced reasoning models performed worse, generating more vulnerabilities than simpler models. These findings highlight a critical gap between LLMs' coding skills and their ability to write secure framework-constrained programs.
• Security and privacy → Software and application security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- Do Users Write More Insecure Code with AI Assistants?Neil Perry, Megha Srivastava, Deepak Kumar, Dan BonehCCS 2023 · 被引用 150 次
- EmPoWeb: Empowering Web Applications with Browser ExtensionsDolière Francis SoméS&P 2019 · 被引用 60 次
- DoubleX: Statically Detecting Vulnerable Data Flows in Browser Extensions at ScaleAurore Fass, Dolière Francis Somé, Michael Backes, Ben StockCCS 2021 · 被引用 35 次
- PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)Mahmoud Nazzal, Issa Khalil, Abdallah Khreishah, NhatHai PhanCCS 2024 · 被引用 16 次
相关 Paper
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesHaoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin 等ACL 2025
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si 等ICML 2026
- Rethinking the Evaluation of Secure Code GenerationShih-Chieh Dai, Jun Xu, Guanhong TaoICSE 2026 · 被引用 1 次
- RMCBench: Benchmarking Large Language Models' Resistance to Malicious CodeJiachi Chen, Qingyuan Zhong, Yanlin Wang, Kaiwen Ning 等ASE 2024 · 被引用 6 次
