FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic Encryption
Mayank Kumar, Jiaqi Xue, Mengxin Zheng, Qian Lou
摘要
Fully Homomorphic Encryption (FHE) is a foundational technology for confidential computing, yet its practical adoption remains limited by the need for specialized cryptographic expertise and error-prone parameter configuration. To lower this barrier, we investigate whether Large Language Model (LLM) agents can reliably generate secure FHE code from natural-language specifications. We present FHE-Coder, a three-phase agentic framework that addresses the key failure modes of FHE code generation: semantic ambiguity, API misuse, and cryptographic insecurity. The framework integrates (1) a Prompt Formalizer that structures user intent and enforces secure parameterization, (2) a specialized retrievalaugmented generation (RAG) module that supplies scheme-specific API and documentation knowledge, and (3) an automated Security Verifier that performs iterative validation and feedback to detect and correct cryptographic flaws. We evaluate FHE-Coder across four leading LLMs on a benchmark of ten FHE programming tasks spanning increasing functional and security complexity. While baseline agents frequently produce code that compiles and passes functional tests, they often violate security constraints or misuse cryptographic parameters. In contrast, FHE-Coder consistently generates solutions that are compilable, functionally correct, and verifiably secure across schemes including TFHE and CKKS. Our work establishes a systematic methodology and benchmark for agentic FHE code generation, providing a practical step toward democratizing secure computation without compromising cryptographic guarantees. Project page: https://fhe-coder.github.io Recent advances extend beyond standalone LLMs toward LLM agents, which integrate planning, retrieval, and tool-use capabilities into iterative reasoning pipelines. In adjacent domains such as High-Level Synthesis (HLS) and Register Transfer Level (RTL) hardware design (Thakur et al., 2023; Liao et al., 2024; Xiong et al., 2024) , LLM-based systems have demonstrated the ability to
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- On the Robustness of Code Generation Techniques: An Empirical Study on GitHub CopilotAntonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi, Matteo Ciniselli 等ICSE 2023 · 被引用 124 次
- AutoQ: Automated Kernel-Wise Neural Network QuantizationQian Lou, Feng Guo, Minje Kim, Lantao Liu 等ICLR 2020 · 被引用 121 次
- Glyph: Fast and Accurately Training Deep Neural Networks on Encrypted DataQian Lou, Bo Feng, Geoffrey Charles Fox, Lei JiangNeurIPS 2020 · 被引用 106 次
- HEMET: A Homomorphic-Encryption-Friendly Privacy-Preserving Mobile Neural Network ArchitectureQian Lou, Lei JiangICML 2021 · 被引用 88 次
相关 Paper
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 被引用 86 次
- Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based BenchmarkingJiexin Wang, Liuwen Cao, Xitong Luo, Yang Cao 等ISSTA 2026 · 被引用 4 次
- A Pair Programming Framework for Code Generation via Multi-Plan Exploration and Feedback-Driven RefinementHuan Zhang, Wei Cheng, Yuhan Wu, Wei HuASE 2024 · 被引用 7 次
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu 等ISSTA 2024 · 被引用 10 次
