FT2: First-Token-Inspired Online Fault Tolerance on Critical Layers for Generative Large Language Models
Yu Sun, Zhu Zhu, Cherish Mulpuru, Roberto Gioiosa, Zhao Zhang, Bo Fang, Lishan Yang
2025Year
7Citations
2Top-tier citations
Abstract
Generative Large Language Models (LLMs) are deployed on large-scale computing systems, where such tasks unavoidably suffer from soft errors, leading to quality degradation of content generated by LLMs. Enhancing LLM resilience is particularly challenging because of its complicated model architecture and tremendous size. State-of-the-art protections have limitations such as high overhead and incomplete coverage, and often require offline profiling.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3d693551-e086-4c1d-b5af-ccee609bc680Cited by top-tier papers2
- Demystifying the Resilience of Large Language Model Inference: An End-to-End PerspectiveYu Sun, Zachary Coalson, Shiyang Chen, Hang Liu et al.SC 2025 · 9 citations
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu et al.HPDC 2026
Related papers
- LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM SystemsWei Liu, Yongchao He, Bohan Zhao, Hongyi Wang et al.SOSP 2026
- ErrorTrace: A Black-Box Traceability Mechanism Based on Model Family Error SpaceChuanchao Zang, Xiangtao Meng, Wenyu Chen, Tianshuo Cong et al.NeurIPS 2025 · 4 citations
- PiLLM: Resource-Efficient LLM Inference Using Workload PredictionYunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang et al.EuroSys 2026
- SymSan: Mitigating Malware Embedding in Large Language Models via Parameter Space SymmetryMing Tan, Wei Li, Tao Hu, Qian Chen et al.CCS 2026
- Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense FrameworkYuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu et al.EMNLP 2025
