Neural Tangent Knowledge Distillation for Optical Convolutional Networks
Jinlin Xiang, Minho Choi, Yubo Zhang, Zhihao Zhou, Arka Majumdar, Eli Shlizerman
Abstract
Hybrid Optical Neural Networks (ONNs, typically consisting of an optical frontend and a digital backend) offer an energy-efficient alternative to fully digital deep networks for real-time, power-constrained systems. However, their adoption is limited by two main challenges: the accuracy gap compared to large-scale networks during training, and discrepancies between simulated and fabricated systems that further degrade accuracy. While previous work has proposed end-to-end optimizations for specific datasets (e.g., MNIST) and optical systems, these approaches typically lack generalization across tasks and hardware designs. To address these limitations, we propose a task-agnostic and hardware-agnostic pipeline that supports image classification and segmentation across diverse optical systems. To assist optical system design before training, we estimate achievable model accuracy based on user-specified constraints such as physical size and the dataset. For training, we introduce Neural Tangent Knowledge Distillation (NTKD), which aligns optical models with electronic teacher networks, thereby narrowing the accuracy gap. After fabrication, NTKD also guides fine-tuning of the digital backend to compensate for implementation errors. Experiments on multiple datasets (e.g., MNIST, CIFAR, Carvana Masking) and hardware configurations show that our pipeline consistently improves ONN performance and enables practical deployment in both pre-fabrication simulations and physical implementations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect TeacherGuangda Ji, Zhanxing ZhuNeurIPS 2020 · 56 citations
- Physics-Constrained Comprehensive Optical Neural NetworksYanbing Liu, Jianwei Qin, Yan Liu, Xi Yue et al.NeurIPS 2024 · 4 citations
- NTK-SAP: Improving neural network pruning by aligning training dynamicsYite Wang, Dawei Li, Ruoyu SunICLR 2023 · 2 citations
- Supervision Complexity and its Role in Knowledge DistillationHrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim et al.ICLR 2023 · 1 citation
Related papers
- LightRidge: An End-to-end Agile Design Framework for Diffractive Optical Neural NetworksYingjie Li, Ruiyang Chen, Minhan Lou, Berardi Sensale Rodriguez et al.ASPLOS 2023 · 6 citations
- Efficient Neural Vision Systems Based on Convolutional Image AcquisitionPedram Pad, Simon Narduzzi, Clément Kündig, Engin Türetken et al.CVPR 2020
- Binary Optical Machine Learning: Million-Scale Physical Neural Networks with Nano NeuronsXueyuan Yang, Zhenlin An, Qingrui Pan, Lei Yang et al.MobiCom 2024 · 3 citations
- Quantized Feature Distillation for Network QuantizationKe Zhu, Yin-Yin He, Jianxin WuAAAI 2023 · 21 citations
- Revisit the Power of Vanilla Knowledge Distillation: from Small Scale to Large ScaleZhiwei Hao, Jianyuan Guo, Kai Han, Han Hu et al.NeurIPS 2023 · 17 citations
