Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval
Valentin Knappich, Anna Hätty, Simon Razniewski, Annemarie Friedrich
摘要
Novelty assessment is a critical yet complex task in the examination process for patent acceptance, requiring examiners to determine whether an invention is disclosed in a prior art document. The process involves intricate matching between specific features of a patent claim and passages in the prior art. While prior work has approached novelty prediction primarily as a binary classification task at the claim level, we argue that this formulation is susceptible to spurious correlations and lacks the granularity required for practical application. In this work, we introduce FiNE-Patents (Fine-grained Novelty Examination of Patents), a novel dataset comprising 3,658 first patent claims annotated with fine-grained, feature-level prior art references extracted from European Search Opinion (ESOP) documents. We propose shifting the evaluation paradigm from simple binary classification to a joint retrieval and abstract reasoning task at the feature level, requiring models to identify specific passages from a prior art document that disclose individual claim features, and to identify which features of a claim make it novel. We implement and evaluate LLM-based workflows that decompose claims into features, analyze each feature against prior art, and finally derive a claim-level novelty prediction. Our experiments demonstrate that these workflows outperform embedding-based baselines on passage retrieval and novel feature identification. Furthermore, we show that unlike trained classifiers, LLMs are robust against spurious correlations present in the claim-level novelty classification task. We release the dataset and code to foster further research into transparent and granular patent analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang 等ICLR 2024 · 被引用 170 次
- A Survey on Patent Analysis: From NLP to Multimodal AIHomaira Huda Shomee, Zhu Wang, Sathya N. Ravi, Sourav MedyaACL 2025 · 被引用 12 次
- Enhancing the Patent Matching Capability of Large Language Models via the Memory GraphQiushi Xiong, Zhipeng Xu, Zhenghao Liu, Mengjia Wang 等SIGIR 2025 · 被引用 4 次
相关 Paper
- Towards Comprehensive Patent Approval Predictions: Beyond Traditional Document ClassificationXiaochen Gao, Zhaoyi Hou, Yifei Ning, Kewen Zhao 等ACL 2022
- LePaRD: A Large-Scale Dataset of Judicial Citations to PrecedentRobert Mahari, Dominik Stammbach, Elliott Ash, Alex PentlandACL 2024 · 被引用 1 次
- PatentLMM: Large Multimodal Model for Generating Descriptions for Patent FiguresShreya Shukla, Nakul Sharma, Manish Gupta, Anand MishraAAAI 2025 · 被引用 6 次
- PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent ClaimsYongmin Yoo, Qiongkai Xu, Longbing CaoEMNLP 2025 · 被引用 1 次
- NSF-SciFy: Mining the NSF Awards Database for Scientific ClaimsDelip Rao, Weiqiu You, Eric Wong, Chris Callison-BurchACL 2026 · 被引用 1 次
