Learning Bill Similarity with Annotated and Augmented Corpora of Bills
Jiseon Kim, Elden Griggs, In Song Kim, Alice Oh
摘要
Bill writing is a critical element of representative democracy. However, it is often overlooked that most legislative bills are derived, or even directly copied, from other bills. Despite the significance of bill-to-bill linkages for understanding the legislative process, existing approaches fail to address semantic similarities across bills, let alone reordering or paraphrasing which are prevalent in legal document writing. In this paper, we overcome these limitations by proposing a 5-class classification task that closely reflects the nature of the bill generation process. In doing so, we construct a human-labeled dataset of 4,721 billto-bill relationships at the subsection-level and release this annotated dataset to the research community. To augment the dataset, we generate synthetic data with varying degrees of similarity, mimicking the complex bill writing process. We use BERT variants and apply multi-stage training, sequentially fine-tuning our models with synthetic and human-labeled datasets. We find that the predictive performance significantly improves when training with both human-labeled and synthetic data. Finally, we apply our trained model to infer section-and bill-level similarities. Our analysis shows that the proposed methodology successfully captures the similarities across legal documents at various levels of aggregation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
相关 Paper
- Understanding the Language of Political Agreement and Disagreement in Legislative TextsMaryam Davoodi, Eric Waltenburg, Dan GoldwasserACL 2020 · 被引用 15 次
- War of Words II: Enriched Models of Law-Making ProcessesVictor Kristof, Aswin Suresh, Matthias Grossglauser, Patrick ThiranWWW 2021 · 被引用 5 次
- Modeling Legal Reasoning: LM Annotation at the Edge of Human AgreementRosamond Elizabeth Thalken, Edward H. Stiglitz, David Mimno, Matthew WilkensEMNLP 2023 · 被引用 9 次
- Explaining Relationships Between Scientific DocumentsKelvin Luu, Xinyi Wu, Rik Koncel-Kedziorski, Kyle Lo 等ACL 2021
- Synthetic Bootstrapped PretrainingZitong Yang, Aonan Zhang, Hong Liu, Tatsunori Hashimoto 等ICLR 2026 · 被引用 3 次
