Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
Chen Yang, Ruping Xu, Ruizhe Li, Bin Cao, Jing Fan
Abstract
Extracting structured procedural knowledge from unstructured business documents is a critical yet unresolved bottleneck in process automation. While prior work has focused on extracting linear action flows from instructional texts (e.g., recipes), it has insufficiently addressed the complex logical structures-such as conditional branching and parallel execution-that are pervasive in real-world regulatory and administrative documents. Furthermore, existing benchmarks are limited by simplistic schemas and shallow logical dependencies, restricting progress toward logic-aware large language models (LLMs). To bridge this "Logic Gap", we introduce BREX, a carefully curated benchmark comprising 409 realworld business documents and 2,855 expertannotated rules. Unlike prior datasets centered on narrow service scenarios, BREX spans over 30 vertical domains, covering scientific, industrial, administrative, and financial regulations. We further propose ExIde, a structure-aware reasoning framework that investigates five distinct prompting strategies, ranging from implicit semantic alignment to executable grounding via pseudo-code generation, enabling explicit modeling of rule dependencies and providing an out-of-the-box framework for different business customers without finetuning their own LLMs. We benchmark Ex-Ide using 13 state-of-the-art LLMs. Our extensive evaluation reveals that: (1) Executable grounding serves as a superior inductive bias, significantly outperforming standard prompts in rule extraction; and (2) Reasoningoptimized models demonstrate a distinct advantage in tracing long-range dependencies and non-linear rule dependencies compared to standard instruction-tuned models. The code and dataset are available at: https://github. com/oYoungCo/Business-as-Rulesual .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 4 citations
- Never Lost in the Middle: Mastering Long-Context Question Answering with Position-Agnostic Decompositional TrainingJunqing He, Kunhao Pan, Xiaoqun Dong, Zhuoyang Song et al.ACL 2024 · 3 citations
Related papers
- StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich TextZhouhong Gu, Haoning Ye, Xingzhou Chen, Zeyang Zhou et al.ACL 2025
- From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document UnderstandingYandi Wang, Libin Zhan, Ziwei Huang, Tiancheng Luo et al.ACL 2026
- Benchmarking and Enhancing Rule Knowledge-Driven Reasoning of Large Language ModelsZijie Xu, Wenjun Ke, Peng Wang, Guozheng Li et al.AAAI 2026
- InductionBench: LLMs Fail in the Simplest Complexity ClassWenyue Hua, Tyler Wong, Fei Sun, Liangming Pan et al.ACL 2025
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and ReasoningAlan Li, Yixin Liu, Arpan Sarkar, Doug Downey et al.ICML 2026 · 4 citations
