Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation
Masato Mita, Soichiro Murakami, Akihiko Kato, Peinan Zhang
Abstract
In response to the limitations of manual ad creation, significant research has been conducted in the field of automatic ad text generation (ATG).However, the lack of comprehensive benchmarks and well-defined problem sets has made comparing different methods challenging.To tackle these challenges, we standardize the task of ATG and propose a first benchmark dataset, CAMERA , carefully designed and enabling the utilization of multi-modal information and facilitating industry-wise evaluations.Our extensive experiments with a variety of nine baselines, from classical methods to state-of-the-art models including large language models (LLMs), show the current state and the remaining challenges.We also explore how existing metrics in ATG and an LLMbased evaluator align with human evaluations.ORIX Card Loan Keyword Diagnosis of instant loan Cards! 3 recommended companies to borrow Landing page (LP) Ad text 1. [Official] Top 3 Popular Card Loans 2. Easily diagnose recommended card loans 3. Diagnose Cards Availbale for Same-Day Borrowing ! 4. Get Financing in as Fast as 30 Mnutes Online !
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- OMS: On-the-fly, Multi-Objective, Self-Reflective Ad Keyword Generation via LLM AgentBowen Chen, Zhao Wang, Shingo TakamatsuEMNLP 2025
- Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following AbilityYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2025
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
Related papers
- MGTBench: Benchmarking Machine-Generated Text DetectionXinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes et al.CCS 2024 · 30 citations
- ARGUS: Hallucination and Omission Evaluation in Video-LLMsRuchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli et al.ICCV 2025 · 1 citation
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng et al.ICLR 2026 · 27 citations
- AIR-Bench: Automated Heterogeneous Information Retrieval BenchmarkJianlyu Chen, Nan Wang, Chaofan Li, Bo Wang et al.ACL 2025
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes et al.USENIX Security 2025
