Multilingual Detection of Personal Employment Status on Twitter
Manuel Tonneau, Dhaval Adjodah, João Palotti, Nir Grinberg, Samuel Fraiberger
摘要
Detecting disclosures of individuals’ employment status on social media can provide valuable information to match job seekers with suitable vacancies, offer social protection, or measure labor market flows. However, identifying such personal disclosures is a challenging task due to their rarity in a sea of social media content and the variety of linguistic forms used to describe them. Here, we examine three Active Learning (AL) strategies in real-world settings of extreme class imbalance, and identify five types of disclosures about individuals’ employment status (e.g. job loss) in three languages using BERT-based classification models. Our findings show that, even under extreme imbalance settings, a small number of AL iterations is sufficient to obtain large and significant gains in precision, recall, and diversity of results compared to a supervised baseline with the same number of labels. We also find that no AL strategy consistently outperforms the rest. Qualitative analysis suggests that AL helps focus the attention mechanism of BERT on core terms and adjust the boundaries of semantic expansion, highlighting the importance of interpretable models to provide greater control and visibility into this dynamic learning process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra 等ACL 2024 · 被引用 14 次
- Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AIIsadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous 等CSCW 2025 · 被引用 6 次
- Probabilistic Reasoning with LLMs for Privacy Risk EstimationJonathan Zheng, Alan Ritter, Sauvik Das, Wei (Coco) XuNeurIPS 2025 · 被引用 3 次
- Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language ModelsChristopher Schröder, Gerhard HeyerEMNLP 2024 · 被引用 2 次
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq 等ACL 2024
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
相关 Paper
- Domain-Guided Task Decomposition with Self-Training for Detecting Personal Events in Social MediaPayam Karisani, Joyce C. Ho, Eugene AgichteinWWW 2020 · 被引用 15 次
- Transfer and Active Learning for Dissonance Detection: Addressing the Rare-Class ChallengeVasudha Varadarajan, Swanie Juhng, Syeda Mahwish, Xiaoran Liu 等ACL 2023 · 被引用 3 次
- #Outage: Detecting Power and Communication Outages from Social NetworksUdit Paul, Alexander Ermakov, Michael Nekrasov, Vivek Adarsh 等WWW 2020 · 被引用 24 次
- ALVIN: Active Learning Via INterpolationMichalis Korakakis, Andreas Vlachos, Adrian WellerEMNLP 2024
- The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active LearningZhuang Li, Lizhen Qu, Philip R. Cohen, Raj Tumuluri 等ACL 2023 · 被引用 4 次
