AAAI2021
Comparing Symbolic Models of Language via Bayesian Inference (Student Abstract)
Annika Heuser, Polina Tsvilodub
Abstract
Given recurring interest in structured representations in computational cognitive models, we extend a Bayesian scoring procedure for comparing symbolic models of language grammar. We conduct a case-study of modeling syntactic principles in German, providing preliminary results consistent with linguistic theory. We also note that dataset and part-of-speech (POS) tagger quality should not be taken for granted. Recent advances in AI have brought back the controversy between symbolic and connectionist approaches to knowledge representation and learning (e.g., Lake et al. 2017). While great progress in AI has been brought about by neural models, they in general require a lot of training data to achieve human-like generalization, while children make correct generalizations and learn their native language based on substantially fewer examples (Tenenbaum et al. 2011; Xu and Tenenbaum 2007; Chomsky 1965) . One position holds that this sample size problem is remedied by resorting to an initial bias. Specifically, many linguists argue that symbolic hierarchical representations and compositionality principles are indispensable for human-like language processing and production (e.g., Crain and Nakayama 1987; Chomsky 1965; Pelletier 1994; Lake et al. 2017 ). These principles could potentially be incorporated as structural constraints in models of natural language. One approach to building and evaluating human-like symbolic language models comes from the increasingly influential domain of Bayesian learning (e.g., Xu and Tenenbaum 2007; Perfors, Tenenbaum, and Regier 2011) . In particular, Perfors, Tenenbaum, and Regier (2011) indicate that a learner equipped with domain-general Bayesian inference capacities favors hierarchical phrase structure over linear phrase structure given linguistic data representing input available to human learners (i.e., from a child-directed speech corpus). Perfors, Tenenbaum, and Regier (2011) represent several explicit hypotheses about syntactic structure as probabilistic grammars which are compared via their Bayesian posterior scores given training data, assuming a meta-grammar which generated those particular hypotheses. We note that the models were all able to parse the entirety of the data they were compared on, making them directly comparable via posterior