<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">2397-1835</journal-id>
<journal-title-group>
<journal-title>Glossa: a journal of general linguistics</journal-title>
</journal-title-group>
<issn pub-type="epub">2397-1835</issn>
<publisher>
<publisher-name>Ubiquity Press</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5334/gjgl.1323</article-id>
<article-categories>
<subj-group>
<subject>Research</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Varying Abstractions: a conceptual vs. distributional view on prepositional polysemy</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-5706-8418</contrib-id>
<name>
<surname>Fonteyn</surname>
<given-names>Lauren</given-names>
</name>
<email>l.fonteyn@hum.leidenuniv.nl</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>Leiden University, Arsenaalstraat 1, 2311 CT Leiden, NL</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2021-07-06">
<day>06</day>
<month>07</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>6</volume>
<issue>1</issue>
<elocation-id>90</elocation-id>
<history>
<date date-type="received" iso-8601-date="2020-05-23">
<day>23</day>
<month>05</month>
<year>2020</year>
</date>
<date date-type="accepted" iso-8601-date="2021-02-16">
<day>16</day>
<month>02</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2021 The Author(s)</copyright-statement>
<copyright-year>2021</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://www.glossa-journal.org/articles/10.5334/gjgl.1323/"/>
<abstract>
<p>The term &#8216;meaning&#8217;, as it is presently employed in Linguistics, is a polysemous concept, covering a broad range of operational definitions. Focussing on two of these definitions, meaning as &#8216;concept&#8217; and meaning as &#8216;context&#8217; (also known as &#8216;distributional semantics&#8217;), this paper explores to what extent these operational definitions lead to converging conclusions regarding the number and nature of distinct senses a polysemous form covers. More specifically, it investigates whether the sense network that emerges from the principled polysemy model of <italic>over</italic> as proposed by Tyler &amp; Evans (<xref ref-type="bibr" rid="B72">2003</xref>; <xref ref-type="bibr" rid="B71">2001</xref>) can be reconstructed by the neural language model BERT. The study assesses whether the contextual information encoded in BERT embeddings can be employed to succesfully (i) recognize the abstract sense categories and (ii) replicate the relative distances between the senses of <italic>over</italic> proposed in the principled polysemy model. The results suggest that, while there is partial convergence, the two models ultimately lead to different global abstractions because the imagistic information that plays a key role in conceptual approaches to prepositional meaning may not be encoded in contextualized word embeddings.</p>
</abstract>
<kwd-group>
<kwd>Prepositions</kwd>
<kwd>Polysemy</kwd>
<kwd>Image Schema</kwd>
<kwd>Metaphor</kwd>
<kwd>Distributional Semantics</kwd>
<kwd>BERT</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>1 Introduction</title>
<p>The aim of the present study is to empirically investigate whether there is any correspondence between the generalizations that emerge from different operational models of meaning representation. More specifically, this study focuses on the operational definition of meaning as &#8216;concept&#8217;, as commonly employed in Cognitive Linguistics, and meaning defined as (or derived from) &#8216;context&#8217;, also known as distributional semantics, which has rapidly gained popularity in Corpus Linguistics and Computational Linguistics/NLP. As a case study, it will home in on the semantics of the English preposition <italic>over</italic>.</p>
<p>Following Brugman (<xref ref-type="bibr" rid="B8">1988</xref>)&#8217;s and Lakoff (<xref ref-type="bibr" rid="B45">1987</xref>)&#8217;s extensive treatment of <italic>over</italic>, prepositions have become default examples in descriptions of the core tenets &#8216;the cognitive approach&#8217; to meaning (see, e.g. <xref ref-type="bibr" rid="B49">Lemmens 2016</xref>). In the cognitive-conceptual approach to semantics, meanings are defined as &#8216;concepts&#8217; which are connected and grounded in complex knowledge structures (e.g. <xref ref-type="bibr" rid="B15">Clausner &amp; Croft 1999</xref>), often described as complex chains or &#8216;networks&#8217; of connected senses. Central to the conceptual approach is that these concepts (and their network of senses) are to a large extent experiential &#8211; that is, they are grounded in the physical or cultural experience of the language user.</p>
<p>When it comes to their discussion of polysemy, such cognitive-conceptual accounts have faced substantial criticism, predominantly aimed at their apparent lack of principled and objective methods to determine how many senses can be distinguished, and how the global design of the polysemy networks is construed. One part of the problem appeared to be that the identification of the core node (or &#8216;prototype&#8217;) of the network seemed to rely solely on subjective, introspective judgements (<xref ref-type="bibr" rid="B66">Sandra &amp; Rice 1995</xref>; <xref ref-type="bibr" rid="B62">Rice 1996: 137</xref>). These criticisms triggered a search for concrete criteria and data-driven tests for prototypicality (e.g. <xref ref-type="bibr" rid="B27">Gilquin &amp; McMichael 2018</xref>; <xref ref-type="bibr" rid="B53">Newman 2011</xref>). A similar discussion also arose regarding the position of derived nodes, with many scholars acknowledging that more objective, non-introspective discussions regarding &#8216;distances&#8217; between senses will remain all but impossible as long as we are not &#8220;able to <italic>measure</italic> the degree of similarity between senses&#8221; (<xref ref-type="bibr" rid="B32">Gries &amp; Divjak 2009: 57</xref>; emphasis mine).</p>
<p>In response to the need for more well-defined and data-driven definitions of word senses (e.g <xref ref-type="bibr" rid="B43">Kilgarriff 2003</xref>), and more objective, measurable ways of determining distances between senses, a number of proposals have been devised that could be subsumed under the header of &#8216;contextual&#8217; or &#8216;distributional&#8217; approaches to semantics. Underlying these approaches is the premise that the meaning of a word can be derived from the context in which it occurs (an idea which dates back to at least Harris (<xref ref-type="bibr" rid="B37">1954</xref>) and Firth (<xref ref-type="bibr" rid="B23">1957</xref>)). While not necessarily equating context to meaning, distributional approaches to semantics are based on the assumption that co-occurrence patterns and other distributional frequencies are indicative of the meaning of a linguistic item (i.e., they serve as &#8220;proxies&#8221; for meaning representation; see, for instance Baroni et al. (<xref ref-type="bibr" rid="B3">2014: 238</xref>)). Hence, distributional similarities between linguistic items in a corpus can be used to approximate a measurement of their functional or semantic similarity, thus placing more prominence on the geometrical relationship between linguistic items (<xref ref-type="bibr" rid="B7">Boleda &amp; Erk 2015</xref>). Particularly in recent years, the approach has been met with great enthusiasm in Computational Linguistics and Machine Learning, as recent incarnations of such distributional semantic models seem to perform astonishingly well on a wide range of NLP, production, and machine translation tasks (<xref ref-type="bibr" rid="B78">Young et al. 2018</xref>; <xref ref-type="bibr" rid="B60">Radford et al. 2018</xref>).</p>
<p>A similar (yet overall more cautious) enthusiasm has been expressed in Linguistics: being almost exclusively corpus-based, distributional approaches have provides a welcome bridge between the &#8220;well-established corpus-linguistic research tradition and Langacker&#8217;s idea that linguistic representations emerge from linguistic usage&#8221; (<xref ref-type="bibr" rid="B68">Stefanowitsch 2010: 370</xref>). Furthermore, because some of the more traditional conceptual models of polysemy already depended on distributional criteria to some extent (<xref ref-type="bibr" rid="B32">Gries &amp; Divjak 2009: 58</xref>; <xref ref-type="bibr" rid="B24">Geeraerts 2016: 241</xref>), the path to a full-fledged distributional approach had already been cleared. Such an approach involves virtually no introspective manual annotation (but see, e.g., comments in Gries &amp; Divjak (<xref ref-type="bibr" rid="B32">2009</xref>) and <xref ref-type="bibr" rid="B38">Heylen et al. (2015: 154)</xref>), and hence it allows the elusive concept of meaning to be studied in a more rigorous, large-scale, and ultimately quantifiable way. The confidence that distributional methods provide an appropriate, objective and fully data-driven alternative to more introspective models appears to be growing steadily, as there have been some successes in replicating experimental, survey-based accounts of polysemy by means of distributional methods (e.g. <xref ref-type="bibr" rid="B4">Berez &amp; Gries 2008</xref>; <xref ref-type="bibr" rid="B33">Gries &amp; Divjak 2010</xref>). In particular in historical linguistics, where native-speaker intuitions regarding the semantics of linguistic items is inevitably inaccessible, distributional methods are now considered a welcome methodological innovation (e.g. <xref ref-type="bibr" rid="B64">Sagi et al. 2011</xref>; <xref ref-type="bibr" rid="B39">Hilpert &amp; Correia Saavedra 2017</xref>; <xref ref-type="bibr" rid="B9">Budts 2020</xref>).</p>
<p>Yet, while the distributional approach and the more traditional, cognitive-conceptual approach essentially share the same goal (that is, to capture the complex internal semantic structure of linguistic items in a rigorous, theoretically motivated, principled manner), it is also clear that there may be a non-trivial epistemological difference between the two approaches. The distributional approach to meaning and sense distinctions is based on a premise that essentially conflicts with one of the core criteria of, for instance, the principled polysemy approach (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>), which treats a linguistic item&#8217;s meaning as precisely that which cannot be directly inferred from contextual cues. Naturally, then, the question arises to what extent the two approaches are related to one another, and whether they can still lead to comparable results. A positive answer to this question (i.e. the results of the two approaches converge) suggests that distributional models are able to use contextual cues to extract and encode the conceptual information that lies at the core of the cognitive-conceptual model. This would mean that we can indeed further fine-tune proposals within the conceptual approach by enabling a quantitative discussion on sense distinctions and possible derivational paths. A negative answer, by contrast, would raise questions about the compatibility of the two approaches, and may even have larger implications. It may seem fair to operate under the &#8220;a priori&#8221; assumption that different operational definitions of meaning are designed to capture the same, unitary phenomenon (<xref ref-type="bibr" rid="B68">Stefanowitsch 2010: 371</xref>), but if the study of meaning is operationalized in two different ways, and we find that the results of a distributional and cognitive-conceptual model lead to different generalizations and abstractions, one may start to question whether they are in fact capturing the same phenomenon (<xref ref-type="bibr" rid="B24">Geeraerts 2016</xref>).</p>
<p>Prioritizing depth over width, this study is set up as a detailed empirical comparison between two operational models of meaning representation, with one serving as a representative of the cognitive semantic (&#8216;concept&#8217;) approach, and one representing the distributional (&#8216;context&#8217;) approach. More specifically, this study focuses on one of the most well-developed cognitive-conceptual proposals &#8211; the principled polysemy model of <italic>over</italic> as set out by Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>) &#8211; and aims to investigate whether the polysemy network that emerges from this theoretical model of meaning representation can be reconstructed by means of a recent neural distributional language model called BERT (&#8216;Bidirectional encoder Representations from Transformers&#8217;). To this end, a stratified sample of 808 contextualized instances of the preposition <italic>over</italic> has been annotated following the criteria outlined in Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>). This annotated data set serves as a &#8216;theory-specific standard&#8217; against which the output of BERT will be assessed. This assessment is targeted at determining whether one can distinguish the same sense categories as proposed in the principled polysemy model by means of BERT embeddings, but it also addresses the question whether there is overlap between the operational models in terms of the suggested similarities and relations between these senses.</p>
<p>What emerges from the analysis is that BERT clearly captures fine-grained, local semantic similarities between tokens. Even with an entirely unsupervised application of BERT, discrete, coherent token groupings can be discerned that correspond relatively well with the sense categories proposed by means of the principled polysemy model. Furthermore, embeddings of <italic>over</italic> also clearly encode information about conceptual domains, as concrete, spatial uses of <italic>over</italic> are neatly distinguished from more abstract, metaphorical extensions (into the conceptual domain of time, or other non-spatial domains). However, there are no indications that BERT embeddings also encode information about the abstract configurational resemblances between tokens across those domains. As such, the global picture of resemblance between sense categories that emerges from the unsupervised application of BERT differs substantially from the theoretical proposal by Tyler &amp; Evans (<xref ref-type="bibr" rid="B72">2003</xref>; <xref ref-type="bibr" rid="B71">2001</xref>), which heavily relies on the language user&#8217;s ability to recognize schematic, imagistic similarities within and across conceptual domains. These findings highlight the fact that such imagistic similarities are not captured by the embeddings of <italic>over</italic>, which provides further insight into the kind of semantic information that can be encoded by means of (unsupervised) BERT embeddings. This can provide an interesting basis for further experimental research (e.g. testing to what extent these different operational models of meaning representation are complementary when assessed against ellicited behavioural data), as well as a discussion on how we can bring about a &#8220;greater cross-fertilization of theoretical and computational approaches&#8221; to the study of meaning (<xref ref-type="bibr" rid="B6">Boleda 2020: 2</xref>; also see, e.g., <xref ref-type="bibr" rid="B2">Baroni &amp; Lenci 2011</xref>; <xref ref-type="bibr" rid="B55">Pater 2019</xref>).</p>
</sec>
<sec>
<title>2 Background</title>
<sec>
<title>2.1 Cognitive-conceptual approaches to prepositional semantics</title>
<p>The interest in prepositional semantics in Cognitive Linguistics stems from the observation that language users are able to use a relatively small set of prepositions refer to an indefinitely large number of relations and scenes because of their cognitive ability to categorize concepts schematically (<xref ref-type="bibr" rid="B44">Kreitzer 1997</xref>). The cognitive-conceptual approach to prepositional semantics relies on two important constructs: the so-called &#8220;embodiment&#8221; of meaning, i.e. &#8220;[t]he idea that the properties of certain categories are a consequence of the nature of human biological capacities and of the experience of functioning in a physical and social environment&#8221; (<xref ref-type="bibr" rid="B45">Lakoff 1987: 12</xref>), and the notion of image schemas, which can be considered as condensed, schematic, recurring patterns of perceptual experience (e.g. <xref ref-type="bibr" rid="B54">Oakley 2010</xref>; <xref ref-type="bibr" rid="B25">Gibbs et al. 1994</xref>).</p>
<p>Some key publications in developing the notion of image schemas and embodiment, and integrating those notions into the discussion of meaning representation, are Brugman (<xref ref-type="bibr" rid="B8">1988</xref>) and Lakoff (<xref ref-type="bibr" rid="B45">1987</xref>). In their analyses of <italic>over</italic>, Brugman and Lakoff distinguish a vast number of distinct image schemas, all of which map onto a distinct &#8216;sense&#8217; of the preposition. The image schema underlying an example such as <italic>Devi lives over the hill</italic>, for instance, conveys a static horizontal spatial configuration, whereby the focal point or &#8220;trajector&#8221; (TR), Devi, is positioned on the other side of the &#8220;landmark&#8221; (LM), the hill. This schema is different from the one underlying an example such as <italic>The helicopter hung over the hill</italic>, which conveys a vertical spatial configuration in which the helicopter (TR) is positioned above the hill (LM). The schema furthermore differs from those underlying examples such as <italic>Devi walks over the hill</italic> or <italic>The helicopter flies over (the hill)</italic>, which involve a (horizontal) path, and so on.</p>
<p>Yet, while they evoke different spatial configurations, the image schemas underlying these examples are still connected to one another, as humans are able to recognize general similarities between abstract image schemas (as demonstrated experimentally by, for instance, <xref ref-type="bibr" rid="B25">Gibbs et al. (1994)</xref>). Thus, a complex yet structured &#8216;network&#8217; of linked senses is formed. Such networks, often termed &#8216;lexical networks&#8217; or &#8216;polysemy networks&#8217;, comprise of nodes which are situated at varying distances from one another, and are centered around a primary sense or prototype (<xref ref-type="bibr" rid="B45">Lakoff 1987</xref>; <xref ref-type="bibr" rid="B62">Rice 1996</xref>). At their core, prepositions are spatial expressions, but the general human ability to apply metaphorical and analogical reasoning allows them to extend the use of prepositions to embody the non-physical domain ubiquitously (<xref ref-type="bibr" rid="B44">Kreitzer 1997: 317</xref>; <xref ref-type="bibr" rid="B48">Lee 1998: 334</xref>; <xref ref-type="bibr" rid="B62">Rice 1996: 135</xref>; <xref ref-type="bibr" rid="B63">Rice 1999: 227</xref>), as demonstrated by examples such as <italic>Devi works over the weekend</italic> (embodiment of time) and <italic>Devi is over her ex-boyfriend</italic> (embodiment of mental state).</p>
<p>As each small modification to an image schema is mapped onto a discrete sense category, the meticulous and comprehensive accounts set out by Brugman and Lakoff are sometimes called the &#8220;full-specification&#8221; approach. In the case of Lakoff (<xref ref-type="bibr" rid="B45">1987</xref>), the full-specification approach led to a fine-grained overview of 24 senses of <italic>over</italic>, which are connected in a sizable polysemy network. In later work, Lakoff&#8217;s proposal was criticized amply for the fact that it leads to a virtually unconstrained number of sense categories, and for lacking methodological rigour (e.g. <xref ref-type="bibr" rid="B62">Rice 1996</xref>; <xref ref-type="bibr" rid="B44">Kreitzer 1997</xref>; <xref ref-type="bibr" rid="B66">Sandra &amp; Rice 1995</xref>; <xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>). This led Sandra (<xref ref-type="bibr" rid="B65">1998: 361</xref>) to coin the term &#8220;polysemy fallacy&#8221; in reference to &#8220;the tendency to look for polysemy even when there is no evidence for it&#8221;.</p>
<p>Subsequent proposals, then, set out ways to tackle the apparent lack of a principled procedure to determine the number of distinct (sub)senses. Two notable examples are the proposal of Kreitzer (<xref ref-type="bibr" rid="B44">1997</xref>), and the &#8220;principled polysemy&#8221; approach advocated by Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>). Drawing strongly on the spatial information encoded in linguistic expression, Kreitzer (<xref ref-type="bibr" rid="B44">1997: 308</xref>) defines a prepositional sense as &#8220;a class of uses sharing a unique relational level image schema&#8221;. In the case of <italic>over</italic>, Kreitzer argues that only three such schemata can be distinguished: (1) a static relation between two points on a vertical axis (<italic>over</italic><sub>1</sub>), (2) a dynamic relation involving a path schema (<italic>over</italic><sub>2</sub>), and (3) a static relation where one point occludes the other (<italic>over</italic><sub>3</sub>):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(1)</td>
<td>The painting hung over the fireplace.</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(2)</td>
<td>The cat jumped over the fence.</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(3)</td>
<td>The mask is over my face.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These three relational schemata are also applicable to non-spatial domains, which, Kreitzer explains, are consistently conceptualized in terms of spatial image schemata: the use of <italic>over</italic> in <italic>I finally got over that relationship</italic> (indicating a path obstructed by an obstacle), for instance, can be motivated by the dynamic schema underlying <italic>over</italic><sub>2</sub>, whereas <italic>over</italic> in <italic>The box is over six feet tall</italic> (indicating excess) is motivated by <italic>over</italic><sub>1</sub>. Yet, as pointed out by Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001: 729</xref>), Kreitzer does not motivate the existence of those three relational image schemas in light of each other, as he &#8220;makes no attempt to account for how <italic>over</italic><sub>1</sub> could give rise to <italic>over</italic><sub>2</sub> and <italic>over</italic><sub>3</sub> respectively&#8221;. Additionally, many senses touched on by Lakoff (<xref ref-type="bibr" rid="B45">1987</xref>) are simply ignored in Kreitzer&#8217;s account.</p>
<p>Addressing these issues, Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>) devised a more encompassing proposal based on slightly different principles. More specifically, they argue that senses can be considered distinct if (and only if) under the following criteria: (i) First, assuming that the primary sense of the preposition involves &#8220;a particular spatial relation between a TR and an LM&#8221;, the distinct sense &#8220;must involve a meaning that is not purely spatial in nature&#8221;, or &#8220;the spatial configuration between the TR and LM is changed vis-a-vis the other senses associated with a particular preposition&#8221; (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 731</xref>). (ii) Second, the sense must exist in examples where it cannot be inferred from the combination of another sense and encyclopedic or contextual knowledge (i.e., they must be instantiated in semantic memory; <xref ref-type="bibr" rid="B22">Evans (2005)</xref>). Following these criteria, there is no reason to assume that the use of <italic>over</italic> involves two distinct and separately stored senses in <italic>The helicopter hovered over the hill</italic> and <italic>The helicopter flew over the hill</italic>, as both examples involve a spatial relation in which the TR (<italic>the helicopter</italic>) is located above the LM (<italic>the hill</italic>). Furthermore, the difference in stativity/dynamicity of the scene can simply be inferred from the lexical verb (<italic>hover</italic> vs. <italic>fly</italic>). As such, Tyler and Evans effectively constrain the full-specification network to a more digestible size.</p>
<p>An issue that remains, however, is that there is still no objective, measurable means of determining the global structure of the polysemy network. Besides further attempts to establish which sense constitutes the core node or prototype (<xref ref-type="bibr" rid="B66">Sandra &amp; Rice 1995</xref>; <xref ref-type="bibr" rid="B62">Rice 1996</xref>; <xref ref-type="bibr" rid="B27">Gilquin &amp; McMichael 2018</xref>; <xref ref-type="bibr" rid="B53">Newman 2011: 137</xref>), there is still much room for discussion regarding the position of derived nodes. To illustrate the issue, we can consider the multitude of possible derivation pathways of the repetitive sense of <italic>over</italic>, as in <italic>She sang the same song over (and over)</italic> (<xref ref-type="bibr" rid="B72">Tyler &amp; Evans 2003: 105&#8211;106</xref>). First, based on their comparable, cyclical image schemas, the repetitive sense can be connected to reflexive uses of <italic>over</italic> (e.g. <italic>She turned the page over/The vase tipped over</italic>). Second, it is possible that repetitive <italic>over</italic> marks an iterative trajectory, in which case the sense could be derived from cases where <italic>over</italic> marks the end of a linear temporal trajectory or process (e.g. <italic>The race is over</italic>). A third possibility is that the repetitive sense constitutes a conceptual blend of reflexivity and trajectory completion, a notion which may equally apply to many other derived senses.</p>
<p>In their accounts, Tyler &amp; Evans choose to remain agnostic on the matter, explaining that &#8220;language does not function like a logical calculus which would allow us to &#8230; establish absolutely a single, precise derivation for each sense&#8221; (<xref ref-type="bibr" rid="B72">Tyler &amp; Evans 2003: 62</xref>). This is not to say that &#8216;anything goes&#8217;, but rather that there is a delimited set of general principles or paths of derivation which may individually or simultaneously give rise to derived senses, and different individuals may draw different connections between senses, if they draw any such connections at all (<xref ref-type="bibr" rid="B47">Langacker 2010: Ch.10</xref>). Yet, even so, the agnostic position is somewhat unsatisfactory if one is interested in, for instance, comparing the general probability of multiple derivational paths across individuals, or even across time. Such queries will remain difficult to address in absence of methods that enable researchers to quantify and measure the degree of similarity between sense categories. The further integration of distributional semantic models into Cognitive Linguistics, then, can at least partially be linked to the research community&#8217;s growing desire to approach the study of polysemy (and synonymy) in a more rigidly corpus-driven and measurable way (<xref ref-type="bibr" rid="B32">Gries &amp; Divjak 2009</xref>; <xref ref-type="bibr" rid="B53">Newman 2011</xref>).</p>
</sec>
<sec>
<title>2.2 Advances in Distributional Semantic Models</title>
<p>At its core, the distributional approach conceptualizes the meaning of a word (or, more generally, of constructions) as a function of its lexical and grammatical context, and as such, meaning can be approached statistically (<xref ref-type="bibr" rid="B70">Turney &amp; Pantel 2010</xref>; <xref ref-type="bibr" rid="B6">Boleda 2020</xref>). Statistical approaches to meaning have a long tradition in corpus linguistics, with functional-semantic classifications into distinct usages being increasingly based on explicit, automatically detectable contextual cues as corpora grew increasingly large.</p>
<p>An interesting observation made by Heylen et al. (<xref ref-type="bibr" rid="B38">2015</xref>) concerns the statistical-manual hybridity of the corpus linguistic tradition. As an example, they take the &#8220;British tradition in corpus linguistics&#8221;, in which lexical collocations and syntactic patterns are employed to capture or approximate word meaning, while the classification of these meanings into categories is conducted manually. By contrast, the more recent application of &#8220;Behavioral Profiles&#8221; (<xref ref-type="bibr" rid="B31">Gries 2006</xref>; <xref ref-type="bibr" rid="B32">Gries &amp; Divjak 2009</xref>), for instance, presents a means of statistically automating the classification by means of hierarchical cluster analysis (or correspondence analysis in <xref ref-type="bibr" rid="B29">Glynn (2010)</xref>). In such cases, a set of tokens can be annotated along a number of variables or dimensions, such as the type of trajector (TR), landmark (LM), as illustrated in <bold><italic><xref ref-type="table" rid="T1">Table 1</xref></italic></bold> from Newman (<xref ref-type="bibr" rid="B53">2011</xref>).</p>
<table-wrap id="T1">
<label>Table 1</label>
<caption>
<p>Example data set adapted from Newman (<xref ref-type="bibr" rid="B53">2011: Table 4</xref>).</p>
</caption>
<table>
<tr>
<th colspan="7"><hr/></th>
</tr>
<tr>
<th align="left" valign="top">context variables of &#8216;over&#8217;</th>
<th align="left" valign="top">dynamicity</th>
<th align="left" valign="top"><italic>TR</italic></th>
<th align="left" valign="top">TR_concrete</th>
<th align="left" valign="top">TR_animate</th>
<th align="left" valign="top">LM</th>
<th align="left" valign="top">&#8230;</th>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>over_1</italic></td>
<td align="left" valign="top">dynamic</td>
<td align="left" valign="top">PERSON</td>
<td align="left" valign="top">concrete</td>
<td align="left" valign="top">animate</td>
<td align="left" valign="top">PLACE</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>over_2</italic></td>
<td align="left" valign="top">stative</td>
<td align="left" valign="top">THING</td>
<td align="left" valign="top">concrete</td>
<td align="left" valign="top">non-animate</td>
<td align="left" valign="top">PLACE</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>over_3</italic></td>
<td align="left" valign="top">dynamic</td>
<td align="left" valign="top">EVENT</td>
<td align="left" valign="top">abstract</td>
<td align="left" valign="top">non-animate</td>
<td align="left" valign="top">TIME</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">&#8230;</td>
<td align="left" valign="top">&#8230;</td>
<td align="left" valign="top">&#8230;</td>
<td align="left" valign="top">&#8230;</td>
<td align="left" valign="top">&#8230;</td>
<td align="left" valign="top">&#8230;</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
</table>
</table-wrap>
<p>Such data frames can subsequently be converted into a table with numeric infomation (e.g. the relative frequency of each example with each label), which can then be used as input for statistical analysis (<xref ref-type="bibr" rid="B34">Gries &amp; Otani 2010</xref>). Focusing only on Trajector-Landmark combinations found in the ICE-GB corpus, Newman (<xref ref-type="bibr" rid="B53">2011</xref>) ultimately identifies seven statistically distinct uses of <italic>over</italic> by means of a Hierarchical Configural Frequency Analysis (HCFA):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(4)</td>
<td>[<sc>AMOUNT</sc>] <italic>over</italic>&#160;<sc>AMOUNT</sc>: &#8230;over 700 farms still cannot sell their meat for human consumption</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(5)</td>
<td><sc>EVENT</sc>&#160;<italic>over</italic>&#160;<sc>TIME</sc>: the blood pressure &lt;unclear-words&gt; at such a level after repeat measurements over a considerable period of time sometimes as long as six months</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(6)</td>
<td><sc>THING</sc>&#160;<italic>over</italic>&#160;<sc>PLACE</sc>: a minute on each side on high and then 5 minutes over a low flame will do it</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(7)</td>
<td><sc>PSYCH-STATE</sc>&#160;<italic>over</italic>&#160;<sc>STIMULUS</sc>: In view of the furore over the transmission of news from the Falklands</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(8)</td>
<td><sc>STATE</sc>&#160;<italic>over</italic>&#160;<sc>DEPENDENT ENTITY</sc>: Abortion is the right of a woman over her own body</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(9)</td>
<td><sc>COMMUNICATION</sc>&#160;<italic>over</italic>&#160;<sc>INSTRUMENT</sc>: When digital data are transmitted over a single parallel interface there is no crosstalk between the codes</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(10)</td>
<td><sc>ATTRIBUTE</sc>&#160;<italic>over</italic>&#160;<sc>STANDARD ITEM</sc>: It offered many advantages over other systems including rapid action</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The appeal of this approach, according to Newman (<xref ref-type="bibr" rid="B53">2011: 541&#8211;542</xref>), is that it &#8220;offers a systematic corpus-based procedure&#8221;, which is &#8220;strongly grounded in facts of usage, complementing any other (intuition-based or experimentally based) methods the researcher might employ&#8221;.</p>
<p>Still, the selection and annotation of the variables (which have been selected and defined by the analyst) is predominantly manual. As a &#8220;logical extension of the statistical state-of-art&#8221; (<xref ref-type="bibr" rid="B38">Heylen et al. 2015: 154</xref>), then, Semantic Vector Space Models were introduced, in which all aspects of semantic analysis are approached statistically. In such models, the contextual properties that are fed into statistical classification models are no longer manually annotated features, but automatically generated numeric representations of syntactic and lexical co-occurence patterns.</p>
<p>Because the distributional approach to meaning is based on a relatively simple, and concrete premise, one may be tempted to assume that studies adopting this approach are highly comparable, if not identical in how they operationalize and model meaning. This would, however, be a mistaken assumption. It would be far beyond the scope of the present paper to survey the many different ways in which &#8216;meaning as context&#8217; has been operationalized (for such a survey, one may consult Turney &amp; Pantel (<xref ref-type="bibr" rid="B70">2010</xref>), Lenci (<xref ref-type="bibr" rid="B50">2018</xref>), Boleda (<xref ref-type="bibr" rid="B6">2020</xref>), or specifically for deep learning based models, <xref ref-type="bibr" rid="B78">Young et al. (2018)</xref>). However, to clarify the model choice in the present study, a brief discussion of two relatively recent developments is warranted. This concerns (i) the rise of models operating with contextualized (or, rather, token-based) semantic vectors, and (ii) the rise of context-predicting models (also known as &#8216;neural language models&#8217;) that create semantic vectors often referred to as &#8216;embeddings&#8217;.</p>
<sec>
<title>2.2.1 Semantic vectors: type vs. token</title>
<p>A first development of note is the gradual turn from models that produce vectors of word types, to models that are able to create token-based (or &#8216;contextualized&#8217;) vectors. The distinction between type-based and token-based is not so much one of whether or not the resulting vector representations include contextual information &#8211; this is the case for both type-level as well as token-level vectors &#8211; but whether or not all contextual occurences of a single word are conflated into a single vector representation.</p>
<p>Type-based models work from the assumption that a word has a single, constant, &#8216;core&#8217; meaning (which can be understood as a prototype, cf. <xref ref-type="bibr" rid="B21">Erk &amp; Pad&#243; (2010: 92)</xref>), thus representing a &#8216;lumped&#8217; approach to meaning representation. Given a number of examples involving the words <italic>cat, mouse</italic>, a type-based model will provide a single numeric representation for all of the context in which these words occur (see <bold><italic><xref ref-type="table" rid="T2">Table 2</xref></italic></bold>).</p>
<table-wrap id="T2">
<label>Table 2</label>
<caption>
<p>Example of contextual input in (based on lexical co-occurrence frequencies) for cat and mouse using a word type representation.</p>
</caption>
<table>
<tr>
<th colspan="7"><hr/></th>
</tr>
<tr>
<th align="left" valign="top">TYPE representation</th>
<th align="left" valign="top"><italic>food</italic></th>
<th align="left" valign="top"><italic>purr</italic></th>
<th align="left" valign="top"><italic>paws</italic></th>
<th align="left" valign="top"><italic>keyboard</italic></th>
<th align="left" valign="top"><italic>computer</italic></th>
<th align="left" valign="top">&#8230;</th>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>cat</italic></td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">2</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>mouse</italic></td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(11)</td>
<td>The <bold>cat</bold> ate some food and purred.</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(12)</td>
<td>Do not pet the paws of a <bold>cat</bold> unless it purrs.</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(13)</td>
<td>The <bold>mouse</bold> held some food between its paws.</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(14)</td>
<td>I bought an external <bold>mouse</bold> and keyboard for my computer.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For a word such as <italic>mouse</italic>, for instance, the contextual information that suggests it is an animal would therefore be conflated with contextual information typical of the object. The problem with such aggregated vector representations is that they may render unsatisfactory or problematic vector representations in cases of polysemy, and, unarguably even more so, in cases of homonymy (<xref ref-type="bibr" rid="B21">Erk &amp; Pad&#243; 2010</xref>; <xref ref-type="bibr" rid="B17">Desagulier 2019</xref>; <xref ref-type="bibr" rid="B16">De Pascale 2019</xref>; &#8216;meaning conflation deficiency&#8217; in Camacho-Collados &amp; Pilehvar).</p>
<p>In response to this issue, models that generate token-specific vector representations were developed. These token-based distributional models &#8211; in which individual vectors are assigned to, for instance, the two different examples of <italic>mouse</italic> as in <bold><italic><xref ref-type="table" rid="T3">Table 3</xref></italic></bold> &#8211; are better equipped to handle the complex internal semantic structure of words, and, hence, are naturally better suited for specific NLP tasks such as word sense disambiguation (see, e.g. ELMo (<xref ref-type="bibr" rid="B58">Peters et al. 2018a</xref>), as well as the model described in, e.g., <xref ref-type="bibr" rid="B38">Heylen et al. (2015)</xref>). Because the aim of the present study is precisely to home in on the differences and similarities between different uses of a single preposition, it evidently employs a distributional model that produces vector presentations at the token level.</p>
<table-wrap id="T3">
<label>Table 3</label>
<caption>
<p>Example of contextual input in (based on lexical co-occurrence frequencies) for <italic>cat</italic> and <italic>mouse</italic> using a word token representation.</p>
</caption>
<table>
<tr>
<th colspan="7"><hr/></th>
</tr>
<tr>
<th align="left" valign="top">TOKEN representation</th>
<th align="left" valign="top"><italic>food</italic></th>
<th align="left" valign="top"><italic>purr</italic></th>
<th align="left" valign="top"><italic>paws</italic></th>
<th align="left" valign="top"><italic>keyboard</italic></th>
<th align="left" valign="top"><italic>computer</italic></th>
<th align="left" valign="top">&#8230;</th>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>cat_1</italic></td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>cat_2</italic></td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>mouse_1</italic></td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">&#8230;</td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>mouse_2</italic></td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">0</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top"></td>
</tr>
<tr>
<td colspan="7"><hr/></td>
</tr>
</table>
</table-wrap>
</sec>
<sec>
<title>2.2.2 Semantic vectors: count vs. predict</title>
<p>Using the terminology employed in Baroni et al. (<xref ref-type="bibr" rid="B3">2014</xref>), I wish to point out that a distinction can be made between &#8216;count models&#8217;, and &#8216;predict(ive) models&#8217; (also see &#8216;explicit&#8217; and &#8216;implicit&#8217; models in Dubossarsky et al. (<xref ref-type="bibr" rid="B20">2017: 1136</xref>)). Count models represent, in a sense, the most straightforward way of operationalizing the distributional hypothesis, in that they make use of numerical vectors that are essentially based on co-occurrence counts (for an accessible explanation of how such vectors are constructed for word types and word tokens, see, for instance, Heylen et al. (<xref ref-type="bibr" rid="B38">2015</xref>) and <xref ref-type="bibr" rid="B39">Hilpert &amp; Correia Saavedra (2017)</xref>). Still, describing count models as such is a severe simplification, as more than often the vectors are optimized in some way (e.g. by changing context window sizes, reweighting function words, leaving out function words, applying dimensionality reduction, etc.).</p>
<p>By contrast, context-predicting models (yet again a cover-term for an extremely varied group of models, including weighted bag-of-words, to more syntactically informed variations, with new types of model architectures being added continuously) are designed to approach the construction of semantic vectors from a training-based angle: &#8220;Instead of first collecting context vectors and then reweighting these vectors based on various criteria, the vector weights are directly set to optimally predict the contexts in which the corresponding words tend to appear&#8221; (<xref ref-type="bibr" rid="B3">Baroni et al. 2014: 238</xref>). In other words, predictive models construct vectors as part of a learning task, which, to some degree, eliminates the vector transformation and optimization process. This, in addition to the performance improvements observed in a range of NLP tasks compared to count models, is why the relatively recent emergence predictive models is often portrayed as an attractive advancement (<xref ref-type="bibr" rid="B3">Baroni et al. 2014</xref>). This is, however, not to say that predictive models involve absolutely no parameter tuning &#8211; and it has been suggested that, given comparable settings and tuning, the vectors created with count models are as effective as the embeddings yielded by predictive models (<xref ref-type="bibr" rid="B51">Levy et al. 2015</xref>). Yet, what does make predictive neural language models particularly appealing for the present study, which focuses on prepositional semantics, is that count models are generally less successful in providing useful representations of function words (e.g. <xref ref-type="bibr" rid="B12">Bullinaria &amp; Levy 2012: 7</xref>), whereas &#8220;recent neural network models do provide usable representations for them&#8221; (<xref ref-type="bibr" rid="B6">Boleda 2020: 7</xref>).</p>
</sec>
</sec>
<sec>
<title>2.3 BERT</title>
<p>In the present study, the distributional approach is represented by a single model architecture: Devlin et al. (<xref ref-type="bibr" rid="B18">2019</xref>)&#8217;s Bidirectional Encoder Representations from Transformers (BERT). First launched in November 2018, BERT quickly became the model to beat due to its impressive performance on a wide range of NLP tasks. Soon after, it also grabbed the attention of computational linguists, who were interested in determining precisely what kind of linguistic information such models acquire and capture (<xref ref-type="bibr" rid="B14">Clark et al. 2019</xref>; <xref ref-type="bibr" rid="B42">Jawahar et al. 2019</xref>; <xref ref-type="bibr" rid="B1">Alishahi et al. 2019</xref>).</p>
<p>In a nutshell, BERT is a deep contextualized model based on a particular type of neural architecture, called &#8220;the Transformer&#8221;, which is entirely based on so-called &#8220;attention mechanisms&#8221; (<xref ref-type="bibr" rid="B75">Vaswani et al. 2017</xref>). A context-predicting model, BERT has been pre-trained on approximately 3.3 billion words (800 million words taken from the BooksCorpus, and 2.5 billion words from English Wikipedia) of unlabelled data over a masked word prediction task (in which the objective is to predict randomly masked input tokens based only on the context in which they occur) and a next sentence prediction task (so that the model will also understand sentence relationships).</p>
<p>Like other Transformers, BERT consists of multiple layers (or &#8216;transformer blocks&#8217;), all of which contain multiple self-attention heads which behave similarly within their layer. The smallest pre-trained model, called BERT<sub>base</sub>, consists of 12 layers with 12 attention heads, whereas the larger model, called BERT<sub>large</sub>, consists of 24 layers with 16 attention heads. Each of these layers captures the <italic>n</italic> tokens in the input sentence (or rather &#8216;sequence&#8217;, as the input need not correspond with what linguists have traditionally defined as a sentence) in compressed numerical vector representations or &#8216;embeddings&#8217;.</p>
<p>The attention heads within BERT&#8217;s layers have been probed for the linguistic phenomena they capture. This revealed that particular heads capture syntactic relations (e.g. valency patterns and dependency relations), while others perform well at coreference resolution (<xref ref-type="bibr" rid="B14">Clark et al. 2019</xref>) &#8211; which is remarkable given that the model has not received any explicit input about syntax or coreference. This &#8220;syntax-aware attention&#8221; (<xref ref-type="bibr" rid="B14">Clark et al. 2019</xref>) may be why BERT is succesful the downstream NLP tasks it has been employed in (cf. <xref ref-type="bibr" rid="B57">Peters et al. 2018b</xref>). Finally, it is important to note that the different layers (and accompanying attention heads) perform slightly differently on different tasks. In various sources, the second-to-last layer (or a concatenation of the last four layers) is suggested to perform best on token-level tasks such as word sense disambiguation (e.g. <xref ref-type="bibr" rid="B18">Devlin et al. 2019</xref>; <xref ref-type="bibr" rid="B77">Wiedemann et al. 2019</xref>), but many applications also operate with the final hidden layer (e.g. <xref ref-type="bibr" rid="B41">Huang et al. 2019</xref>; <xref ref-type="bibr" rid="B5">Blevins &amp; Zettlemoyer 2020</xref>).</p>
<p>With respect to linguistic investigation into polysemy and sense disambiguation, the contextualized embeddings produced by BERT have thus far not been explored. One reason may be that neural models have grown into increasingly intransparant systems (<xref ref-type="bibr" rid="B52">Linzen et al. 2019: iii</xref>), making linguists more reluctant to rely on them for lingusitic analysis. Still, it is worth investigating to what extent they could be employed as analytic tools in linguistic research, as neural language models like (but consistently outperformed by) BERT have already been shown to capture very nuanced aspects of meaning, and they even seem to provide usable representations for function words (see Boleda (<xref ref-type="bibr" rid="B6">2020: 18</xref>), in reference to <xref ref-type="bibr" rid="B58">Peters et al. (2018a)</xref>). Furthermore, a model such as BERT also unites the strengths of different types of token-based distributional methods. First, the fact that the model is syntax-aware agrees with the cognitive-linguistic (and constructionist) view that differences in syntactic structures reflect differences in meaning (<xref ref-type="bibr" rid="B46">Langacker 1991</xref>; <xref ref-type="bibr" rid="B30">Goldberg 1995</xref>). As such, its syntax-awareness sets BERT apart from bag-of-words approaches to contextualized vectors (e.g. <xref ref-type="bibr" rid="B38">Heylen et al. 2015</xref>), and thus makes it more akin to, for example, the Behavioural Profiles approach. Second, the application of BERT to the study of polysemy does not involve any manual annotation, and neither does it involve making an a priori selection of syntactic features to be included, making it a fully data-driven approach to the question at hand.</p>
</sec>
</sec>
<sec sec-type="methods">
<title>3 Data and Methodology</title>
<p>In the present study, BERT<sub>base</sub> has been used to create contextualized embeddings for all occurrences of <italic>over</italic> in the final decade of the Corpus of Historical American English (COHA, 2000&#8211;2010). In total, embeddings were created for 39,834 tokens of <italic>over</italic> using the Spacy implemenation of BERT<sub>base</sub> (which, at the time the analysis was conducted, only offered access to the final hidden layer). The embeddings of the target tokens were created with a context window set to 20 words preceding and 20 words following <italic>over</italic>. In principle, the performance of the model in the task at hand could still be improved by experimenting with different hyperparameter settings or by fine-tuning the model to specific tasks or corpus data, but no such operations were undertaken.</p>
<p>Of the 39,834 tokens, 808 examples were manually annotated by two human annotators, following the sense description in Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>). Note that the proposal by Tyler &amp; Evans does not consitute a &#8216;gold standard&#8217;, as its cognitive reality remains to be tested against ellicited, experimental data. Yet, their proposal was chosen as a point of comparison because (i) it is well-documented, (ii) is firmly grounded in and motivated by linguistic theory, (iii) and presents the most comprehensive assessment of all possible senses of <italic>over</italic> since Lakoff (<xref ref-type="bibr" rid="B45">1987</xref>). Furthermore, because their proposal focusses not only on motivating the number of distinct senses, but also on motivating the connections between those senses by foregrounding the importance of image schemas, it lends itself well to an assessment of whether such imagistic information is encoded in corpus data and captured by word embeddings.</p>
<p>In total, 16 different sense categories were distinguished. Following the example of Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>), the categories are given a label that corresponds with their status as a discrete sense (i.e., 1, 3, etc.) and subsense (i.e., A, B, etc.). Note that the 808 examples constitute a stratified sample: first, a random sample of 300 tokens was manually annotated. Subsequently, the sample was expanded with further examples until each sense category was represented by at least 10 tokens. The token frequencies per sense category are listed in <bold><italic><xref ref-type="table" rid="T4">Table 4</xref></italic></bold>.</p>
<table-wrap id="T4">
<label>Table 4</label>
<caption>
<p>Token Frequencies per category.</p>
</caption>
<table>
<tr>
<th colspan="4"><hr/></th>
</tr>
<tr>
<th align="left" valign="top">Category</th>
<th align="left" valign="top">Tokens</th>
<th align="left" valign="top">Category</th>
<th align="left" valign="top">Tokens</th>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>1. Protoscene &#8216;above&#8217;</italic></td>
<td align="left" valign="top">152</td>
<td align="left" valign="top"><italic>4A. Focus-of-attention</italic></td>
<td align="left" valign="top">62</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>2A. On-the-other-side</italic></td>
<td align="left" valign="top">122</td>
<td align="left" valign="top"><italic>5A. More</italic></td>
<td align="left" valign="top">39</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>2B. Excess &#8216;beyond&#8217;</italic></td>
<td align="left" valign="top">24</td>
<td align="left" valign="top"><italic>5B. Control</italic></td>
<td align="left" valign="top">50</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>2C. Completion</italic></td>
<td align="left" valign="top">37</td>
<td align="left" valign="top"><italic>5C. Preference</italic></td>
<td align="left" valign="top">33</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>2D. Transfer</italic></td>
<td align="left" valign="top">35</td>
<td align="left" valign="top"><italic>6. Reflexive</italic></td>
<td align="left" valign="top">31</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>2E. Time span</italic></td>
<td align="left" valign="top">50</td>
<td align="left" valign="top"><italic>6A. Repetition</italic></td>
<td align="left" valign="top">32</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>3. Covering</italic></td>
<td align="left" valign="top">49</td>
<td align="left" valign="top"><italic>7. Communication line</italic></td>
<td align="left" valign="top">23</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><italic>4. Examining</italic></td>
<td align="left" valign="top">32</td>
<td align="left" valign="top"><italic>8. Hangover</italic></td>
<td align="left" valign="top">11</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="left" valign="top"><italic>Indeterminate</italic></td>
<td align="left" valign="top">26</td>
</tr>
<tr>
<td colspan="4"><hr/></td>
</tr>
</table>
</table-wrap>
<p>Between the two human annotators, inter-rater agreement was found to be very good (Fleiss&#8217; Kappa = 0.867). In the statistical analyses presented below, 26 were examples were excluded because they were considered indeterminate between multiple categories. Further information on the sense categories is provided in Section 3.1.</p>
<p>Ultimately, the sense categorization proposed by Tyler and Evans was created in response to models that are too fine-grained, and hence lack what Tyler &amp; Evans consider to be meaningful, principled abstractions. Thus, the question we are in fact asking is to what extent these abstractions are also &#8216;meaningful&#8217; to models such as BERT, which approach prepositional meaning by compressing contextual data. To address this question, I adopt a procedure based on the Varying Abstraction Model (<xref ref-type="bibr" rid="B74">Vanpaemel &amp; Storms 2008</xref>). Originally, the Varying Abstraction Model (VAM) was designed in response to the debate in psychology on how the classification accuracy of exemplar-based models (involving no abstraction) compares to that of prototype models (involving complete abstraction) as well as models involving intermediate levels of abstraction. The procedure adopted here is based on the <italic>k</italic>-means variant of the VAM (<xref ref-type="bibr" rid="B76">Verbeemen et al. 2005</xref>), where a particular level of abstraction is operationalized by the degree to which category members are clustered.</p>
<p>The VAM conducts a series of evaluation tasks, where it predicts the category label of unseen test tokens against a manually assigned label. The series start with the prediction of the category label of an unseen test token based on its nearest neighbour embedding in a labelled training set. At this level, none of the token embeddings in the training set have been clustered, which, one could argue, means that the number of &#8216;clusters&#8217; <italic>k</italic> equals the number of tokens <italic>n</italic> in the training set (<italic>k</italic> = <italic>n</italic>). This is also called the &#8216;exemplar level&#8217;. In other word sense disambiguation studies, the performance of models such as BERT is commonly assessed solely at this level (see, e.g. <xref ref-type="bibr" rid="B77">Wiedemann et al. 2019</xref>). In the present study, however, the assessment is also taken beyond the exemplar level: the VAM will subsequently attempt the same classification task again, but instead of using all the embeddings of the concrete tokens in the training set as a reference set, it will create a slightly higher level of &#8216;abstraction&#8217; by merging the embeddings of some concrete tokens of the same sense category into a slightly more schematic, averaged representation. At every step of the VAM procedure, an increasing number of token embeddings are merged, until all embeddings of all training tokens that belong the same sense category are merged into a single averaged embedding. A schematic representation of the different steps of abstraction is presented in <bold><italic><xref ref-type="fig" rid="F1">Figure 1</xref></italic></bold>.</p>
<fig id="F1">
<label>Figure 1</label>
<caption>
<p>Schematic representation of abstraction continuum. At the lowest level of abstraction, the items in the training set that can be used for classifiying an unseen token are the embeddings of concrete tokens in that set. At intermediate levels, the number of items that can be used to classify an unseen token is gradually reduced, as an increasing number of token embeddings are merged into an averaged embedding. When complete abstraction is reached, all items of the same category are merged into a single, averaged &#8216;sense embedding&#8217;.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60708/"/>
</fig>
<p>At the highest level of abstraction, then, the classification task of the unseen test tokens is attempted by means of a clustered or averaged representation of all embeddings assigned to that category. This averaged embedding can hence be thought of as a &#8216;contextualized sense embedding&#8217;. Because I consider the 16 sense categories of the principled polysemy model (described in Section 3.1) to represent the highest level of abstraction, the number of clustered representations by means of which classification is attempted at this level is 16 (<italic>k</italic> = <italic>16</italic>). Note that the level of abstraction could be increased further by reducing the number of clusters to 8 (to attempt the classification task using clustered representations of, for instance all tokens labelled as sense 5A, 5B and 5C), but no such actions were undertaken in this study.</p>
<p>The series of classification tasks (from <italic>k</italic> = <italic>n</italic> to <italic>k</italic> = <italic>16</italic>) is evaluated against a test set (20% of the data; 100 iterations per level; see <bold><italic><xref ref-type="fig" rid="F2">Figure 2</xref></italic></bold>), and will be expressed in an accuracy score (F<sub>1</sub>-score, between 0 and 1, with 1 representing perfect accuracy). The resulting series of classification accuracy scores allows us to assess the following: if the sense classification task goes well at the lowest level of abstraction (the exemplar level), we find that the contextual information encoded in the BERT embeddings of <italic>over</italic> encodes and captures local similarities between concrete tokens of the same sense category. As the level of abstraction increases, the classification task will involve classifiying unseen tokens not by means of other, concrete tokens, but by means of averaged contextual representations of multiple tokens that have been assigned the same label. In other words, the model will attempt the classification of unseen tokens by means of contextual representations that are decreasingly concrete and increasingly schematic (that is, representing abstractions over multiple tokens in the same sense category). If classification accuracy of the unseen test tokens remains high when all training tokens of the same category are averaged into a &#8216;sense embedding&#8217;, we find that these abstract contextual representations are helpful tools to categorize new, unseen tokens. In that case, we could say that the &#8216;meaningful abstractions&#8217; or sense categories proposed by Tyler and Evans also make sense in terms of the contextual information encoded in BERT embeddings.</p>
<fig id="F2">
<label>Figure 2</label>
<caption>
<p>Schematic representation of VAM procedure. Starting from no clustered items, the VAM uses different configurations of clustered or averaged embeddings (which represent different levels of abstraction) in a training set (80% of the data) of the data to classify an unseen test token set (20% of the data).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60709/"/>
</fig>
<p>Besides assessing to what extent BERT embeddings can be used to distinguish the sense categories proposed in the principled polysemy model, I will also discuss the global structure of the network proposed by Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>). To discuss distances between sense categories, I use the cosine similarity between the embeddings (see, e.g., Bullinaria &amp; Levy (<xref ref-type="bibr" rid="B11">2007</xref>); Heylen et al. (<xref ref-type="bibr" rid="B38">2015</xref>); <xref ref-type="bibr" rid="B57">Peters et al. (2018b)</xref>).</p>
<sec>
<title>3.1 Sense categories</title>
<p>In what follows, I will describe the categories distinguished in Tyler &amp; Evans, illustrating them with examples from the data set. For an in-depth description of the sense categories and a full argumentation as to why these (and only these) categories have been distinguished, I refer to Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>).</p>
<sec>
<title>3.1.1 Sense 1: &#8216;above&#8217;</title>
<p>The first category contains all examples in which <italic>over</italic> signals that the TR is located above the LM. This category is considered to be the primary sense or &#8216;protoscene&#8217; from which all other senses can be derived (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 735&#8211;737</xref>). The relation expressed is an atemporal, spatial relation, where the TR is typically in close proximity to the LM. In many cases, TR is typically movable and smaller than the LM (as in (15)), but immovable (e.g. (16)) and larger (e.g. (17)) TRs occur as well.</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(15)</td>
<td>I noticed a painting hanging <bold>over</bold> the piano (COHA, 2006)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(16)</td>
<td>He was bleeding from a cut <bold>over</bold> his eye (COHA, 2003)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(17)</td>
<td>And, so that&#8217;s how I got into the apartment <bold>over</bold> the garage. (COHA, 2005)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.1.2 Sense (group) 2: A-B-C trajectory</title>
<p>Besides the protoscene, Tyler &amp; Evans also distinguish a number of derived senses. Four of these can be conceived of as a &#8216;cluster&#8217; of senses where <italic>over</italic> marks a trajectory from a starting point (A), a midpoint (B), and an endpoint (C). While not all senses in this cluster put equal focus on all points in the trajectory, the uniting factor seems to be that there is a certain linearity to the expressed relation.</p>
<p><bold>ON-THE-OTHER-SIDE-OF (2A)</bold> In examples (18) and (19), the TR is portrayed as being not above, but on the other side of the LM:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(18)</td>
<td>I&#8217;d grown up only a few hours away, <bold>over</bold> the Kentucky line. (COHA, 2007)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(19)</td>
<td>God, let this be the peak. Let us be <bold>over</bold> the mountain (COHA, 2007)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Note that the verb itself does not trigger the trajectory reading (when it does, the example will be assigned to the protoscene).</p>
<p>While the examples in (18) and (19) both function as prepositions, a large group of tokens in this category function as an adprep. In some cases, such as (20), the verb is combined with an adverbial phrase that indicates the endpoint of the trajectory. Thus, it could be argued that the combination of the verb and the endpoint adverbial already implies movement along a trajectory. The addition of <italic>over</italic>, then, seems to have a mere emphatic function. However, in other examples, such as (21), the adprep is non-optional if a trajectory is to be evoked. All adprep trajectory uses are assigned to category 2A.</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(20)</td>
<td>After he left us, he drove <bold>over</bold> to my brother Jacob &#8216;s apartment (COHA, 2006)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(21)</td>
<td>Smiling, she hurries <bold>over</bold> (COHA, 2004)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Examples such as (22), where a look is thrown at an explicit endpoint, were also assigned to category 2A:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(22)</td>
<td>He looked <bold>over</bold> at the computer. (2007, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, a number of examples assigned to this category do not refer to a spatial relation. If we consider examples such as (23), where the LM represents an obstacle or hurdle, one can metaphorically extend the use of <italic>over</italic> to non-physical obstacles (often relating to past relationships), as in (24) and (25):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(23)</td>
<td>&#8230; that old and painful relationship. But Mike had seemed okay with it, as if he was completely <bold>over</bold> Lindsey (COHA, 2009).</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(24)</td>
<td>I had a thing with her a bunch of years ago, and I guess I never got <bold>over</bold> the attraction. (COHA, 2007)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(25)</td>
<td>Memphis had gotten <bold>over</bold> her steering problems. (COHA, 2004)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>&#8216;EXCESS&#8217;: ABOVE-AND-BEYOND (2B)</bold> In this category, we find cases where the TR moves above the LM, and thus misses or exceeds a point it should not have crossed. What is key about Sense 2B, and what distinguishes it from Sense 1 and Sense 2A, is the implicature that &#8220;the LM represents an intended goal or target and that the TR moved beyond the intended or desired point&#8221; (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 749</xref>), as in (26):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(26)</td>
<td>Your article is <bold>over</bold> the page limit. (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 749</xref>)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>With such an example, it is difficult to say whether interpretation of excess is entirely &#8216;context-free&#8217; and not evoked or supported by the lexeme <italic>limit</italic>. Similarly, the &#8216;excess&#8217; implicature in examples (27) and (28) may be triggered by <italic>fault</italic> and <italic>illegally</italic>. Still, the decision to include a separate category of Sense 2B was maintained.</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(27)</td>
<td>&#8230; seldom-called violations in tennis &#8211; the foot fault. It occurs when a player&#8217;s foot brushes or goes <bold>over</bold> the baseline when serving. (COHA, 2006)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(28)</td>
<td>He&#8217;s illegally parked. His ass is <bold>over</bold> the white line. (2003, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Furthermore included in this category are cases such as (29), which portray a situation where the &#8216;missed target&#8217; is a person in line for a reward (usually in the form of a job offer or promotion, as in (30)). The implicature here is that the reward was expected or deserved, but those expectations were not fulfilled.</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(29)</td>
<td>&#8230; his monumental 1957 paper on the origins of elements, for which &#8211; to his annoyance &#8211; he was passed <bold>over</bold> for a nobel prize. (COHA, 2001)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(30)</td>
<td>&#8230; a lot of times he&#8217;ll pass <bold>over</bold> the most talented and put someone in with the biggest heart. (COHA, 2003)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>COMPLETION (2C)</bold> This category contains examples such as (31) and (32), where <italic>over</italic> indicates that something is finished or completed.</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(31)</td>
<td>My school days are finally officially <bold>over</bold>. (2007, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(32)</td>
<td>All the decisions had been made, the story was <bold>over</bold>. (2006, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This category solely contains examples where <italic>over</italic> functions as an adprep, and consistently combines with the verb <italic>be</italic>.</p>
<p><bold>TRANSFER (2D)</bold> Another category that solely includes adpreps is category 2D, which includes examples such as (33):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(33)</td>
<td>The woodsman reached in his pocket, pulled out the thirty euros, and handed the two bills <bold>over</bold> to the new man. (2003, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The only difference between these examples and examples of 2A is that some sort of transfer has taken place. However, one could argue that such a transaction is encoded by the verb used in the same construction. Consider, for instance, the example in (34):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(34)</td>
<td>The clerk lifted the bill from Peterson&#8217;s hand and took it <bold>over</bold> to the second clerk sitting at the desk (COHA, 2003).</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Here, an object is indeed moved by a clerk to another clerk, but there is no explicit indication that the object was given to the second clerk. While different in lexical material, the example in (34) is structurally identical Tyler &amp; Evans&#8217; examples of transfer (e.g. <italic>The teller handed the money over to the investigating officer</italic>). The key element that triggers the meaning of transfer is therefore perhaps not <italic>over</italic>, but the verb <italic>hand</italic>, which complicates the suggestion that &#8216;transfer&#8217; constitutes an encoded sense somewhat. Still, examples such as (33) were placed in a separate category 2D.</p>
<p>Finally, Non-physical transfers, as illustrated in (35) are also considered as instances of 2D:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(35)</td>
<td>&#8230; the London office had grown considerably in the last eight years. Boyd wouldn&#8217;t half mind taking <bold>over</bold> the running of it. (2008, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As non-physical transfers almost exclusively involve a transfer of control or authority, these examples are, at times, difficult to distinguish from examples of 5B (see below).</p>
<p><bold>TIME SPAN (2E)</bold> In a fairly large number of examples, <italic>over</italic> &#8220;mediates a temporal relation of concurrence between a process or activity and the times during which the process or activity elapses&#8221; (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 748&#8211;749</xref>), as illustrated in (36) and (37):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(36)</td>
<td>The war on witchcraft intensified <bold>over</bold> the next 200 years, sending millions of cats, not to mention humans, to their deaths. (2001, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(37)</td>
<td>Geologists and biologists before Darwin noted that the Earth and its inhabitants change <bold>over</bold> time. (2004, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Note that the question may be raised whether these examples do in fact constitute a distinct sense of <italic>over</italic>, as the temporal reading is inferable from the fact that the LM consistently involves a noun that refers to a time-related concept. The choice was made to create a separate label for these examples, but if the inferrability criterion is adhered to more strictly, these examples could perhaps be classified as instances of Sense 2A.</p>
</sec>
<sec>
<title>3.1.3 Sense 3: Covering</title>
<p>Like Lakoff, Tyler &amp; Evans also distinguish a category of examples such as (38), in which there is &#8220;an understood viewpoint from which the TR is blocking accessibility of vision to at least some part of the LM&#8221; (<xref ref-type="bibr" rid="B45">Lakoff 1987: 429</xref>). In these cases, the TR is not located above the LM from the vantage point of the viewer (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 752</xref>):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(38)</td>
<td>A ratty leather jacket gaped open to reveal a white button-front shirt <bold>over</bold> an ample but not outrageous bosom (2005, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The &#8216;covering&#8217; sense also includes a examples there is a &#8220;multiplex trajector&#8221; (<xref ref-type="bibr" rid="B45">Lakoff 1987: 428</xref>) that is scattered over the LM, as in (39), or where the TR has covered a path consisting of multiple points over the LM, as in (40):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(39)</td>
<td>He searches through the papers scattered <bold>over</bold> the desk, but finds nothing (COHA, 2005)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(40)</td>
<td>I can get into the crawlspace from my closet and climb all <bold>over</bold> the house. (COHA, 2005)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.1.4 Sense (group) 4: Proximity</title>
<p>As in sense 3, the TR is no longer necessarily positioned above the LM from the perspective of the viewer. Instead, <italic>over</italic> conveys that there is close proximity between the TR and LM. This proximity goes beyond the spatial realm, and is manifested in the attention paid by the TR to the LM.</p>
<p><bold>EXAMINING (4)</bold> The first category in Sense group 4 includes examples where the TR is examining the LM. The majority of cases involve the verb <italic>look</italic> (or near-synonyms such as <italic>glance</italic>), as in (41), but other verbs (e.g. <italic>read, go</italic>) occur as well:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(41)</td>
<td>Chad got out and walked around the truck, looking it <bold>over</bold>. (2008, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(42)</td>
<td>After studying my folder and going <bold>over</bold> the exact sequence of what to speak on, I allow myself the pleasure of flipping on the news (2004, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>FOCUS-OF-ATTENTION (4A)</bold> In the majority of examples included in this category, the LM is the focus of the TRs attention. In these examples, <italic>over</italic> is equivalent to <italic>about</italic>, and in some cases, the LM can be considered the cause of the TRs actions:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(43)</td>
<td>That debate, <bold>over</bold> how fast and how far to cut emissions, was the right battle to have (2004, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(44)</td>
<td>Ben was pictured displaying great emotion by crying <bold>over</bold> her loss. (2001, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In discussing sense 4A, Tyler &amp; Evans also mention examples such as (45):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(45)</td>
<td>John Stewart presides <bold>over</bold> Comedy Central&#8217;s The Daily Show, a blessed wedding of performer and format. (2001, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Note that, because the verb <italic>preside</italic> is used, it is also implied that the TR controls the LM. The same could also be said for examples such as (46), where the the notion of control or authority is not encoded in the lexical verb:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(46)</td>
<td>Francis watched <bold>over</bold> the boy&#8217;s education (2004, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Examples such as (45) and (46) were classified as instances of sense 4A, but it must be noted here that the distinction between these examples and examples of sense 5B, which are discussed below, is difficult to maintain.</p>
</sec>
<sec>
<title>3.1.5 Sense (group) 5: Up</title>
<p>Four further senses fall under &#8220;the <italic>up</italic> cluster&#8221;, which are suggested to derive &#8220;from construing a TR located physically higher than the LM as being vertically elevated or up relative to the LM&#8221; (<xref ref-type="bibr" rid="B71">Tyler &amp; Evans 2001: 755</xref>).</p>
<p><bold>MORE (5A)</bold> The first (and most frequently occurring) sense is 5A. In all examples in this category, <italic>over</italic> indicates that a quantity is higher than the quantity expressed in the LM:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(47)</td>
<td>Jerry and I were parents to <bold>over</bold> fifty foster kids in our thirty years of marriage. (2004, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Tyler &amp; Evans distinguish one further sense, sense 5A.1, where the TR is understood as something that is contained by, but exceeds the capacity of, the LM. The only clear example discussed is <italic>overtired</italic>, in <italic>The child was overtired and thus had difficulty falling asleep</italic>. The data set in the present paper does not include compounds with <italic>over</italic>. In a footnote, Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001: 757</xref>) explain that it is often possible to &#8220;construct a &#8216;more&#8217; conceptualization&#8221; alongside &#8220;an &#8216;excess&#8217; interpretation&#8221;. In practice, this seemed to apply to nearly all examples in the data set. As such, no distinction was made between sense 5A and 5A.1.</p>
<p>It is, in many cases, also extremely difficult to distinguish cases of &#8216;excess&#8217; as crossing a target point, or excess as exceeding an amount or capacity. Consider, for instance, the example in (48):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(48)</td>
<td>But for kids <bold>over</bold> age 5, as the portion size got larger, so did the amount they ate. (2003, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>While suitable for the &#8216;up&#8217; conceptualization, it is not inconceivable that examples such as (34) could also be classified under 2B (as time is a linear concept rather than a container). Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001: 758</xref>) also address this issue, stating that their network of senses &#8220;should be thought of as a semantic continuum, in which complex conceptualizations can draw on meanings from distinct nodes as well as the range of points between nodes, which provide nuanced semantic values&#8221;. For simplicity&#8217;s sake, the choice was made to assign all cases where a numeric threshold was exceeded to sense 5A.</p>
<p><bold>CONTROL (5B)</bold> The <italic>up</italic>-cluster further includes a category for all cases where <italic>over</italic> is used to mark that the &#8220;TR exerts influence, or control over the LM&#8221; Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001: 758</xref>), as in (49) and (50):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(49)</td>
<td>She was moved by her power <bold>over</bold> me. I would have fallen down for her any day. (2002, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(50)</td>
<td>But Nolan has final say <bold>over</bold> all personnel decisions, including the draft. (2005, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As noted earlier, there is also a sense of &#8216;control&#8217; in examples classified as 2D, where control is transferred from one party to another, and examples classified under 5B.</p>
<p><bold>PREFERENCE (5C)</bold> The final group of examples in the <italic>up</italic>-cluster convey a preference of one option, the TR, over another, the LM. In many examples, the notion of preference is encoded by the verb (e.g. <italic>prefer</italic>), but this need not be the case, as shown in (51) and (52):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(51)</td>
<td>We haven&#8217;t switched to a local pediatrician, believing irrationally in Manhattan doctors <bold>over</bold> Brooklyn doctors. (2001, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(52)</td>
<td>His name was Miguel Santiago, and he insisted on being called Miguel <bold>over</bold> Mike. (2005, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.1.6 Sense (group) 6: Reflexivity and Repetition</title>
<p>Two more senses distinguished by Tyler &amp; Evans are the reflexive use of <bold>over</bold>, as in (53), and the repetitive use, as in (54) and (55):</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(53)</td>
<td>Elaine pulls her leg back and kicks the grill. The coals fly up and out, the grill tips <bold>over</bold>. (2000, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(54)</td>
<td>Sometimes even horrible memories play <bold>over</bold> and <italic>over</italic> in my mind (2003, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(55)</td>
<td>I hope someday soon we can begin again &#8230; start <bold>over</bold>. (2002, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Examples such as (53) are categorized as Sense 6, whereas (54) and (55) are classified as Sense 6A. All instances of reflexive and repetitive <italic>over</italic> function as adpreps.</p>
</sec>
<sec>
<title>3.1.7 Other Senses</title>
<p>Besides the senses discerned by Tyler &amp; Evans, a few further categories were distinguished.</p>
<p><bold>INDETERMINATE</bold> First, when the precise sense of <italic>over</italic> in a given example was considered vague or ambiguous between multiple readings, the example was classified as &#8216;indeterminate&#8217;. Consider, for instance, example (56), in which it is unclear whether the interpretation is reflexive (&#8216;she bent/tilted forward&#8217;), or whether a trajectory is implied (&#8216;she leaned over (to us)&#8217;).</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(56)</td>
<td>She leaned <bold>over</bold> and talked with excitement. (2009, COHA)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, 49 examples were not classifiable along the categories set out above. These examples seem to fall within two categories: &#8216;means of communication&#8217; and &#8216;hangover&#8217;.</p>
<p><bold>MEANS OF COMMUNICATION</bold> The first category concerns examples such as (57) and (58), in which the TR (if expressed) is a person who uses a particular channel or means of communication (the LM). In these examples, <italic>over</italic> seems to be paraphrasable as &#8216;by means of&#8217;:</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(57)</td>
<td>You could break the news <bold>over</bold> the phone (COHA, 2005)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(58)</td>
<td>In Napster&#8217;s case the transfers took place <bold>over</bold> the internet (COHA, 2006)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>HANGOVER</bold> The second category concerns cases where <italic>over</italic> is used in an idiomatic expression with <italic>hung</italic>, indicating the (unpleasant) after-effects of excessive substance abuse, as in (59). In these cases, <italic>over</italic> is an adprep.</p>
<table-wrap>
<table content-type="example">
<tbody>
<tr>
<td>(59)</td>
<td>&#8230; that he had gotten drunk the night before and that he was still horribly hung <bold>over</bold>. (COHA, 2003)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec>
<title>4 Results</title>
<sec>
<title>4.1 Sense Distinctions</title>
<p>As a first point of enquiry, it is investigated whether BERT indeed recognises the sense categories proposed in the principled polysemy model in a relatively distinct and coherent manner. Following Kilgarriff (<xref ref-type="bibr" rid="B43">2003: 108</xref>), I define &#8216;senses&#8217; as &#8220;abstractions over clusters of word usages&#8221;. In other words, if the abstract, conceptual sense categories proposed in the principled polysemy model are recognized by BERT, we would expect to find that the geometrical distance (operationalized as the cosine distance) between the embeddings of all tokens labelled as Sense 1, for instance, is shorter than the distance between those tokens and tokens with a different category label, thus forming a cluster.</p>
<p>To visualize the local embedding clusters and the global positioning of those clusters relative to one another, a two-dimensional representation of the token embeddings was created based on the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm for dimensionality reduction of high-dimensional data (<xref ref-type="bibr" rid="B73">van der Maaten &amp; Hinton 2008</xref>). In <bold><italic><xref ref-type="fig" rid="F3">Figure 3</xref></italic></bold>, each token is represented by a dot, which has been coloured according to its manually assigned label. Note that the embeddings were created solely based on the contextual information surrounding <italic>over</italic>, and that the manual labels were assigned separately. The distributional model was hence not fed any human-defined knowledge about the number or nature of labelled sense categories. The two-dimensional plot below therefore visualizes the overall correspondence between clustered token embeddings, which can be conceptualized as &#8216;distributionally defined senses&#8217;, and the proposed &#8216;conceptually defined&#8217; sense labels.</p>
<fig id="F3">
<label>Figure 3</label>
<caption>
<p>t-SNE embeddings of <italic>over</italic>, perplexity = 20, KL divergence after 1,000 iterations: 0.457.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60710/"/>
</fig>
<p>From eyeballing <bold><italic><xref ref-type="fig" rid="F3">Figure 3</xref></italic></bold>, it looks like the distributional model proposes a fair number of distinct token-clusters that correspond relatively well with the suggested sense categories: in the majority of cases, the local clusters (or cluster areas) consist of tokens that were assigned the same label. At the bottom of <bold><italic><xref ref-type="fig" rid="F3">Figure 3</xref></italic></bold>, for instance, a local cluster area was marked, which comprises entirely of examples of Sense 2A (more specifically, those cases where the TR has mentally overcome an obstacle or past relationship). Yet, at the same time, the correspondence between the two models seems to become weaker at the global level. In some cases, such as Sense 2C (&#8216;completion&#8217;), 2E (&#8216;time span&#8217;), 4 (&#8216;examining&#8217;), 7 (&#8216;means of communication&#8217;) and 8 (&#8216;hangover&#8217;), it appears that all tokens with the same label are assigned to the same local cluster area, but in others, such as 2D (&#8216;transfer&#8217;), tokens are grouped in separate, relatively distant areas. Thus, the question arises how we can assess the degree to which there is correspondence between the distributionally defined and conceptually defined senses.</p>
<p>To address this question in a way that goes beyond eyeballing a visualization, this study uses a series of classification tasks, which help quantify the extent to which there is correspondence between the sense categories emerging from the two models (VAM). <bold><italic><xref ref-type="fig" rid="F4">Figure 4</xref></italic></bold> presents the results of the VAM when applied to all categories (over 100 iterations). The x-axis represents the abstraction continuum, which starts at no abstraction (the exemplar level, where classification of unseen tokens (20% of the data) is attempted by means of the nearest neighbour embeddings of concrete tokens) and reaches up to the target level (the highest level of abstraction, where all items within the 16 labelled categories are clustered into averaged &#8216;sense embeddings&#8217;). The y-axis represents the classification accuracy of the distributional model (F<sub>1</sub>-score, i.e. the harmonic mean of precision and recall when applied to the test set).</p>
<fig id="F4">
<label>Figure 4</label>
<caption>
<p>VAM output over all data.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60711/"/>
</fig>
<p>At the lowest level of abstraction, the model&#8217;s classification accuracy is quite high at 0.95, remaining relatively stable at this level until approximately 200 tokens have been clustered. Subsequently, its accuracy gradually drops to 0.8 (at approx. 300 tokens), after which it drops below 0.7 at the highest levels of abstraction. These findings imply that the BERT embeddings do encode the similarities between members of the categories proposed in the principled polysemy model, but only up to a certain point. Beyond that point, the proposed abstractions no longer optimally fit the output of the distributional model.</p>
<p>When we assess the models classification accuracy per sense category (<bold><italic><xref ref-type="fig" rid="F5">Figure 5</xref></italic></bold>), we also find that the model is more successful in &#8216;recognizing&#8217; some sense abstractions than others. Overall, the embeddings of <italic>over</italic> clearly encode the similarities between concrete tokens, and the performance of the model remains high at lower-intermediate levels of abstraction where only some of the concrete token embeddings are merged into slightly more schematic representations. Yet, whether higher levels of abstraction are still meaningfully encoded in the BERT embeddings of <italic>over</italic> seems to depend on the sense category under scrutiny. Unsurprisingly, perhaps, it is precisely the sense categories that occur in fixed syntactic configurations and have clear collocational preferences (e.g. Sense 2C &#8216;completion&#8217;, in which <italic>over</italic> consistently functions as an adprep in combination with a form of <italic>BE</italic>), are easier to group than more &#8216;schematic&#8217; senses.</p>
<fig id="F5">
<label>Figure 5</label>
<caption>
<p>VAM output per sense category.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60712/"/>
</fig>
<p>There are, however, a number of notable cases where the model&#8217;s classification accuracy drops sharply when more more items of the same category are merged. This pertains to Sense 3, where the model does not recognize similarities between, for instance, <italic>Spread a tablecloth over the table</italic> and <italic>He received votes from all over the floor</italic>. Reassuringly, this ties in with other analyses that have argued that &#8216;covering&#8217; does not adequately capture the interpretation of <italic>all over</italic> (<xref ref-type="bibr" rid="B59">Queller 2001</xref>; <xref ref-type="bibr" rid="B69">Taylor 2006</xref>; <xref ref-type="bibr" rid="B56">Pawelec 2010: 98&#8211;101</xref>). Furthermore, Sense 4A (&#8216;focus of attention&#8217;) and 5B (&#8216;control&#8217;) suffer from the high number of false positives of the other category the model wishes to assign to them (see Section 3.3.4&#8211;3.3.5 on the difficulty of distinguishing 4A and 5B). Finally, drops in performance can also be witnessed for category 2A and 2D, as the model seems to have difficulties in relating examples describing physical and non-physical scenes.</p>
</sec>
<sec>
<title>4.2 Relations between senses</title>
<p>Having established that there is some correspondence between clustered BERT embeddings and the proposed sense categories (up to a certain point), we can now turn to the question whether the geometrical distances between the various senses of <italic>over</italic> (as emergent from the distributional semantic model) correspond with the semantic relationships proposed in the prepositional polysemy network of <italic>over</italic> proposed by Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001: 746</xref>), reproduced in <bold><italic><xref ref-type="fig" rid="F6">Figure 6</xref></italic></bold>.</p>
<fig id="F6">
<label>Figure 6</label>
<caption>
<p>Polysemy Network of <italic>over</italic> as presented in Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001: 746</xref>).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60713/"/>
</fig>
<p>The dark, full nodes in the suggested network representation in <bold><italic><xref ref-type="fig" rid="F6">Figure 6</xref></italic></bold> constitute what are considered to be separate senses, whereas the empty nodes are included as abstractions over a proposed cluster of related, derived senses. At the centre, we find the protoscene (Sense 1). Because the representation in Tyler &amp; Evans has been constructed based on theoretical principles, it makes little sense to assess the representation based on the absolute geometrical distances between embeddings derived from the distributional model. It does make sense, however, to operationalize the directness of node linkage in the proposed network to relative distances between embeddings. More specifically, we could hypothesize that, if Sense 4A is not directly derived from the protoscene (Sense 1), but emerged as a further extension derived from Sense 4, we would expect that the relative distance between Sense 4A and Sense 1 is bigger than the relative distance between Sense 4A and Sense 4. Similarly, if Sense 2A, 2B, 2C and 2D form a cluster of related senses, we would expect the relative distance between, for instance, Sense 2A and 2B or 2A and 2D, to be shorter than the relative distance between Sense 2A and Sense 3.</p>
<p>To compare the principled polysemy network to the output of the distributional model by means of cosine distances, two approaches can be taken. First, we may approach the comparison by taking the senses proposed by Tyler &amp; Evans (<xref ref-type="bibr" rid="B71">2001</xref>; <xref ref-type="bibr" rid="B72">2003</xref>) as given, and rely on manually assigned labels to create an averaged sense embedding &#8211; that is, a summary embedding similar to the clusters created at the highest level of abstraction in the VAM. Subsequently, we can calculate cosine similarities between these sense embeddings. If the result of this assessment turns out to be that the relative distance between the sense embeddings maps onto the suggested relative distances in the network in <bold><italic><xref ref-type="fig" rid="F6">Figure 6</xref></italic></bold>, we could conclude that both models arrive at the same network representation in a relatively straightforward manner.</p>
<p>However, as explained in Section 4.1, the proposed sense categories do not always correspond with the way in which the tokens of a manually labelled category cluster, with categories such as 2A falling apart into multiple rather distinct groupings. A second approach, then, would be to adhere less strictly to the sense categories proposed by Tyler &amp; Evans, and determine the geometrical distance between token clusters proposed by the distributional model. In what follows, I restrict myself to this second approach.</p>
<p><bold><italic><xref ref-type="fig" rid="F7">Figure 7</xref></italic></bold> presents a hierarchical cluster tree of the annotated tokens. Of the 808 annotated examples, 26 were examples were excluded because they were considered unclear or ambiguous between multiple readings (cf. example (42)). The clustering presented in <bold><italic><xref ref-type="fig" rid="F7">Figure 7</xref></italic></bold> is based on the cosine distance between the embeddings of the remaining 782 examples. The coloured areas represent clusters of neighbouring tokens that were assigned to the same category. In order for a group of tokens to be considered a cluster, it was decided that there should be at least 5 neighbouring tokens of the same sense category. As such, the smallest clusters represented in <bold><italic><xref ref-type="fig" rid="F7">Figure 7</xref></italic></bold> are based on at least 5 examples, and the largest (i.e. cluster 2E, &#8216;time span&#8217;) contains 51 examples. In total, 47 examples did not have at least 4 neighbours of the same type.</p>
<fig id="F7">
<label>Figure 7</label>
<caption>
<p>Cluster tree (distance = cosine) with representative examples.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/5470/file/60714/"/>
</fig>
<p>As was briefly pointed out in Section 4.1, seven of the conceptual categories are &#8216;consistent&#8217;, with all tokens clustering together. This includes 2E (&#8216;time span&#8217;), 2C (&#8216;completion&#8217;), 4 (&#8216;examining&#8217;), 5B (&#8216;control&#8217;), 6 (&#8216;reflexive&#8217;), 7 (&#8216;communication channel&#8217;) and 8 (&#8216;hangover&#8217;). For the remaining nine categories, the model suggests that there are at least two different clusters. In some cases, the separate clusters are still part of the same higher-order branch (e.g. 1 &#8216;above&#8217;, 4A &#8216;focus-of-attention&#8217;, 6A &#8216;repetition&#8217;), whereas for others, the embeddings are less closely related (e.g. 2A &#8216;other-side&#8217;, 2D &#8216;transfer&#8217;, 5A &#8216;excess&#8217;). All in all, the suggested distances between the sense groupings differ substantially from the proposal put forward by Tyler &amp; Evans: not only does the distributional distributional model suggest fairly large distances between token groupings that Tyler &amp; Evans would have assigned to the same category (based on their shared underlying image schema), the proposed relative distances in the sense network are also not reflected in the geometrical distances between the (groupings) of embeddings.</p>
<p>Given that BERT does not group tokens of <italic>over</italic> according to abstract similarities in spatial configurations between them, the question that remains is what kind of groupings the model does suggest, and whether any (other) meaningful abstractions can be made. To address this question, we could examine the token groupings illustrated in <bold><italic><xref ref-type="fig" rid="F7">Figure 7</xref></italic></bold>.</p>
<p><bold>GROUP 1 &#8211; SPACE</bold> The first group that appears to be recognized appears to be a collection of spatial uses of <italic>over</italic>. Tokens grouped in Group 1 include all tokens assigned to Sense 1 (&#8216;above&#8217;), and some tokens of the spatial &#8216;excess&#8217; sense, 2B, in which a spatial border or threshold is exceeded. Note, however, that this cluster does not correspond with Kreitzer (<xref ref-type="bibr" rid="B44">1997</xref>)&#8217;s static <italic>over</italic><sub>1</sub>: Group 1 also includes a subgroup of tokens of Sense 3 (&#8216;covering&#8217;, <italic>over</italic><sub>3</sub>), (stative) uses of Sense 2A (&#8216;other-side-of&#8217;, <italic>over</italic><sub>2</sub>), and examples involving dynamic verbs (e.g. <italic>the cat jumped over the fence, over</italic><sub>2</sub>).</p>
<p><bold>GROUP 2 &#8211; TIME SPAN and GROUP 3 &#8211; COMPLETION</bold> The two temporal uses of &#8216;over&#8217;, 2C (&#8216;completion&#8217;) and 2E (&#8216;time span&#8217;), constitute fairly separate categories. The coherent cluster of 2E tokens will henceforth be referred to as Group 2. The coherent cluster of 2C will henceforth be referred to as Group 3. In all examples in Group 2, <italic>over</italic> functions as a preposition, whereas in all examples in Group 3 function as an adprep.</p>
<p><bold>GROUP 4 &#8211; MIND AND PERCEPTION</bold> As the closest neighbour of Group 2, we can discern a very large grouping of tokens where <italic>over</italic> consistently marks a non-spatial relation (Group 4). In Group 4, we find all examples categorized as 4A (&#8216;focus of attention&#8217;). One cluster of examples of Sense 4A involves verbs such as <italic>preside</italic> and <italic>watch</italic>. These are closely associated with examples of 2D, in which there is a transfer of power or authority (e.g. <italic>He took over the business</italic>). These are, in turn, closely related to 5B (&#8216;control&#8217;). The other token cluster of 4A (&#8216;focus-of-attention&#8217;), which includes all examples where <italic>over</italic> can be paraphrased with &#8216;about&#8217; or &#8216;because of&#8217; (e.g. <italic>He agonized over it</italic>), is most closely linked to the subgroup of tokens in category 2A, where the TR has mentally overcome or lost interest in the LM (e.g. <italic>I am over the drama</italic>). We furthermore find all tokens of 5C (&#8216;preference&#8217;) in Group 4, closely positioned next to a subgroup of 2B, where the LM is omitted or skipped in a selection procedure (possibly implying absence of preference, e.g. <italic>She was passed over for the job</italic>). Group 4 also contains all tokens of 4 (&#8216;examining&#8217;). Given that the latter group of examples frequently (but not exclusively) involves the phrasal verb combination <italic>look over</italic>, it is not surprising that these tokens are positioned relatively close to (but, notably, are not confused with) cases where a glance is cast (classified as 2A). Finally, we also find all tokens of Sense 7 (&#8216;means of communication&#8217;) in Group 4. While perhaps more loosely related to the &#8216;mind&#8217; and &#8216;perception&#8217; relations, Sense 7 also involves animate TRs and a non-spatial interpretation (i.e. in a sentence like <italic>they spoke over the phone</italic>, the preposition does not capture a physical, spatial positioning of the TR and LM).</p>
<p><bold>GROUP 5 &#8211; PATH</bold> Related to Group 3, we find a cluster of tokens where <italic>over</italic> again has a spatial interpretation. Yet, unlike the tokens in Group 1, <italic>over</italic> functions as an adprep, and involves movement along a path (and hence partially overlaps with Kreitzer (<xref ref-type="bibr" rid="B44">1997</xref>)&#8217;s dynamic <italic>over</italic><sub>2</sub>). These include examples of 2D (&#8216;transfer&#8217;) as well as examples of 2A where the TR moves to a different location (e.g. <italic>I made my way over (to the computer)</italic>).</p>
<p><bold>GROUP 6 &#8211; EXCESS (IN NUMBERS)</bold> In the remaining set of tokens, then, a clustering of 5A tokens can be distinguished, where the LM is an amount or quantitative threshold (e.g. She mentioned that I donated <italic>over</italic> $100,000 to Katrina victims (2006, COHA)) or an amount (this cluster is also identified by <xref ref-type="bibr" rid="B53">Newman 2011</xref>).</p>
<p><bold>GROUP 7 &#8211; NON-SPATIAL ADPREPS AND FIXED PHRASES</bold> While it is not easy to make sense of Group 7, it is interesting to note that all tokens classified as Sense 6 (&#8216;reflexive&#8217;) and 6A (&#8216;repetition&#8217;) are part of this group. Yet, rather than clustering together, they are positioned closely to token groupings classified as Sense 3 (in the fixed combination <italic>all over</italic>, e.g. <italic>I feel pain all over</italic>), Sense 5A (non-literal adprep uses, as in <italic>their personal involvement will spill over into their workplace interaction</italic>), and Sense 8 (&#8216;hangover&#8217;). The relation between these groupings seems formal rather than semantic, as the majority of groupings present cases where <italic>over</italic> functions as an adprep, and occurs in relatively fixed phrases.</p>
<p>If we approach the complex internal semantic structure of <italic>over</italic> by means of BERT embeddings, then, it appears that global clusters are formed based on in intersection of similarities in, on the one hand, conceptual domain (i.e., spatial, temporal, mental, etc.), and syntactic resemblance on the other. Overall, the lack of correspondence between the suggested network configuration in <bold><italic><xref ref-type="fig" rid="F6">Figure 6</xref></italic></bold> and the global distances between the grouped embeddings does not necessarily indicate that the global groupings are not interpretable, or that no abstractions can be made &#8211; rather, it suggests that the use of an embedding-based approach leads to different abstractions.</p>
</sec>
</sec>
<sec>
<title>5 Discussion and Conclusion</title>
<p>Within Cognitive Linguistics, there has been no shortage of proposals for modelling polysemy networks, in which the syntactic configurations, collocations, and the notion of underlying image schemas are of central concern. To minimize the arguably subjective nature of further proposals, linguists are increasingly turning to the use of distributional, statistical methods, and, most recently, to deep contextualized neural language models. In the present study, I investigated to what extent the output of a fully unsupervised application of BERT (a &#8216;meaning as context&#8217; model) corresponds with the sense network of <italic>over</italic> proposed by Tyler &amp; Evans (<xref ref-type="bibr" rid="B72">2003</xref>; <xref ref-type="bibr" rid="B71">2001</xref>) (a &#8216;meaning as concept&#8217; model). The analyses reveal that, while there are interesting correspondences the two approaches, they ultimately lead to different abstractions. Which of these abstractions most closely approximate the abstractions that emerge from ellicited, experimental data (or the extent to which the &#8216;context&#8217; and &#8216;concept&#8217; models are complementary) remains an open question that needs to be addressed by means of behavioural studies. However, because Tyler &amp; Evans&#8217; proposal foregrounds the importance of sense connections via image schemas, the analysis presented in this study does provide some insight into the extent to which such imagistic information may be encoded in BERT embeddings.</p>
<p>What emerges from the preceding analyses is that the extent to which the models appear to converge varies depending on whether we rely on BERT embeddings to detect local (concrete, token-based) or global (schematic) similarities between examples. At the local level, BERT&#8217;s focus on collocational and syntactic patterns helps it in &#8216;recognizing&#8217; similarities between tokens of the same category, resulting in relatively coherent local clusters. When the consistency of these local clusters does deviate from what is proposed by the principled polysemy model, it is furthermore reassuring that we can come up with a reasonable explanation for the divergence. For instance, when the embeddings of the tokens of the same principled polysemy category are split in separate clusters, the split can be motivated semantically: examples of Sense 2A (&#8216;other-side&#8217;) that are used in a literal, physical sense (e.g. <italic>I got over the hill</italic>) are distinguished from more metaphorical uses (e.g. <italic>I got over my puberty weirdness</italic>). Notably, BERT&#8217;s &#8216;recognition&#8217; of such metaphorical uses extends into cases where surrounding context words are themselves used metaphorically (e.g. <italic>We will</italic> move on <italic>and get over these</italic> rough patches), indicating an unexpected aptitude for coping with elaborated metaphors (cf. <xref ref-type="bibr" rid="B16">De Pascale 2019: 157</xref>).</p>
<p>Yet, the fact that BERT is apt at recognizing metaphorical uses of <italic>over</italic> does not necessarily imply that it also recognizes that they are, in fact, metaphorical extensions of a literal, spatial source. This becomes evident when we consider the model&#8217;s output at the global level. When we examine the relations between the token groupings emergent from the cluster analysis presented in Section 4.2, we find that the geometrical distance between the embedding of a particular spatial use of <italic>over</italic>, which may have given rise to a particular non-spatial use via metaphorical extension, is not shorter than the geometrical distance between the embeddings of two literal spatial uses (e.g. Group 1) or two metaphorical uses (e.g. Group 4) with different underlying image schemas. As such, if there is indeed a close connection between a spatial configuration and a non-physical scene it embodies, there may be a discrepancy between the geometric distances between the token embeddings of <italic>over</italic> and the actual conceptual similarity between those tokens.</p>
<p>The observation BERT does not immediately capture similarities in terms of image-schema resemblances can be understood in light of the fact that the model has been trained on linguistic data alone, and has no experience with (physical) non-linguistic, perceptual information (such as spatial configurations, but also, for example, visual properties such as colours: <xref ref-type="bibr" rid="B67">Sommerauer &amp; Fokkens 2018</xref>). Hence, BERT embeddings pick up fine-grained semantic distinctions based on collocational and morpho-syntactic cues, and can be employed to successfully group senses in distinct domains. However, BERT embeddings seem less equipped to flag abstract configurational resemblances in image schemas across those domains, which helps highlight what sort of semantic information is (and is not) encoded in contextualized embeddings.</p>
<p>Note that, if the perceptual, imagistic information that motivates the abstractions and sense connections made in the cognitive-conceptual model is still somehow encoded in contextual information (as appears to be suggested by <xref ref-type="bibr" rid="B35">Gromann &amp; Hedblom (2017)</xref>), such information could be brought to the fore by further experimentation with the model&#8217;s hyperparameters (e.g. different context window sizes, different (combinations of) layers), or fine-tuning the model to a sense classification task by exposing it to manually labelled examples in training. Yet, if such perceptual information is not represented in context embeddings and requires extra-linguistic knowledge, an interesting avenue to pursue is, for instance, to train language models based on coupled textual and visual input (<xref ref-type="bibr" rid="B13">Chrupa&#322;a et al. 2015</xref>). Of course, whether such additional training and supervision is desirable depends entirely on the question the researcher wishes to address, and which facets of meaning they deem relevant within their study or theoretical framework. In Cognitive Linguistics, researchers may be inclined to say that a model of meaning representation should capture the global resemblance between the underlying image schemas of prepositions, as image schemas (and embodiment) are part of the core tenets of the framework (e.g. <xref ref-type="bibr" rid="B54">Oakley 2010</xref>; <xref ref-type="bibr" rid="B26">Gibbs &amp; Matlock 2001: 233</xref>), and play an important role in, for instance, studies of semantic change and grammaticalization (e.g. <xref ref-type="bibr" rid="B61">Rhee 2002</xref>).</p>
<p>As a final concluding remark, I wish to add that the findings presented in this study have important implications for the integration of neural language models &#8211; and perhaps, more generally, the application of Semantic Vector Space Models &#8211; in theoretical linguistic research, and in particular, to research on semantic change. In a recent publication, Boleda (<xref ref-type="bibr" rid="B6">2020</xref>) surveys a number of studies that have applied either count or predictive models to historical and diachronic corpus data. Such studies, which involve examination of nearest neighbours and cosine similarities between type- and/or token-vectors, have provided the key to detecting, as well as describing the diachronic trajectory of lexical and, albeit less commonly, grammatical semantic changes (e.g. <xref ref-type="bibr" rid="B40">Hilpert &amp; Perek 2015</xref>; <xref ref-type="bibr" rid="B36">Hamilton et al. 2016</xref>; <xref ref-type="bibr" rid="B19">Dubossarsky 2018</xref>; <xref ref-type="bibr" rid="B64">Sagi et al. 2011</xref>; <xref ref-type="bibr" rid="B39">Hilpert &amp; Correia Saavedra 2017</xref>; <xref ref-type="bibr" rid="B10">Budts &amp; Petr&#233; 2020</xref>; <xref ref-type="bibr" rid="B28">Giulianelli et al. 2020</xref>). In some of these studies, it is argued that distributional semantic models could also be employed to detect different types of semantic change (and, by extension, I could add that they may also help assess competing hypotheses regarding the mechanisms of change at play in a particular diachronic development). Some steps have already been taken in this direction (e.g. the automated detection of semantic broadening and narrowing in Sagi et al. (<xref ref-type="bibr" rid="B64">2011</xref>); <xref ref-type="bibr" rid="B28">Giulianelli et al. (2020)</xref>), and indeed, BERT could be an excellent tool for detecting metaphorical extensions of linguistic items in diachronic corpora (<xref ref-type="bibr" rid="B28">Giulianelli et al. 2020</xref>). However, it should be clear that, at least when left entirely unsupervised, BERT does not seem to pick up that there may be abstract, imagistic similarities between domains. As such, researchers interested in studying metaphorical extensions (of prepositions or otherwise) should take into consideration that unsupervised BERT will be great at indicating <italic>that</italic> a metaphorical extension has occurred from one domain to another, but they do not reveal which perceptual similarity pattern is the most likely source of the extension. It could be possible, however, to tackle these issues by experimenting with additional supervision and different model architectures, and, crucially, by accelerating the dialogue on how to integrate these models in theoretical linguistic research, and vice versa.</p>
</sec>
<sec>
<title>Data accessibility statement</title>
<p>All data and code can be found at <italic><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/LFonteyn/Glossa_over">https://github.com/LFonteyn/Glossa_over</ext-link></italic>.</p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>I wish to thank Charlotte Maekelberghe for acting as the second annotator. I furthermore thank all anonymous reviewers as well as Folgert Karsdorp and Stefano De Pascale for their insightful feedback on earlier versions of this manuscript.</p>
</ack>
<sec>
<title>Competing interests</title>
<p>The author has no competing interests to declare.</p>
</sec>
<ref-list>
<ref id="B1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Alishahi</surname>, <given-names>Afra</given-names></string-name>, <string-name><given-names>Grzegorz</given-names> <surname>Chrupa&#322;a</surname></string-name> &amp; <string-name><given-names>Tal</given-names> <surname>Linzen</surname></string-name>. <year>2019</year>. <article-title>Analyzing and interpreting neural networks for nlp: A report on the first blackboxnlp workshop</article-title>. <source>Natural Language Engineering</source> <volume>25</volume>(<issue>4</issue>). <fpage>543</fpage>&#8211;<lpage>557</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S135132491900024X</pub-id></mixed-citation></ref>
<ref id="B2"><label>2</label><mixed-citation publication-type="confproc"><string-name><surname>Baroni</surname>, <given-names>Marco</given-names></string-name> &amp; <string-name><given-names>Alessandro</given-names> <surname>Lenci</surname></string-name>. <year>2011</year>. <article-title>How we BLESSed distributional semantic evaluation</article-title>. In <conf-name>Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Language Semantics</conf-name>, <fpage>1</fpage>&#8211;<lpage>10</lpage>. <conf-loc>Edinburgh, UK</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>.</mixed-citation></ref>
<ref id="B3"><label>3</label><mixed-citation publication-type="confproc"><string-name><surname>Baroni</surname>, <given-names>Marco</given-names></string-name>, <string-name><given-names>Georgiana</given-names> <surname>Dinu</surname></string-name> &amp; <string-name><given-names>Germ&#225;n</given-names> <surname>Kruszewski</surname></string-name>. <year>2014</year>. <article-title>Don&#8217;t count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectors</article-title>. In <conf-name>Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</conf-name>, <fpage>238</fpage>&#8211;<lpage>247</lpage>. <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.3115/v1/P14-1023</pub-id></mixed-citation></ref>
<ref id="B4"><label>4</label><mixed-citation publication-type="confproc"><string-name><surname>Berez</surname>, <given-names>Andrea L.</given-names></string-name> &amp; <string-name><given-names>Stefan Th.</given-names> <surname>Gries</surname></string-name>. <year>2008</year>. <article-title>In defense of corpus-based methods: A behavioral profile analysis of polysemous get in English</article-title>. In <conf-name>Proceedings of the 24th NWLC</conf-name>, <fpage>157</fpage>&#8211;<lpage>166</lpage>. <conf-loc>Seattle, WA</conf-loc>.</mixed-citation></ref>
<ref id="B5"><label>5</label><mixed-citation publication-type="confproc"><string-name><surname>Blevins</surname>, <given-names>Terra</given-names></string-name> &amp; <string-name><given-names>Luke</given-names> <surname>Zettlemoyer</surname></string-name>. <year>2020</year>. <article-title>Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encoders</article-title>. In <conf-name>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</conf-name>, <fpage>1006</fpage>&#8211;<lpage>1017</lpage>. Online: <conf-sponsor>>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.95</pub-id></mixed-citation></ref>
<ref id="B6"><label>6</label><mixed-citation publication-type="journal"><string-name><surname>Boleda</surname>, <given-names>Gemma</given-names></string-name>. <year>2020</year>. <article-title>Distributional Semantics and Linguistic Theory</article-title>. <source>Annual Review of Linguistics</source> <volume>6</volume>(<issue>1</issue>). <fpage>213</fpage>&#8211;<lpage>234</lpage>. DOI: <pub-id pub-id-type="doi">10.1146/annurev-linguistics-011619-030303</pub-id></mixed-citation></ref>
<ref id="B7"><label>7</label><mixed-citation publication-type="confproc"><string-name><surname>Boleda</surname>, <given-names>Gemma</given-names></string-name> &amp; <string-name><given-names>Katrin</given-names> <surname>Erk</surname></string-name>. <year>2015</year>. <article-title>Distributional Semantic Features as Semantic Primitives &#8211; Or Not</article-title>. In <conf-name>Aaai spring symposium on knowledge representation and reasoning</conf-name>, <fpage>2</fpage>&#8211;<lpage>5</lpage>. <conf-loc>USA</conf-loc>: <conf-sponsor>>Stanford University</conf-sponsor>.</mixed-citation></ref>
<ref id="B8"><label>8</label><mixed-citation publication-type="book"><string-name><surname>Brugman</surname>, <given-names>Claudia M.</given-names></string-name> <year>1988</year>. <source>The Story of Over: Polysemy, Semantics, and the Structure of the Lexicon</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Garland</publisher-name>.</mixed-citation></ref>
<ref id="B9"><label>9</label><mixed-citation publication-type="thesis"><string-name><surname>Budts</surname>, <given-names>Sara</given-names></string-name>. <year>2020</year>. <source>On periphrastic do and the modal auxiliaries: a connectionist approach to language change</source>. <publisher-loc>Antwerp</publisher-loc>: <publisher-name>Universiteit Antwerpen</publisher-name> PhD dissertation.</mixed-citation></ref>
<ref id="B10"><label>10</label><mixed-citation publication-type="book"><string-name><surname>Budts</surname>, <given-names>Sara</given-names></string-name> &amp; <string-name><given-names>Peter</given-names> <surname>Petr&#233;</surname></string-name>. <year>2020</year>. <chapter-title>Putting connections centre stage in diachronic construction grammar</chapter-title>. In <string-name><given-names>Lotte</given-names> <surname>Sommerer</surname></string-name> &amp; <string-name><given-names>Elena</given-names> <surname>Smirnova</surname></string-name> (eds.), <source>Nodes and Networks in Diachronic Construction Grammar</source>, <fpage>317</fpage>&#8211;<lpage>352</lpage>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>John Benjamins</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1075/cal.27.09bud</pub-id></mixed-citation></ref>
<ref id="B11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Bullinaria</surname>, <given-names>John A.</given-names></string-name> &amp; <string-name><given-names>Joseph P.</given-names> <surname>Levy</surname></string-name>. <year>2007</year>. <article-title>Extracting semantic representations from word co-occurrence statistics: A computational study</article-title>. <source>Behavior Research Methods</source> <volume>39</volume>(<issue>3</issue>). <fpage>510</fpage>&#8211;<lpage>526</lpage>. DOI: <pub-id pub-id-type="doi">10.3758/BF03193020</pub-id></mixed-citation></ref>
<ref id="B12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Bullinaria</surname>, <given-names>John A.</given-names></string-name> &amp; <string-name><given-names>Joseph P.</given-names> <surname>Levy</surname></string-name>. <year>2012</year>. <article-title>Extracting semantic representations from word co-occurrence statistics: stop-lists, stemming, and SVD</article-title>. <source>Behavior Research Methods</source> <volume>44</volume>(<issue>3</issue>). <fpage>890</fpage>&#8211;<lpage>907</lpage>. DOI: <pub-id pub-id-type="doi">10.3758/s13428-011-0183-8</pub-id></mixed-citation></ref>
<ref id="B13"><label>13</label><mixed-citation publication-type="confproc"><string-name><surname>Chrupa&#322;a</surname>, <given-names>Grzegorz</given-names></string-name>, <string-name><given-names>&#193;kos</given-names> <surname>K&#225;d&#225;r</surname></string-name> &amp; <string-name><given-names>Afra</given-names> <surname>Alishahi</surname></string-name>. <year>2015</year>. <article-title>Learning language through pictures</article-title>. In <conf-name>Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)</conf-name>, <fpage>112</fpage>&#8211;<lpage>118</lpage>. <conf-loc>Beijing, China</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.3115/v1/P15-2019</pub-id></mixed-citation></ref>
<ref id="B14"><label>14</label><mixed-citation publication-type="journal"><string-name><surname>Clark</surname>, <given-names>Kevin</given-names></string-name>, <string-name><given-names>Urvashi</given-names> <surname>Khandelwal</surname></string-name>, <string-name><given-names>Omer</given-names> <surname>Levy</surname></string-name> &amp; <string-name><given-names>Christopher D.</given-names> <surname>Manning</surname></string-name>. <year>2019</year>. <article-title>What Does BERT Look At? An Analysis of BERT&#8217;s Attention</article-title>. <source>arXiv:1906.04341</source>. DOI: <pub-id pub-id-type="doi">10.18653/v1/W19-4828</pub-id></mixed-citation></ref>
<ref id="B15"><label>15</label><mixed-citation publication-type="journal"><string-name><surname>Clausner</surname>, <given-names>Timothy C.</given-names></string-name> &amp; <string-name><given-names>William</given-names> <surname>Croft</surname></string-name>. <year>1999</year>. <article-title>Domains and image schemas</article-title>. <source>Cognitive Linguistics</source> <volume>10</volume>(<issue>1</issue>). <fpage>1</fpage>&#8211;<lpage>31</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/cogl.1999.001</pub-id></mixed-citation></ref>
<ref id="B16"><label>16</label><mixed-citation publication-type="thesis"><string-name><surname>De Pascale</surname>, <given-names>Stefano</given-names></string-name>. <year>2019</year>. <source>Token-based vector space models as semantic control in lexical sociolectometry</source>. <publisher-loc>Leuven</publisher-loc>: <publisher-name>KU Leuven</publisher-name> PhD dissertation.</mixed-citation></ref>
<ref id="B17"><label>17</label><mixed-citation publication-type="journal"><string-name><surname>Desagulier</surname>, <given-names>Guillaume</given-names></string-name>. <year>2019</year>. <article-title>Can word vectors help corpus linguists?</article-title> <source>Studia Neophilologica</source> <volume>91</volume>(<issue>2</issue>). <fpage>219</fpage>&#8211;<lpage>240</lpage>. DOI: <pub-id pub-id-type="doi">10.1080/00393274.2019.1616220</pub-id></mixed-citation></ref>
<ref id="B18"><label>18</label><mixed-citation publication-type="confproc"><string-name><surname>Devlin</surname>, <given-names>Jacob</given-names></string-name>, <string-name><given-names>Ming-Wei</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>Kenton</given-names> <surname>Lee</surname></string-name> &amp; <string-name><given-names>Kristina</given-names> <surname>Toutanova</surname></string-name>. <year>2019</year>. <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>. In <conf-name>Proceedings of NAACL-HLT 2019</conf-name>, <fpage>4171</fpage>&#8211;<lpage>4186</lpage>. <conf-loc>Minneapolis, Minnesota</conf-loc>.</mixed-citation></ref>
<ref id="B19"><label>19</label><mixed-citation publication-type="thesis"><string-name><surname>Dubossarsky</surname>, <given-names>Haim</given-names></string-name>. <year>2018</year>. <source>Semantic change at large: A computational approach for semantic change</source>. <publisher-loc>Jerusalem</publisher-loc>: <publisher-name>the Senate of the Hebrew University of Jerusalem</publisher-name> PhD dissertation.</mixed-citation></ref>
<ref id="B20"><label>20</label><mixed-citation publication-type="confproc"><string-name><surname>Dubossarsky</surname>, <given-names>Haim</given-names></string-name>, <string-name><given-names>Daphna</given-names> <surname>Weinshall</surname></string-name> &amp; <string-name><given-names>Eitan</given-names> <surname>Grossman</surname></string-name>. <year>2017</year>. <article-title>Outta Control: Laws of Semantic Change and Inherent Biases in Word Representation Models</article-title>. In <conf-name>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</conf-name>, <fpage>1136</fpage>&#8211;<lpage>1145</lpage>. <conf-loc>Copenhagen, Denmark</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/D17-1118</pub-id></mixed-citation></ref>
<ref id="B21"><label>21</label><mixed-citation publication-type="confproc"><string-name><surname>Erk</surname>, <given-names>Katrin</given-names></string-name> &amp; <string-name><given-names>Sebastian</given-names> <surname>Pad&#243;</surname></string-name>. <year>2010</year>. <article-title>Exemplar-Based Models for Word Meaning in Context</article-title>. In <conf-name>Proceedings of the ACL 2010 Conference Short Papers</conf-name>, <fpage>92</fpage>&#8211;<lpage>97</lpage>. <conf-loc>Uppsala, Sweden</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.1017/S0022226704003056</pub-id></mixed-citation></ref>
<ref id="B22"><label>22</label><mixed-citation publication-type="journal"><string-name><surname>Evans</surname>, <given-names>Vyvyan</given-names></string-name>. <year>2005</year>. <article-title>The meaning of time: polysemy, the lexicon and conceptual structure</article-title>. <source>Journal of Linguistics</source> <volume>41</volume>(<issue>1</issue>). <fpage>33</fpage>&#8211;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="B23"><label>23</label><mixed-citation publication-type="book"><string-name><surname>Firth</surname>, <given-names>John Rupert</given-names></string-name>. <year>1957</year>. <source>Papers in Linguistics 1934&#8211;1951</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</mixed-citation></ref>
<ref id="B24"><label>24</label><mixed-citation publication-type="book"><string-name><surname>Geeraerts</surname>, <given-names>Dirk</given-names></string-name>. <year>2016</year>. <chapter-title>Sense individuation</chapter-title>. In <string-name><given-names>Nick</given-names> <surname>Riemer</surname></string-name> (ed.), <source>The Routledge Handbook of Semantics</source>, <fpage>233</fpage>&#8211;<lpage>247</lpage>. <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>.</mixed-citation></ref>
<ref id="B25"><label>25</label><mixed-citation publication-type="journal"><string-name><surname>Gibbs</surname>, <given-names>Raymond W.</given-names></string-name>, <string-name><given-names>Dinara A.</given-names> <surname>Beitel</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Harrington</surname></string-name> &amp; <string-name><given-names>Paul E.</given-names> <surname>Sanders</surname></string-name>. <year>1994</year>. <article-title>Taking a Stand on the Meanings of <italic>Stand</italic>: Bodily Experience as Motivation for Polysemy</article-title>. <source>Journal of Semantics</source> <volume>11</volume>(<issue>4</issue>). <fpage>231</fpage>&#8211;<lpage>251</lpage>. DOI: <pub-id pub-id-type="doi">10.1093/jos/11.4.231</pub-id></mixed-citation></ref>
<ref id="B26"><label>26</label><mixed-citation publication-type="book"><string-name><surname>Gibbs</surname>, <given-names>Raymond W.</given-names></string-name> &amp; <string-name><given-names>Teenie</given-names> <surname>Matlock</surname></string-name>. <year>2001</year>. <chapter-title>Psycholinguistic perspectives on polysemy</chapter-title>. In <string-name><given-names>Hubert</given-names> <surname>Cuyckens</surname></string-name> &amp; <string-name><given-names>Britta</given-names> <surname>Zawada</surname></string-name> (eds.), <source>Polysemy in Cognitive Linguistics</source>, <fpage>213</fpage>&#8211;<lpage>239</lpage>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>John Benjamins</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1075/cilt.177.10gib</pub-id></mixed-citation></ref>
<ref id="B27"><label>27</label><mixed-citation publication-type="journal"><string-name><surname>Gilquin</surname>, <given-names>Ga&#235;tanelle</given-names></string-name> &amp; <string-name><given-names>Andrew</given-names> <surname>McMichael</surname></string-name>. <year>2018</year>. <article-title>Through the prototypes of through: A corpus-based cognitive analysis</article-title>. <source>Yearbook of the German Cognitive Linguistics Association</source> <volume>6</volume>(<issue>1</issue>). <fpage>43</fpage>&#8211;<lpage>70</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/gcla-2018-0003</pub-id></mixed-citation></ref>
<ref id="B28"><label>28</label><mixed-citation publication-type="confproc"><string-name><surname>Giulianelli</surname>, <given-names>Mario</given-names></string-name>, <string-name><given-names>Marco</given-names> <surname>Del Tredici</surname></string-name> &amp; <string-name><given-names>Raquel</given-names> <surname>Fern&#225;ndez</surname></string-name>. <year>2020</year>. <article-title>Analysing lexical semantic change with contextualised word representations</article-title>. In <conf-name>Proceedings of the 58th annual meeting of the association for computational linguistics</conf-name>, <fpage>3960</fpage>&#8211;<lpage>3973</lpage>. Online: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.365</pub-id></mixed-citation></ref>
<ref id="B29"><label>29</label><mixed-citation publication-type="book"><string-name><surname>Glynn</surname>, <given-names>Dylan</given-names></string-name>. <year>2010</year>. <chapter-title>Testing the hypothesis. Objectivity and verification in usage-based Cognitive Semantics</chapter-title>. In <string-name><given-names>Dylan</given-names> <surname>Glynn</surname></string-name> &amp; <string-name><given-names>Kerstin</given-names> <surname>Fischer</surname></string-name> (eds.), <source>Quantitative Methods in Cognitive Semantics: Corpus-Driven Approaches</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1515/9783110226423.239</pub-id></mixed-citation></ref>
<ref id="B30"><label>30</label><mixed-citation publication-type="book"><string-name><surname>Goldberg</surname>, <given-names>Adele E.</given-names></string-name> <year>1995</year>. <source>Constructions: A Construction Grammar Approach to Argument Structure</source>. <publisher-loc>Chicago</publisher-loc>: <publisher-name>University of Chicago Press</publisher-name>.</mixed-citation></ref>
<ref id="B31"><label>31</label><mixed-citation publication-type="book"><string-name><surname>Gries</surname>, <given-names>Stefan Th.</given-names></string-name> <year>2006</year>. <chapter-title>Corpus-based methods and cognitive semantics: The many senses of to run</chapter-title>. In <string-name><given-names>Stefan Th.</given-names> <surname>Gries</surname></string-name> &amp; <string-name><given-names>Anatol</given-names> <surname>Stefanowitsch</surname></string-name> (eds.), <source>Trends in Linguistics. Studies and Monographs [TiLSM]</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>Mouton de Gruyter</publisher-name>.</mixed-citation></ref>
<ref id="B32"><label>32</label><mixed-citation publication-type="book"><string-name><surname>Gries</surname>, <given-names>Stefan Th.</given-names></string-name> &amp; <string-name><given-names>Dagmar</given-names> <surname>Divjak</surname></string-name>. <year>2009</year>. <chapter-title>Behavioral profiles: A corpus-based approach to cognitive semantic analysis</chapter-title>. In <string-name><given-names>Vyvyan</given-names> <surname>Evans</surname></string-name> &amp; <string-name><given-names>St&#233;phanie</given-names> <surname>Pourcel</surname></string-name> (eds.), <source>Human Cognitive Processing</source> <volume>24</volume>, <fpage>57</fpage>&#8211;<lpage>75</lpage>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>John Benjamins Publishing Company</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1075/hcp.24.07gri</pub-id></mixed-citation></ref>
<ref id="B33"><label>33</label><mixed-citation publication-type="book"><string-name><surname>Gries</surname>, <given-names>Stefan Th.</given-names></string-name> &amp; <string-name><given-names>Dagmar</given-names> <surname>Divjak</surname></string-name>. <year>2010</year>. <chapter-title>Quantitative approaches in usagebased Cognitive Semantics: Myths, erroneous assumptions, and a proposal</chapter-title>. In <string-name><given-names>Dylan</given-names> <surname>Glynn</surname></string-name> &amp; <string-name><given-names>Kerstin</given-names> <surname>Fischer</surname></string-name> (eds.), <source>Quantitative Methods in Cognitive Semantics: Corpus-Driven Approaches</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1515/9783110226423.331</pub-id></mixed-citation></ref>
<ref id="B34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Gries</surname>, <given-names>Stefan Th</given-names></string-name> &amp; <string-name><given-names>Naoki</given-names> <surname>Otani</surname></string-name>. <year>2010</year>. <article-title>Behavioral profiles: A corpus-based perspective on synonymy and antonymy</article-title>. <source>ICAME journal</source> <volume>34</volume>. <fpage>30</fpage>.</mixed-citation></ref>
<ref id="B35"><label>35</label><mixed-citation publication-type="journal"><string-name><surname>Gromann</surname>, <given-names>Dagmar</given-names></string-name> &amp; <string-name><given-names>Maria M.</given-names> <surname>Hedblom</surname></string-name>. <year>2017</year>. <article-title>Kinesthetic Mind Reader: A Method to Identify Image Schemas in Natural Language</article-title>. <source>Advances in Cognitive Systems</source> <volume>5</volume>. <fpage>14</fpage>.</mixed-citation></ref>
<ref id="B36"><label>36</label><mixed-citation publication-type="confproc"><string-name><surname>Hamilton</surname>, <given-names>William L.</given-names></string-name>, <string-name><given-names>Jure</given-names> <surname>Leskovec</surname></string-name> &amp; <string-name><given-names>Dan</given-names> <surname>Jurafsky</surname></string-name>. <year>2016</year>. <article-title>Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change</article-title>. In <conf-name>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</conf-name>, <fpage>1489</fpage>&#8211;<lpage>1501</lpage>. <conf-loc>Berlin, Germany</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/P16-1141</pub-id></mixed-citation></ref>
<ref id="B37"><label>37</label><mixed-citation publication-type="journal"><string-name><surname>Harris</surname>, <given-names>Zellig S.</given-names></string-name> <year>1954</year>. <article-title>Distributional Structure</article-title>. <source>WORD</source> <volume>10</volume>(<issue>2&#8211;3</issue>). <fpage>146</fpage>&#8211;<lpage>162</lpage>. DOI: <pub-id pub-id-type="doi">10.1080/00437956.1954.11659520</pub-id></mixed-citation></ref>
<ref id="B38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Heylen</surname>, <given-names>K.</given-names></string-name>, <string-name><given-names>T.</given-names> <surname>Wielfaert</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Speelman</surname></string-name> &amp; <string-name><given-names>D.</given-names> <surname>Geeraerts</surname></string-name>. <year>2015</year>. <article-title>Monitoring polysemy: Word space models as a tool for large-scale lexical semantic analysis</article-title>. <source>Lingua</source> <volume>157</volume>. <fpage>153</fpage>&#8211;<lpage>172</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.lingua.2014.12.001</pub-id></mixed-citation></ref>
<ref id="B39"><label>39</label><mixed-citation publication-type="journal"><string-name><surname>Hilpert</surname>, <given-names>Martin</given-names></string-name> &amp; <string-name><given-names>David Correia</given-names> <surname>Saavedra</surname></string-name>. <year>2017</year>. <article-title>Using token-based semantic vector spaces for corpus-linguistic analyses: From practical applications to tests of theoretical claims</article-title>. <source>Corpus Linguistics and Linguistic Theory</source> <volume>16</volume>(<issue>2</issue>). <fpage>393</fpage>&#8211;<lpage>424</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/cllt-2017-0009</pub-id></mixed-citation></ref>
<ref id="B40"><label>40</label><mixed-citation publication-type="journal"><string-name><surname>Hilpert</surname>, <given-names>Martin</given-names></string-name> &amp; <string-name><given-names>Florent</given-names> <surname>Perek</surname></string-name>. <year>2015</year>. <article-title>Meaning change in a petri dish: constructions, semantic vector spaces, and motion charts</article-title>. <source>Linguistics Vanguard</source> <volume>1</volume>(<issue>1</issue>). <fpage>339</fpage>&#8211;<lpage>350</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/lingvan-2015-0013</pub-id></mixed-citation></ref>
<ref id="B41"><label>41</label><mixed-citation publication-type="confproc"><string-name><surname>Huang</surname>, <given-names>Luyao</given-names></string-name>, <string-name><given-names>Chi</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Xipeng</given-names> <surname>Qiu</surname></string-name> &amp; <string-name><given-names>Xuanjing</given-names> <surname>Huang</surname></string-name>. <year>2019</year>. <article-title>GlossBERT: BERT for Word Sense Disambiguation with Gloss Knowledge</article-title>. In <conf-name>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</conf-name>, <fpage>3509</fpage>&#8211;<lpage>3514</lpage>. <conf-loc>Hong Kong, China</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/D19-1355</pub-id></mixed-citation></ref>
<ref id="B42"><label>42</label><mixed-citation publication-type="confproc"><string-name><surname>Jawahar</surname>, <given-names>Ganesh</given-names></string-name>, <string-name><given-names>Beno&#238;t</given-names> <surname>Sagot</surname></string-name> &amp; <string-name><given-names>Djam&#233;</given-names> <surname>Seddah</surname></string-name>. <year>2019</year>. <article-title>What Does BERT Learn about the Structure of Language?</article-title> In <conf-name>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</conf-name>, <fpage>3651</fpage>&#8211;<lpage>3657</lpage>. <conf-loc>Florence, Italy</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/P19-1356</pub-id></mixed-citation></ref>
<ref id="B43"><label>43</label><mixed-citation publication-type="book"><string-name><surname>Kilgarriff</surname>, <given-names>Adam</given-names></string-name>. <year>2003</year>. <chapter-title>&#8220;I don&#8217;t believe in word senses&#8221;</chapter-title>. In <string-name><given-names>Brigitte</given-names> <surname>Nerlich</surname></string-name>, <string-name><given-names>Zazie</given-names> <surname>Todd</surname></string-name>, <string-name><given-names>Vimala</given-names> <surname>Herman</surname></string-name> &amp; <string-name><given-names>David D.</given-names> <surname>Clarke</surname></string-name> (eds.), <source>Polysemy</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>.</mixed-citation></ref>
<ref id="B44"><label>44</label><mixed-citation publication-type="journal"><string-name><surname>Kreitzer</surname>, <given-names>Anatol</given-names></string-name>. <year>1997</year>. <article-title>Multiple levels of schematization: A study in the conceptualization of space</article-title>. <source>Cognitive Linguistics</source> <volume>8</volume>(<issue>4</issue>). <fpage>291</fpage>&#8211;<lpage>326</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/cogl.1997.8.4.291</pub-id></mixed-citation></ref>
<ref id="B45"><label>45</label><mixed-citation publication-type="book"><string-name><surname>Lakoff</surname>, <given-names>George</given-names></string-name>. <year>1987</year>. <source>Women, fire, and dangerous things: what categories reveal about the mind</source>. <publisher-loc>Chicago</publisher-loc>: <publisher-name>The University of Chicago Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.7208/chicago/9780226471013.001.0001</pub-id></mixed-citation></ref>
<ref id="B46"><label>46</label><mixed-citation publication-type="book"><string-name><surname>Langacker</surname>, <given-names>Ronald W.</given-names></string-name> <year>1991</year>. <source>Foundations of Cognitive Grammar 2: Descriptive Application</source>. <publisher-loc>Stanford, CA</publisher-loc>: <publisher-name>Stanford University Press</publisher-name>.</mixed-citation></ref>
<ref id="B47"><label>47</label><mixed-citation publication-type="book"><string-name><surname>Langacker</surname>, <given-names>Ronald W.</given-names></string-name> <year>2010</year>. <source>Concept, Image, and Symbol: The Cognitive Basis of Grammar</source>. <publisher-name>Walter de Gruyter</publisher-name>.</mixed-citation></ref>
<ref id="B48"><label>48</label><mixed-citation publication-type="journal"><string-name><surname>Lee</surname>, <given-names>David</given-names></string-name>. <year>1998</year>. <article-title>A Tour through through</article-title>. <source>Journal of English Linguistics</source> <volume>26</volume>(<issue>4</issue>). <fpage>333</fpage>&#8211;<lpage>351</lpage>. DOI: <pub-id pub-id-type="doi">10.1177/007542429802600404</pub-id></mixed-citation></ref>
<ref id="B49"><label>49</label><mixed-citation publication-type="book"><string-name><surname>Lemmens</surname>, <given-names>Maarten</given-names></string-name>. <year>2016</year>. <chapter-title>Cognitive semantics</chapter-title>. In <string-name><given-names>Nick</given-names> <surname>Riemer</surname></string-name> (ed.), <source>The Routledge Handbook of Semantics</source>, <fpage>90</fpage>&#8211;<lpage>105</lpage>. <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>.</mixed-citation></ref>
<ref id="B50"><label>50</label><mixed-citation publication-type="journal"><string-name><surname>Lenci</surname>, <given-names>Alessandro</given-names></string-name>. <year>2018</year>. <article-title>Distributional Models of Word Meaning</article-title>. <source>Annual Review of Linguistics</source> <volume>4</volume>(<issue>1</issue>). <fpage>151</fpage>&#8211;<lpage>171</lpage>. DOI: <pub-id pub-id-type="doi">10.1146/annurev-linguistics-030514-125254</pub-id></mixed-citation></ref>
<ref id="B51"><label>51</label><mixed-citation publication-type="journal"><string-name><surname>Levy</surname>, <given-names>Omer</given-names></string-name>, <string-name><given-names>Yoav</given-names> <surname>Goldberg</surname></string-name> &amp; <string-name><given-names>Ido</given-names> <surname>Dagan</surname></string-name>. <year>2015</year>. <article-title>Improving Distributional Similarity with Lessons Learned from Word Embeddings</article-title>. <source>Transactions of the Association for Computational Linguistics</source> <volume>3</volume>. <fpage>211</fpage>&#8211;<lpage>225</lpage>. DOI: <pub-id pub-id-type="doi">10.1162/tacl_a_00134</pub-id></mixed-citation></ref>
<ref id="B52"><label>52</label><mixed-citation publication-type="confproc"><string-name><surname>Linzen</surname>, <given-names>Tal</given-names></string-name>, <string-name><given-names>Grzegorz</given-names> <surname>Chrupala</surname></string-name>, <string-name><given-names>Yonatan</given-names> <surname>Belinkov</surname></string-name> &amp; <string-name><given-names>Dieuwke</given-names> <surname>Hupkes</surname></string-name> (eds.). <year>2019</year>. <conf-name>Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP</conf-name>. <conf-loc>Florence, Italy</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>.</mixed-citation></ref>
<ref id="B53"><label>53</label><mixed-citation publication-type="journal"><string-name><surname>Newman</surname>, <given-names>John</given-names></string-name>. <year>2011</year>. <article-title>Corpora and cognitive linguistics</article-title>. <source>Revista Brasileira de Lingu&#237;stica Aplicada</source> <volume>11</volume>(<issue>2</issue>). <fpage>521</fpage>&#8211;<lpage>559</lpage>. DOI: <pub-id pub-id-type="doi">10.1590/S1984-63982011000200010</pub-id></mixed-citation></ref>
<ref id="B54"><label>54</label><mixed-citation publication-type="book"><string-name><surname>Oakley</surname>, <given-names>Todd</given-names></string-name>. <year>2010</year>. <chapter-title>Image Schemas</chapter-title>. In <string-name><given-names>Dirk</given-names> <surname>Geeraerts</surname></string-name> &amp; <string-name><given-names>Hubert</given-names> <surname>Cuyckens</surname></string-name> (eds.), <source>Handbook of Cognitive Linguistics</source>, <fpage>214</fpage>&#8211;<lpage>235</lpage>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford Univeristy Press</publisher-name>.</mixed-citation></ref>
<ref id="B55"><label>55</label><mixed-citation publication-type="journal"><string-name><surname>Pater</surname>, <given-names>Joe</given-names></string-name>. <year>2019</year>. <article-title>Generative linguistics and neural networks at 60: Foundation, friction, and fusion</article-title>. <source>Language</source> <volume>95</volume>(<issue>1</issue>). <fpage>e41</fpage>&#8211;<lpage>e74</lpage>. DOI: <pub-id pub-id-type="doi">10.1353/lan.2019.0009</pub-id></mixed-citation></ref>
<ref id="B56"><label>56</label><mixed-citation publication-type="book"><string-name><surname>Pawelec</surname>, <given-names>Andrzej</given-names></string-name>. <year>2010</year>. <source>Prepositional network models a hermeneutical case study</source>. <publisher-loc>Krakow</publisher-loc>: <publisher-name>Jagiellonian University Press</publisher-name>.</mixed-citation></ref>
<ref id="B57"><label>57</label><mixed-citation publication-type="confproc"><string-name><surname>Peters</surname>, <given-names>Matthew</given-names></string-name>, <string-name><given-names>Mark</given-names> <surname>Neumann</surname></string-name>, <string-name><given-names>Luke</given-names> <surname>Zettlemoyer</surname></string-name> &amp; <string-name><given-names>Wen-tau</given-names> <surname>Yih</surname></string-name>. <year>2018b</year>. <article-title>Dissecting Contextual Word Embeddings: Architecture and Representation</article-title>. In <conf-name>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</conf-name>, <fpage>1499</fpage>&#8211;<lpage>1509</lpage>. <conf-loc>Brussels, Belgium</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/D18-1179</pub-id></mixed-citation></ref>
<ref id="B58"><label>58</label><mixed-citation publication-type="confproc"><string-name><surname>Peters</surname>, <given-names>Matthew</given-names></string-name>, <string-name><given-names>Mark</given-names> <surname>Neumann</surname></string-name>, <string-name><given-names>Mohit</given-names> <surname>Iyyer</surname></string-name>, <string-name><given-names>Matt</given-names> <surname>Gardner</surname></string-name>, <string-name><given-names>Christopher</given-names> <surname>Clark</surname></string-name>, <string-name><given-names>Kenton</given-names> <surname>Lee</surname></string-name> &amp; <string-name><given-names>Luke</given-names> <surname>Zettlemoyer</surname></string-name>. <year>2018a</year>. <article-title>Deep Contextualized Word Representations</article-title>. In <conf-name>Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers</conf-name>), <fpage>2227</fpage>&#8211;<lpage>2237</lpage>. <conf-loc>New Orleans, Louisiana</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/N18-1202</pub-id></mixed-citation></ref>
<ref id="B59"><label>59</label><mixed-citation publication-type="book"><string-name><surname>Queller</surname>, <given-names>Kurt</given-names></string-name>. <year>2001</year>. <chapter-title>A usage-based approach to modeling and teaching the phrasal lexicon</chapter-title>. In <string-name><given-names>Dirk</given-names> <surname>Geeraerts</surname></string-name>, <string-name><given-names>Ren&#233;</given-names> <surname>Dirven</surname></string-name>, <string-name><given-names>John R.</given-names> <surname>Taylor</surname></string-name> &amp; <string-name><given-names>Ronald W.</given-names> <surname>Langacker</surname></string-name> (eds.), <source>Applied Cognitive Linguistics, II, Language Pedagogy</source>. <publisher-loc>Berlin, Boston</publisher-loc>: <publisher-name>De Gruyter</publisher-name>.</mixed-citation></ref>
<ref id="B60"><label>60</label><mixed-citation publication-type="book"><string-name><surname>Radford</surname>, <given-names>Alec</given-names></string-name>, <string-name><given-names>Jeffrey</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Rewon</given-names> <surname>Child</surname></string-name>, <string-name><given-names>David</given-names> <surname>Luan</surname></string-name>, <string-name><given-names>Dario</given-names> <surname>Amodei</surname></string-name> &amp; <string-name><given-names>Ilya</given-names> <surname>Sutskever</surname></string-name>. <year>2018</year>. <chapter-title>Language Models are Unsupervised Multitask Learners</chapter-title>. <publisher-loc>San Francisco, California, United States</publisher-loc>: <publisher-name>OpenAI</publisher-name>.</mixed-citation></ref>
<ref id="B61"><label>61</label><mixed-citation publication-type="journal"><string-name><surname>Rhee</surname>, <given-names>Seongha</given-names></string-name>. <year>2002</year>. <article-title>Semantic Changes of English Preposition against A Grammaticalization Perspective</article-title>. <source>Language Research</source> <volume>38</volume>(<issue>2</issue>). <fpage>563</fpage>&#8211;<lpage>583</lpage>.</mixed-citation></ref>
<ref id="B62"><label>62</label><mixed-citation publication-type="book"><string-name><surname>Rice</surname>, <given-names>Sally</given-names></string-name>. <year>1996</year>. <chapter-title>Prepositional prototypes</chapter-title>. In <string-name><given-names>Martin</given-names> <surname>P&#252;tz</surname></string-name> &amp; <string-name><given-names>Ren&#233;</given-names> <surname>Dirven</surname></string-name> (eds.), <source>The Construal of Space in Language and Thought</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>.</mixed-citation></ref>
<ref id="B63"><label>63</label><mixed-citation publication-type="book"><string-name><surname>Rice</surname>, <given-names>Sally A.</given-names></string-name> <year>1999</year>. <chapter-title>Aspects of prepositions and prepositional aspect</chapter-title>. In <string-name><given-names>Leon</given-names> <surname>de Stadler</surname></string-name> &amp; <string-name><given-names>Christoph</given-names> <surname>Eyrich</surname></string-name> (eds.), <source>Issues in Cognitive Linguistics</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>.</mixed-citation></ref>
<ref id="B64"><label>64</label><mixed-citation publication-type="book"><string-name><surname>Sagi</surname>, <given-names>Eyal</given-names></string-name>, <string-name><given-names>Stefan</given-names> <surname>Kaufmann</surname></string-name> &amp; <string-name><given-names>Brady</given-names> <surname>Clark</surname></string-name>. <year>2011</year>. <chapter-title>Tracing semantic change with Latent Semantic Analysis</chapter-title>. In <string-name><given-names>Kathryn</given-names> <surname>Allan</surname></string-name> &amp; <string-name><given-names>Justyna A.</given-names> <surname>Robinson</surname></string-name> (eds.), <source>Current Methods in Historical Semantics</source>. <publisher-loc>Berlin, Boston</publisher-loc>: <publisher-name>De Gruyter</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1515/9783110252903.161</pub-id></mixed-citation></ref>
<ref id="B65"><label>65</label><mixed-citation publication-type="journal"><string-name><surname>Sandra</surname>, <given-names>Dominiek</given-names></string-name>. <year>1998</year>. <article-title>What linguists can and can&#8217;t tell you about the human mind: A reply to Croft</article-title>. <source>Cognitive Linguistics</source> <volume>9</volume>(<issue>4</issue>). <fpage>361</fpage>&#8211;<lpage>378</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/cogl.1998.9.4.361</pub-id></mixed-citation></ref>
<ref id="B66"><label>66</label><mixed-citation publication-type="journal"><string-name><surname>Sandra</surname>, <given-names>Dominiek</given-names></string-name> &amp; <string-name><given-names>Sally</given-names> <surname>Rice</surname></string-name>. <year>1995</year>. <article-title>Network analyses of prepositional meaning: Mirroring whose mind&#8212;the linguist&#8217;s or the language user&#8217;s?</article-title> <source>Cognitive Linguistics</source> <volume>6</volume>(<issue>1</issue>). <fpage>89</fpage>&#8211;<lpage>130</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/cogl.1995.6.1.89</pub-id></mixed-citation></ref>
<ref id="B67"><label>67</label><mixed-citation publication-type="confproc"><string-name><surname>Sommerauer</surname>, <given-names>Pia</given-names></string-name> &amp; <string-name><given-names>Antske</given-names> <surname>Fokkens</surname></string-name>. <year>2018</year>. <article-title>Firearms and Tigers are Dangerous, Kitchen Knives and Zebras are Not: Testing whether Word Embeddings Can Tell</article-title>. In <conf-name>Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP</conf-name>, <fpage>276</fpage>&#8211;<lpage>286</lpage>. <conf-loc>Brussels, Belgium</conf-loc>: <conf-sponsor>Association for Computational Linguistics</conf-sponsor>. DOI: <pub-id pub-id-type="doi">10.18653/v1/W18-5430</pub-id></mixed-citation></ref>
<ref id="B68"><label>68</label><mixed-citation publication-type="book"><string-name><surname>Stefanowitsch</surname>, <given-names>Anatol</given-names></string-name>. <year>2010</year>. <chapter-title>Empirical cognitive semantics: Some thoughts</chapter-title>. In <string-name><given-names>Dylan</given-names> <surname>Glynn</surname></string-name> &amp; <string-name><given-names>Kerstin</given-names> <surname>Fischer</surname></string-name> (eds.), <source>Quantitative Methods in Cognitive Semantics: Corpus-Driven Approaches</source>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1515/9783110226423.355</pub-id></mixed-citation></ref>
<ref id="B69"><label>69</label><mixed-citation publication-type="book"><string-name><surname>Taylor</surname>, <given-names>John R.</given-names></string-name> <year>2006</year>. <chapter-title>Polysemy and the lexicon</chapter-title>. In <string-name><given-names>Gitte</given-names> <surname>Kristiansen</surname></string-name>, <string-name><given-names>Michel</given-names> <surname>Achard</surname></string-name>, <string-name><given-names>Ren&#233;</given-names> <surname>Dirven</surname></string-name> &amp; <string-name><given-names>Francisco J.</given-names> <surname>Ruiz Mendoza Ib&#225;&#241;ez</surname></string-name> (eds.), <source>Cognitive Linguistics: Current Applications and Future Perspectives</source>, <fpage>51</fpage>&#8211;<lpage>80</lpage>. <publisher-loc>Berlin</publisher-loc>: <publisher-name>Mouton De Gruyter</publisher-name>.</mixed-citation></ref>
<ref id="B70"><label>70</label><mixed-citation publication-type="journal"><string-name><surname>Turney</surname>, <given-names>Peter D.</given-names></string-name> &amp; <string-name><given-names>Patrick</given-names> <surname>Pantel</surname></string-name>. <year>2010</year>. <article-title>From Frequency to Meaning: Vector Space Models of Semantics</article-title>. <source>Journal of Artificial Intelligence Research</source> <volume>37</volume>. <fpage>141</fpage>&#8211;<lpage>188</lpage>. DOI: <pub-id pub-id-type="doi">10.1613/jair.2934</pub-id></mixed-citation></ref>
<ref id="B71"><label>71</label><mixed-citation publication-type="journal"><string-name><surname>Tyler</surname>, <given-names>Andrea</given-names></string-name> &amp; <string-name><given-names>Vyvyan</given-names> <surname>Evans</surname></string-name>. <year>2001</year>. <article-title>Reconsidering Prepositional Polysemy Networks: The Case of Over</article-title>. <source>Language</source> <volume>77</volume>(<issue>4</issue>). <fpage>724</fpage>&#8211;<lpage>765</lpage>. DOI: <pub-id pub-id-type="doi">10.1353/lan.2001.0250</pub-id></mixed-citation></ref>
<ref id="B72"><label>72</label><mixed-citation publication-type="book"><string-name><surname>Tyler</surname>, <given-names>Andrea</given-names></string-name> &amp; <string-name><given-names>Vyvyan</given-names> <surname>Evans</surname></string-name>. <year>2003</year>. <source>The Semantics of English Prepositions: Spatial Scenes, Embodied Meaning, and Cognition</source>. <publisher-name>Cambridge University Press</publisher-name> <edition>1st edn.</edition> DOI: <pub-id pub-id-type="doi">10.1017/CBO9780511486517</pub-id></mixed-citation></ref>
<ref id="B73"><label>73</label><mixed-citation publication-type="journal"><string-name><surname>van der Maaten</surname>, <given-names>Laurens</given-names></string-name> &amp; <string-name><given-names>Geoffrey</given-names> <surname>Hinton</surname></string-name>. <year>2008</year>. <article-title>Visualizing Data using t-SNE</article-title>. <source>Journal of Machine Learning Research</source> <volume>9</volume>. <fpage>2579</fpage>&#8211;<lpage>2605</lpage>.</mixed-citation></ref>
<ref id="B74"><label>74</label><mixed-citation publication-type="journal"><string-name><surname>Vanpaemel</surname>, <given-names>Wolf</given-names></string-name> &amp; <string-name><given-names>Gert</given-names> <surname>Storms</surname></string-name>. <year>2008</year>. <article-title>In search of abstraction: The varying abstraction model of categorization</article-title>. <source>Psychonomic Bulletin &amp; Review</source> <volume>15</volume>(<issue>4</issue>). <fpage>732</fpage>&#8211;<lpage>749</lpage>. DOI: <pub-id pub-id-type="doi">10.3758/PBR.15.4.732</pub-id></mixed-citation></ref>
<ref id="B75"><label>75</label><mixed-citation publication-type="confproc"><string-name><surname>Vaswani</surname>, <given-names>Ashish</given-names></string-name>, <string-name><given-names>Noam</given-names> <surname>Shazeer</surname></string-name>, <string-name><given-names>Niki</given-names> <surname>Parmar</surname></string-name>, <string-name><given-names>Jakob</given-names> <surname>Uszkoreit</surname></string-name>, <string-name><given-names>Llion</given-names> <surname>Jones</surname></string-name>, <string-name><given-names>Aidan N.</given-names> <surname>Gomez</surname></string-name>, <string-name><given-names>Lukasz</given-names> <surname>Kaiser</surname></string-name> &amp; <string-name><given-names>Illia</given-names> <surname>Polosukhin</surname></string-name>. <year>2017</year>. <article-title>Attention Is All You Need</article-title>. In <conf-name>31st conference on neural information processing systems (nips 2017)</conf-name>. <conf-loc>CA, USA</conf-loc>: <conf-sponsor>Long Beach</conf-sponsor>.</mixed-citation></ref>
<ref id="B76"><label>76</label><mixed-citation publication-type="confproc"><string-name><surname>Verbeemen</surname>, <given-names>Timothy</given-names></string-name>, <string-name><given-names>Gert</given-names> <surname>Storms</surname></string-name> &amp; <string-name><given-names>Tom</given-names> <surname>Verguts</surname></string-name>. <year>2005</year>. <article-title>Varying Abstraction in Categorization: a K-means Approach</article-title>. In <conf-name>Proceedings of the 27th Annual Conference of the Cognitive Science Society</conf-name>, <fpage>2301</fpage>&#8211;<lpage>2306</lpage>. <conf-loc>Mahwah, NJ</conf-loc>: <conf-sponsor>Erlbaum</conf-sponsor>.</mixed-citation></ref>
<ref id="B77"><label>77</label><mixed-citation publication-type="journal"><string-name><surname>Wiedemann</surname>, <given-names>Gregor</given-names></string-name>, <string-name><given-names>Steffen</given-names> <surname>Remus</surname></string-name>, <string-name><given-names>Avi</given-names> <surname>Chawla</surname></string-name> &amp; <string-name><given-names>Chris</given-names> <surname>Biemann</surname></string-name>. <year>2019</year>. <article-title>Does BERT Make Any Sense? Interpretable Word Sense Disambiguation with Contextualized Embeddings</article-title>. <source>arXiv:1909.10430</source>.</mixed-citation></ref>
<ref id="B78"><label>78</label><mixed-citation publication-type="journal"><string-name><surname>Young</surname>, <given-names>Tom</given-names></string-name>, <string-name><given-names>Devamanyu</given-names> <surname>Hazarika</surname></string-name>, <string-name><given-names>Soujanya</given-names> <surname>Poria</surname></string-name> &amp; <string-name><given-names>Erik</given-names> <surname>Cambria</surname></string-name>. <year>2018</year>. <article-title>Recent Trends in Deep Learning Based Natural Language Processing</article-title>. <source>arXiv:1708.02709</source>. DOI: <pub-id pub-id-type="doi">10.1109/MCI.2018.2840738</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>