1 Background
Filler-gap dependencies are syntactic constructions in which a filler phrase is displaced from its canonical position, leaving behind a gap (typically represented by an underscore). However, as first observed by Ross (1967), certain syntactic structures block the formation of such dependencies. These structures are known as islands, and extraction from them typically results in degraded acceptability. Consider the following examples of island violations:
- (1)
- a.
- Complex noun phrase constraint(Ross 1967)
- *What did you hear [the rumor that John bought __]?
- b.
- Adjunct island(Huang 1982)
- *What did John leave [because Mary was reading __]?
- c.
- Wh-island(Chomsky 1973)
- *What did you wonder [whether Mary bought __]?
The source of island effects remains a topic of central debate. Early accounts attribute these effects to violations of syntactic constraints (see, e.g., Chomsky 1973; 1986; Huang 1982). Specifically, filler phrases located within island structures are claimed to be inaccessible to the syntactic operations that establish filler-gap dependencies (i.e. movement operations). Under this traditional view, island-violating dependencies are predicted to be sharply degraded, reflecting a clear contrast between structurally licit and illicit configurations.
However, the acceptability of island-violating dependencies appears to be gradient, with some violations judged more acceptable than others; numerous counterexamples in the literature show that apparent violations of classic island constraints can yield relatively acceptable sentences (see Ross 1967; Boeckx 2012). These observations have led many researchers to propose that island effects may instead arise from processing difficulties (Hofmeister & Sag 2010; Liu et al. 2022). In other words, these accounts argue that the low acceptability of island-violating dependencies stems not from the violations of syntactic constraints, but from the difficulty of processing these structures in real time.1
A dominant proposal within this approach is the memory-based theory (Kluender 1991; Kluender & Kutas 1993a; 1993b; Hofmeister & Sag 2010), which attributes the unacceptability of island violations to excessive demands on working memory. Building a filler–gap dependency requires the parser to hold the filler phrase in working memory until it reaches the gap site. At that point, the filler is retrieved and interpreted in its canonical position (Kluender 1991; Kluender & Kutas 1993a; 1993b). In island contexts, this process becomes more taxing for working memory resources. The parser must not only keep the filler phrase active while searching for a potential gap, but also process intervening clause boundaries and handle the semantic complexity of island structures. Together, these demands consume additional cognitive resources and make dependency resolution more difficult (Sag & Hofmeister & Casasanto 2013). As a result, the parser’s limited working memory capacity is overloaded, leading to processing breakdown and, consequently, degraded acceptability judgments.2
To disentangle grammatical constraints from extra-grammatical sources of difficulty, experimental work on island effects has developed paradigms designed to factor out independent processing costs associated with filler–gap dependencies. A prominent example is the factorial, or superadditivity, paradigm (Sprouse 2007; Kush & Lohndal & Sprouse 2018), which seeks to quantify the contribution of processing costs—such as dependency length and structural complexity—and assess whether island violations incur an additional penalty beyond these factors. Although superadditive effects have often been interpreted as evidence for a grammatical contribution to island effects, subsequent critiques have emphasized that such interactions are not diagnostic of syntactic constraints per se, as similar patterns could arise from discourse-related limitations or overloading the processing system (Liu et al. 2022).
Additionally, online processing studies have examined how comprehenders construct filler–gap dependencies in real time. According to the active filler strategy, encountering a filler prompts the parser to search for a suitable gap as soon as a grammatically plausible position becomes available (Stowe 1986; Frazier & d’Arcais 1989; Traxler & Pickering 1996). Evidence for this strategy comes from filled-gap effects, in which reading times increase when an anticipated gap position is unexpectedly occupied. Importantly, however, such effects are often absent in island configurations. While syntactic accounts interpret this pattern as evidence that grammatical knowledge constrains online parsing, processing-based accounts attribute it to the suspension of active dependency formation when processing costs become too high (Phillips 2013).
A central empirical finding often cited in support of the memory-based accounts is the phenomenon of island amelioration—the observation that the acceptability of island-violating dependencies improves by changing non-structural factors. For instance, discourse-linked (D-linked) fillers such as which student are judged more acceptable than bare wh-words like what or who. This effect has been consistently observed across experimental studies: Hofmeister & Sag (2010) and Hofmeister (2011) demonstrated that semantically rich filler phrases significantly improve both reading times and acceptability ratings of island-violating dependencies in English; Goodall (2014) further observed that D-linked wh-phrases improve acceptability in structurally complex dependencies, including those that cross island boundaries; and Alexopoulou & Keller (2013) found similar results in English and Greek. Recent findings from Baha Arabic (Hayyas 2020) replicated this pattern.
Memory-based accounts argue that this amelioration effect arises from the greater accessibility in memory of referentially salient, semantically rich and individual-denoting fillers (Alexopoulou & Keller 2013; Sag et al. 2013), all of which eases the retrieval and integration process at the gap site. Crucially, these results of island amelioration studies challenge structural accounts and support the view that processing-related factors—such as memory accessibility—contribute to the unacceptability of island-violating dependencies.
Beyond D-linking, another linguistic feature that may influence the acceptability and processing of island-violating dependencies is animacy. Because animate fillers are referentially salient and individual denoting (i.e. identify individuals rather than kinds) (Alexopoulou & Keller 2013), they may be easier to maintain and retrieve during processing, thereby reducing memory demands. Alexopoulou & Keller (2013) found that animate fillers led to higher acceptability ratings in whether-island contexts in both English and Greek. Similarly, Atkinson et al. (2016) show that the semantic distinctness conferred by animacy improves the acceptability of wh-island violations, though the effect is selective and relatively weak. Converging evidence from Villata & Franck (2024) indicates that animacy similarity exerts a mild, graded influence on wh-island judgments, consistent with a processing-based explanation. Together, these findings suggest that animacy can facilitate dependency formation in island environments, but that its contribution is modest and likely processing-related rather than structural.
To date, the role of animacy in island contexts has rarely been tested cross-linguistically, with the notable exception of Alexopoulou & Keller’s (2013) study on English and Greek and Villata & Franck’s (2024) study on French. In the context of Arabic theoretical syntax, many analyses distinguish between Arabic who and what not only in terms of animacy, but also in terms of their internal syntactic structure: who is analyzed as bearing a [+D] feature, whereas what patterns as a bare wh-NP (Aoun et al. 2009). See (2).
- (2)
- a.
- what

- b.
- which-phrases and who

Importantly, in the Arabic syntactic literature this distinction is not treated as a purely semantic one but is argued to have concrete morphosyntactic consequences. Wh-expressions that project a DP layer—such as who and which-N phrases—are able to participate in binding-based dependencies, whereas bare wh-NPs like what lack this option.
This distinction has been proposed primarily on the basis of differences in how wh-expressions participate in dependency formation, especially in their interaction with resumptive pronouns (Aoun et al. 2009). Across Arabic varieties, resumption is licensed with who and which-N fillers, but is systematically disallowed with what (Aoun et al. 2009).3 See (3).
- (3)
- Lebanese Arabic
- a.
- *šu
- what
- Talabit-o
- ordered.3SG.F-it
- laila
- Laila
- b-l-maTʕam?
- in-the-restaurant
- ‘What did Laila order at the restaurant?’
- b.
- miin/ʔayya
- who/which
- maariḍ
- patient
- zarit-u
- visit.3SG.F-him
- naadia?
- Nadia
- ‘Who/which patient did Nadia visit?’
- Modern Standard Arabic
- a.
- *maaðaa
- what
- ʔištarat-hu
- bought.3SG.F-it
- laila
- Laila
- min al-maktabati?
- from the-bookstore
- ‘What did Laila buy from the bookstore?’
- b.
- man/ʔayya
- Who/which
- mariiDin
- patient
- zaarat-hu
- visited.3SG.F-him
- naadia?
- Nadia
- ‘Who/which patient did Nadia visit?’
- (Aoun et al. 2009: 132, 135)
Thus, the dominant view holds that what-questions violating island constraints are ungrammatical, regardless of whether the dependency terminates with a gap or a resumptive pronoun: gap dependencies violate movement constraints, while resumptive pronouns remain unavailable because what lacks the DP-layer required to license a binding relation with the pronoun; who-questions violating island constraints are also highly unacceptable due to violations of movement constraints—unless the dependency is resolved by a resumptive pronoun, which licenses a binding relation rather than illicit movement (Aoun et al. 2009).4 See (4).
- (4)
- a.
- Gap (ungrammatical)
- *miin/ʔayya
- who/which
- mariiḍ
- patient
- btaʕrfo
- know.2PL
- l-mara
- the-woman
- yalli
- that
- zeerit __?
- visited.3SG.F__
- (‘Who/which patient do you know the woman that visited __?’)
- b.
- Resumptive pronoun (grammatical)
- miin/ʔayya
- who/which
- mariiḍ
- patient
- btaʕrfo
- know.2PL
- l-mara
- the-woman
- yalli
- that
- zeerit-o?
- visited.3SG.F-him
- (‘Who/which patient do you know the woman that visited him?’)
- (Aoun et al. 2009: 146)
Despite this theoretical claim, experimental studies on island effects in Arabic have generally treated animate and inanimate fillers (e.g., who vs. what) in context of islands as a uniform class of bare wh-elements, without examining whether they differ in syntactic status or processing behaviour (see, e.g., Tucker et al. (2019) on Modern Standard Arabic, and Al-Aqarbeh & Sprouse (2023) on Jordanian Arabic). If our results show that animacy modulates acceptability or processing in the same way that D-linked fillers do, this would suggest that previous studies may have overlooked an important variable; the failure to control for animacy could undermine the generalizability of their conclusions.
Although previous experimental work points to an animacy effect, the mechanism underlying this effect remains uncertain. It is not clear whether animacy facilitates the real-time processing of long-distance dependencies, as predicted by memory-based accounts (a possibility that, notably, has not been directly tested), whether it contributes to licensing a syntactic binding relation with resumptive pronouns, or whether it primarily enhances interpretability at a later, discourse-level stage. We return to this latter possibility in the Discussion.
Furthermore, prior evidence for animacy-related modulation of acceptability has come primarily from studies of wh-islands, a domain in which gradient acceptability and partial amelioration effects are well documented (e.g., Almeida 2014). In this literature, animate fillers have been argued to improve acceptability either by increasing discourse prominence—since they refer to individuated entities rather than kinds (Alexopoulou & Keller 2013)—or by enhancing cue distinctiveness during dependency retrieval (Atkinson et al. 2016; Villata & Franck 2024). These findings raise the possibility that animacy plays a general role in mitigating island effects. At the same time, because this work has focused primarily on wh-islands, it leaves open the question of whether animacy-based amelioration extends to other island configurations, including adjunct islands.
Adjunct islands have long been treated as among the strongest extraction domains, both because they are structurally opaque and because they typically involve backgrounded or non-argumental material (Huang 1982; Cinque 1990). However, experimental work has increasingly suggested that adjunct islands are not uniformly resistant to amelioration. For example, extraction from adjuncts can improve under discourse-supporting contexts (Gibson et al. 2021) and extraction from adjunct clauses is often judged more acceptable in relative clause constructions than in interrogatives (Sprouse et al. 2016; Kush et al. 2018). Moreover, repeated exposure has been shown to yield satiation effects in adjunct islands (Chaves & Putnam 2020). These findings suggest that adjunct island effects are not uniformly categorical and are sensitive to non-structural factors.
Importantly, research on Arabic varieties indicates that adjunct islands are among the island types most susceptible to improvement under resumption. Al-Aqarbeh & Sprouse (2023) report that in Jordanian Arabic, resumptive pronouns yield their strongest amelioration effects in adjunct island environments relative to other island configurations. Similarly, Tucker et al. (2019) show that adjunct islands are more strongly ameliorated by resumption in D-linked wh-questions than subject, complex NP, or coordinate structure islands. Together, these findings suggest that adjunct islands constitute a particularly informative testing ground for examining whether structural or processing factors modulate island effects.
At the same time, prior work has shown that acceptability in adjunct island configurations is sensitive to multiple interacting factors, including properties of the adjunct and the discourse relation between the adjunct and the matrix event (Truswell 2011; Chaves & Putnam 2020). These considerations motivate careful experimental control, which is implemented in the design described below.
The present study builds on previous literature by examining animacy effects in adjunct islands in Saudi Arabic, thereby extending prior work on wh-islands to a class of island configurations that has not previously been investigated in this context. More specifically, we ask whether animacy influences the acceptability and processing of adjunct island-violating dependencies in Saudi Arabic—and whether this effect depends on the type of tail involved (gap vs. RP). Crucially, we aim to uncover not just if animacy ameliorates island effects, but how: is it helping to structurally license the dependency, or simply making the sentence easier to process?
According to syntactic accounts of Arabic, such as Aoun et al. (2009), the grammaticality of wh-dependencies is determined by the mechanism used to establish the dependency.5 Dependencies formed via movement are subject to island constraints and are therefore predicted to be unacceptable when they cross island boundaries. By contrast, dependencies established via binding—specifically through resumptive pronouns—are not subject to these movement constraints and may be licit in island contexts.
Within this framework, wh-questions that involve gap dependencies crossing adjunct island boundaries are structurally illicit, irrespective of the animacy or referential specificity of the filler. Because the dependency is established through movement, such configurations are predicted to be unacceptable in adjunct island environments. Wh-phrases like who, which are argued to project a DP layer (Aoun et al. 2009), differ in that they can enter into a binding relation with a resumptive pronoun. Only such binding-based dependencies are therefore predicted to yield acceptable island-violating structures.
Crucially, under this account, acceptability is determined by the availability of a licit dependency-forming mechanism rather than by properties of the filler itself. As a result, the theory predicts a clear contrast between structurally licit, binding-based dependencies (who-questions with resumptive pronouns) and illicit movement-based dependencies (all other conditions), with no independent role for animacy in modulating the acceptability of gap-based island violations.
In contrast, memory-based accounts predict that animate fillers, by virtue of their referential salience, individual-denoting properties, and increased accessibility in working memory (Alexopoulou & Keller 2013), should give rise to partial amelioration of island violations. On this view, animacy is expected to facilitate the retrieval and integration of the filler at the dependency integration site, resulting in modestly higher acceptability relative to inanimate fillers. This effect is predicted to be most pronounced when the dependency terminates in a resumptive pronoun, which may support dependency resolution by overtly marking the integration site and facilitating the linking of the filler to its interpretive position (McDaniel & Cowart 1999; Chacón 2019; Hammerly 2021). Importantly, resumptive pronouns in Saudi Arabic are morphologically specified for gender and number but not animacy and therefore do not themselves provide animacy-based retrieval cues. Nevertheless, the presence of an overt integration site provided by the resumptive pronoun is expected to strengthen the facilitatory effect of animacy in RP conditions compared to gap conditions.
To evaluate these predictions, we conducted a self-paced reading experiment alongside an acceptability judgment task. We limited our investigation to adjunct island–violating dependencies, as they constitute a syntactic environment in which movement-based dependencies are disallowed and island effects are predicted to be robust under structural accounts. This makes adjunct islands a particularly informative testing ground for assessing whether animacy contributes to the structural licensing of the dependency—via resumption—or instead facilitates dependency resolution through processing mechanisms that give rise to partial amelioration in acceptability.
2 The experiment
2.1 Materials
All experimental items were wh-questions in which the wh-dependency crossed a temporal adjunct island boundary. Adjunct type was held constant across all materials to avoid confounds associated with variation across adjunct classes (e.g., temporal vs. manner or reason adjuncts), which have been shown to differ in their susceptibility to extraction and amelioration effects (Chaves & Putnam 2020).
The materials were further constructed to ensure discourse relevance between the matrix predicate and the event described in the adjunct clause. Across all conditions, the adjunct event plausibly motivated the emotional or evaluative predicate in the matrix clause, yielding coherent event interpretations and minimizing the likelihood that pragmatic infelicity or event incoherence independently influenced acceptability judgments or processing measures (Truswell 2011; Chaves & Putnam 2020).
The experiment was not designed to investigate satiation effects. Lexicalization, repetition, and spacing between critical items were not controlled in a manner that would permit assessment of exposure-based changes in acceptability (Chaves & Putnam 2020); accordingly, no claims regarding satiation are made on the basis of the present data.
The experiment employed a 2 × 2 within-subjects design crossing the factors Animacy (animate vs. inanimate) and Tail (gap vs. resumptive pronoun), yielding four experimental conditions. The materials consisted of eight items (hereafter referred to as lexicalizations in the statistical analyses), each corresponding to a single base sentence frame with identical lexical content across conditions. Each item was instantiated in all four experimental conditions, resulting in 32 experimental sentences in total. Each participant was presented with 16 experimental sentences (four per condition), together with 32 filler sentences, for a total of 48 trials per participant.
All experimental sentences were bi-clausal to control for structural complexity and to keep dependency length constant across conditions. The same lexicalizations were used across conditions to isolate the effects of the manipulated factors. A three-word adjunct phrase (preposition + DP) was included following the dependency tail as a controlled spillover region, allowing potential reading-time effects extending beyond the integration site to be captured.
A summary of the experimental conditions is illustrated in Example (5), and the full list of experimental sentences is presented in Appendix A. Sentences predicted to be unacceptable under the Arabic theoretical literature are marked with an asterisk (*).
- (5)
- Representative experimental stimuli
- a.
- Animate filler – RP
- meen
- who
- tafaʔʒaʔti
- got.surprised-you
- lamma
- when
- al-mudīra
- the-principal
- intaqadat-hu
- criticized-him
- fī
- in
- ijtimāʕ
- meeting
- al-idāra?
- the-administration
- (‘Who did you get surprised about when the principal criticized him in the administration meeting?’)
- b.
- Animate filler – Gap
- *meen
- who
- tafaʔʒaʔti
- got.surprised-you
- lamma
- when
- al-mudīra
- the-principal
- intaqadat __
- criticized __
- fī
- in
- ijtimāʕ
- meeting
- al-idāra?
- the-administration
- (‘Who did you get surprised about when the principal criticized __ in the administration meeting?’)
- c.
- Inanimate filler – RP
- *ʔayš
- what
- tafaʔʒaʔti
- got.surprised-you
- lamma
- when
- al-mudīra
- the-principal
- intaqadat-hu
- criticized-it
- fī
- in
- ijtimāʕ
- meeting
- al-idāra?
- the-administration
- (‘What did you get surprised about when the principal criticized it in the administration meeting?’)
- d.
- Inanimate filler – Gap
- *ʔayš
- what
- tafaʔʒaʔti
- got.surprised-you
- lamma
- when
- al-mudīra
- the-principal
- intaqadat __
- criticized __
- fī
- in
- ijtimāʕ
- meeting
- al-idāra?
- the-administration
- (‘What did you get surprised about when the principal criticized __ in the administration meeting?’)
The filler items were also bi-clausal and matched the experimental sentences in length and structure. Half were fully acceptable (e.g., 6), and half were ungrammatical, containing a range of violations such as subject–verb agreement errors (7) and subcategorization violations (8). Including both grammatical and ungrammatical fillers ensured that participants were exposed to the full range of acceptability, encouraging them to use the entire rating scale and reducing the likelihood of response bias.
- (6)
- Acceptable filler sentence
- ʔēsh
- what
- al-ʕilāj
- the-treatment
- illi
- that
- ad-duktūrah
- the-doctor.F
- waṣafat-hu
- prescribed-3SG.F-it
- li-Sumayya?
- for-Sumayya
- ‘What is the treatment that the doctor prescribed for Sumayya?’
- (7)
- Subject–verb agreement error
- *ʔayy
- which
- bint
- girl
- fāz
- won.M
- fī
- in
- as-sibāq
- the-race
- al-nihāʔī
- the-final
- ʕalā
- on
- kass
- cup
- al-malik?
- the-king
- (Intended: ‘Which girl won(F) the final race for the King’s Cup?’)
- (8)
- Subcategorization violation
- *ʔummi
- my.mother
- tibġa
- wants
- tiʕrif
- to.know
- meen
- who
- ḏākart-i-ha
- studied-you-it
- maʕ
- with
- ṣāḥibāt-ik
- friends-your
- ams?
- yesterday
- (Intended: ‘My mother wants to know who you studied her with your friends yesterday.’)
- Violation: The object clitic -ha (‘her’) introduces an argument with the wrong semantic type: the verb ḏākara ‘study’ does not license a human object, yielding a subcategorization (selectional restriction) violation.
In total, 32 filler items were included. The fillers were evenly divided between grammatical (n = 16) and ungrammatical (n = 16), yielding a 1:1 ratio of grammatical-to-ungrammatical fillers. This ensured participants were regularly exposed to both acceptable and unacceptable sentences.6
2.2 Participants
Eighty-three native speakers of Saudi Arabic (aged 18–25) participated in the study. All participants were undergraduate students enrolled in an English-language program and received course credit for their participation. None had received formal training in linguistics at the time of participation, as linguistics courses are taken in later stages of the degree.
Participants were fluent speakers who reported using Saudi Arabic regularly in their daily lives. They were informed about the purpose of the study and their rights as participants, and all gave informed consent before beginning the task.
2.3 Methodology
The study employed a repeated-measures design, with each participant exposed to all four experimental conditions. To control for lexical repetition, carryover, and order effects, we implemented Latin Square counterbalancing. The 32 critical sentences (8 per condition) were distributed across eight pseudo-randomized lists, ensuring that no item ‘lexicalization’ appeared more than once within any list.
Each participant completed one of two online experiments, each consisting of four lists. Across these lists, each item ‘lexicalization’ appeared in two different forms, but no participant saw more than one version of a given item ‘lexicalisation’ in one list. Trial order was randomized individually within each list. All sentences (critical and fillers) were presented word-by-word in the SPR task and were followed by an acceptability judgment.
The experiment was administered online using PsychoPy (version 2023.2.1) (Peirce et al. 2019) and hosted on Pavlovia (Open Science Tools Ltd. 2023). Participants completed the task in approximately 25 minutes. The self-paced reading task used a moving-window design, where sentences were read one word at a time. After each sentence, participants rated its acceptability on a 7-point Likert scale (1 = completely unacceptable, 7 = fully acceptable). The instructions emphasized that participants should base their judgments on their own intuitions.
2.4 Results
2.4.1 Acceptability judgment analysis
Before turning to the experimental results, we note that participants performed the task as expected: grammatical fillers received high acceptability ratings (mean = 6.43), whereas ungrammatical fillers were rated substantially lower (mean = 2.77). A full analysis of filler responses is reported at the end of this section.
We now turn to the descriptive results of the acceptability judgment task. The two factors manipulated in this study were Animacy (Animate vs. Inanimate) and Tail (Gap vs. RP). The primary research question was whether animacy improves the acceptability of island-violating dependencies, and whether this improvement is modulated by the type of dependency tail. Table 1 presents the mean acceptability ratings across the four experimental conditions.
Table 1: Mean acceptability ratings by Animacy and Tail.
| Animacy | Tail | Mean | SD | N |
| Animate | Gap | 2.64 | 1.93 | 332 |
| Animate | RP | 3.20 | 2.17 | 332 |
| Inanimate | Gap | 2.50 | 1.89 | 332 |
| Inanimate | RP | 2.42 | 1.80 | 332 |
As shown, Animate–RP conditions yielded the highest mean acceptability, while Inanimate–RP conditions received the lowest scores. Gap conditions were rated lower overall, but ratings for Animate–Gap were slightly higher than for Inanimate–Gap.
To visualize these trends, raw ratings were plotted for each Animacy × Tail condition in Figure 1.
To test whether the observed patterns were statistically reliable, we fitted a series of ordinal cumulative link mixed models (CLMMs) using the ordinal package in RStudio. Likelihood ratio tests showed that adding Tail significantly improved model fit (p = .002), adding Animacy provided a further improvement (p < .001), and including the Animacy × Tail interaction yielded an additional significant improvement (p = .0004). Random intercepts were included for Participant and Lexicalisation. Post-hoc comparisons confirmed that adding random slopes did not improve model fit (all p > .55), so the simpler random-intercepts-only model was retained. The final model thus included both main effects and their interaction, with random intercepts:
Acceptability ~ Animacy * Tail + (1|Participant) + (1|Lexicalisation)
A summary of the fixed effects is given in Table 2.
Table 2: Fixed effects from the best-fitting CLMM (Inanimate and Gap were reference levels).
| Predictor | Estimate | Std. Error | z-value | p-value |
| Tail: RP | –0.053 | 0.156 | –0.342 | 0.732 |
| Animacy: Animate | 0.210 | 0.155 | 1.355 | 0.175 |
| RP × Animate Interaction | 0.770 | 0.219 | 3.522 | 0.0004 |
To aid interpretation, Figure 2 plots the predicted probabilities of each acceptability rating (1–7) across conditions. Each vertical line corresponds to a specific Animacy × Tail condition. Points where vertical lines intersect the rating curves represent the model-estimated probability of each rating value.
Figure 2: Ordinal cumulative link model showing predicted probabilities of each acceptability rating (1–7) for every Animacy × Tail condition. The horizontal axis indexes discrete experimental conditions; the curves represent model-estimated probabilities rather than trends over a continuous dimension. Conditions that appear close in the figure have similar predicted probability distributions across rating levels.
The plot revealed a clear pattern:
Inanimate–RP and Inanimate–Gap were the least acceptable conditions, with a high probability of receiving the lowest rating.
Animate–Gap was rated slightly higher; the probability of receiving the lowest rating decreased compared to Inanimate–RP and Inanimate–Gap.
Animate–RP produced the lowest probability of low ratings.
Post-hoc Tukey-adjusted pairwise comparisons confirmed this pattern (Table 3). Animate–RP was rated significantly higher than Animate–Gap (p < .0001) and Inanimate–RP (p < .0001), whereas no reliable differences emerged between Inanimate–Gap and Inanimate–RP, or between Inanimate–Gap and Animate–Gap.
Table 3: Post-hoc contrasts.
| Contrast | Estimate | SE | z-ratio | p-value |
| Inanimate Gap – Animate Gap | –0.210 | 0.155 | –1.355 | 0.528 |
| Inanimate Gap – Inanimate RP | 0.053 | 0.156 | 0.342 | 0.986 |
| Animate Gap – Animate RP | –0.717 | 0.152 | –4.716 | <.0001 |
| Inanimate RP – Animate RP | –0.980 | 0.154 | –6.346 | <.0001 |
For completeness, estimated marginal means illustrating this interaction are provided in Appendix B.
Random effects revealed that most variance was attributable to participants (τ₀₀ = 2.79), while lexical items contributed relatively little (τ₀₀ = 0.11). The intra-class correlation (ICC = 0.47) indicates that nearly half the variance stemmed from participant-level clustering. Thus, while the animacy effect was broadly consistent across items, ratings varied considerably across individuals.
Figure 3 illustrates this variability. It plots raw ratings across conditions, highlighting broad dispersion but also a tendency toward higher ratings in the Animate–RP condition. This underscores that the interaction effect is not driven by outliers but reflects a robust group-level trend.
To further illustrate the distribution of acceptability judgments across participants, histograms of participant-level mean ratings by condition are provided in Figure C1 in Appendix C.
The full statistical output of the CLMM is reported in Appendix D. We further tested a model with the maximally converging random effects structure (Animacy * Tail | Participant; Animacy * Tail | Lexicalisation), since models using only random intercepts are known to elicit increased Type 1 error rates (Barr et al. 2013). This model converged and yielded a fixed-effect pattern that was qualitatively consistent with the simpler random-intercepts model. In particular, the Animacy × Tail interaction remained significant and in the same direction in both models (β = 0.77, p < .001 in the intercept-only model; β = 0.98, p < .001 in the maximal model). Full results of the maximal model are reported in Appendix E.
For the filler sentences, participants rated the grammatical items consistently at the top of the scale, with most responses clustering at 7. By contrast, the ungrammatical fillers, regardless of the type of violation illustrated in (3)–(5), were rated toward the lower end, with the majority of responses at 1–2 and progressively fewer at higher values. The distributional plots (Figure 4) show that participants used the full 1–7 scale and that fillers produced the expected contrast between grammatical and ungrammatical sentences.
2.4.2 Self-paced reading (SPR) analysis
We evaluated reading times at the spillover region immediately following the RP or gap, which consisted of a three-word adjunct phrase. We targeted the spillover region rather than the integration site itself because verbs hosting resumptive pronouns involve cliticization, which independently inflates reading times relative to gap conditions and could therefore obscure effects of interest. An example stimulus is shown below, with the spillover region in bold:
- (9)
- meen / ʔayš
- who / what
- tafaʔʒaʔt-i
- got.surprised-you
- lamma
- when
- al-mudīra
- the-principal
- intaqadat __/-uh
- criticized __/(him/it)
- fߓī
- in
- ijtimāʕ
- meeting
- al-idāra?
- the-administration
- (‘Who/what did you get surprised about when the principal criticized __ / him in the administration meeting?’)
Table 4 illustrates how region labels (N1–N4, RT0–RT3) correspond to individual words in the sentence.
Table 4: Illustration of region labels (N1–N4, RT0–RT3) used in the self-paced reading analysis, mapped onto the words of an example experimental sentence.
| Region | Word (example sentence) | Word (English translation) |
| N1 | meen / ʔayš | Who / What |
| N2 | tafaʔʒaʔt-i | got.surprised-you |
| N3 | lamma | when |
| N4 | al-mudīra | the-principal |
| RT0 (integration site) | Intaqadat__/-uh | criticized-/(him/it) |
| RT1 (spillover region) | fī | in |
| RT2 (spillover region) | ijtimāʕ | meeting |
| RT3 (spillover region) | al-idāra | the-administration |
Reaction times (RTs) were cleaned at the region level by removing values below 100 ms and above 3000 ms for all regions. Extremely fast responses were taken to reflect accidental key presses, while very slow ones likely signalled distraction. After trimming, reaction time at the spillover region was computed as the sum of the RTs of the three words following the integration site, so that this measure was based on the same cleaned trials. This procedure kept about 78% of the data, which is typical for self-paced reading studies, and left us with a dataset that better reflects natural sentence processing.
Figure 5 shows the region-by-region reading times across conditions. Reading times in the pre-critical regions (N1–N4) were comparable across conditions, indicating that participants were engaged with the task and that differences did not emerge prior to the integration site. At the integration region (RT0), which corresponds to the subcategorizing verb, RTs were elevated in RP-animate condition, but it could be due to the cliticization process. In the following spillover regions (RT1–RT3), reading times were systematically longer in the animate conditions compared to the inanimate conditions, while no consistent differences were observed between gap and RP conditions.
Figure 5: Region-by-region reading times across conditions (Left panel: gap conditions; right panel: RP conditions). Mean RTs (in seconds) are plotted for pre-critical regions (N1–N4), the integration region (RT0), and spillover regions (RT1–RT3). Lines indicate condition means (Animate–Gap, Animate–RP, Inanimate–Gap, Inanimate–RP), with 95% confidence intervals shown as vertical bars. Reading times were comparable across conditions in the pre-critical regions, elevated at the integration region (RT0) in Animate-RP condition, and consistently longer for animate fillers in the spillover regions.7
Figure 6 presents the mean raw reading times for the three spillover words (RT_Sum) across conditions (the sum of RT1, RT2, RT3). The descriptive pattern suggests longer reading times in the Animate conditions relative to the Inanimate ones.
To correct for skewness in the RT_Sum distribution, a Box–Cox transformation was applied. The estimated λ was –0.18, close to zero, indicating that a log transformation provides a suitable approximation (Box & Cox 1964). Distributional diagnostics of the transformed RTs are provided in Appendix F.
Transformed reaction times at the spillover region (i.e. RT_Sum) were analyzed using Generalized Additive Models (GAMs) (Wood 2011) implemented in the mgcv package in RStudio (Version 1.1.419). While Box–Cox transformation reduces skewness in reaction times, it does not address other non-linear properties of reading data, such as practice and fatigue effects, spillover between regions, and individual variability in temporal dynamics. GAMs allow flexible smooths for these predictors and random smooths for participants, making them well suited for psycholinguistic time-course data (Baayen et al. 2016).
Likelihood ratio comparisons between nested GAMs showed that the best-fitting model included Animacy as a fixed effect, with random smooths for Participant and Lexicalisation. The smooth for Participant was highly significant (edf ≈ 70, F = 8.86, p < .001), indicating substantial individual variability, whereas the smooth for Lexicalisation was not (edf ≈ 2, F = 0.40, p = .20). Adding Tail or the Animacy × Tail interaction did not improve model fit. Thus, Animacy emerged as the sole significant predictor of reading times.
Specifically, sentences with inanimate fillers were read significantly faster than those with animate fillers (β = –0.020, SE = 0.005, t = –4.05, p < .001). The model explained 45% of the deviance, with an adjusted R2 of 0.409 (REML = –1066.4). This effect is visualized in Figure 7, which shows that median RTs were consistently lower for inanimate fillers.
Figure 7: Distribution of Box–Cox–transformed reading times (RTs) by animacy condition. Each violin plot shows the full distribution of RTs, with an embedded boxplot marking the interquartile range and median. Median RTs are consistently lower for inanimate fillers, reflecting the significant animacy effect identified in the GAM analysis.
Model diagnostics supported the adequacy of the fit. Fitted versus observed values showed a strong linear trend with no signs of heteroscedasticity, and residuals were symmetrically distributed, further confirming the reliability of the model (see Appendix G).
To ensure the robustness of these results, we also ran supplementary linear mixed-effects models using the lme4 package (Bates et al. 2015). The results of these models, reported in Appendix H, converge with the GAM analysis in showing a reliable animacy effect and no evidence for a tail effect or interaction.
To assess the reliability of the self-paced reading data under the acceptability-judgment task, we conducted an independent filler-based analysis. This analysis shows that sentence-level reading times reliably distinguish grammatical from ungrammatical filler items (see Appendix I).
Furthermore, we examined whether variation in offline acceptability judgments across participants was systematically related to variation in online processing difficulty. To this end, we conducted a participant-level correlation analysis between each participant’s mean acceptability ratings and their mean spillover-region reading times. This analysis revealed no reliable association between acceptability and reading times (Pearson’s r = .19, p = .097). The same conclusion was obtained using a rank-based correlation (Spearman’s ρ = .14, p = .218). These results suggest that acceptability judgments do not systematically track online processing difficulty at the participant level.
3 Discussion
This study investigated whether animacy affects the acceptability and processing of island-violating dependencies in Saudi Arabic, and whether this effect is modulated by the type of dependency tail (gap vs. resumptive pronoun). The experiment crossed filler animacy (animate vs. inanimate) with tail type (gap vs. RP) in adjunct island contexts. The results reveal a misalignment between acceptability judgment data and online processing data. While animacy slightly but significantly improved acceptability ratings in the RP condition, there was no such benefit in the gap condition. Self-paced reading data, however, revealed longer reading times for animate fillers across both tail types. These findings pose a challenge to both a purely syntactic account as well as a purely memory-based processing account; instead, it motivates a multivariate account that views island effects as emerging from the interplay of syntax, processing, and interpretation.
3.1 Syntactic accounts
The findings are not predicted by categorical syntactic accounts of island effects. Within theoretical syntax, island-violating wh-dependencies are assumed to be ungrammatical due to violations of movement constraints (Chomsky 1973; 1986; Huang 1982). In the context of Arabic syntax, Aoun et al. (2009) argue that while gap dependencies across islands violate movement constraints, RPs may rescue grammaticality through establishing binding relations, but only with wh-expressions that project a DP-layer (see Section 2), as with animate who but not inanimate what.
Our results diverge from the predictions of this account. The analysis in Aoun et al. (2009) predicts a sharp contrast between licit and illicit dependencies: animate–RP conditions should be judged acceptable, whereas all gap conditions and inanimate–RP conditions should be unacceptable. Instead, we observe a gradient pattern of acceptability. Animate–RP sentences receive higher ratings than the other conditions, but their acceptability remains substantially below full acceptability. The modest, graded improvement observed in the acceptability of animate–RP sentences challenges the syntactic claim that resumptive pronouns in Arabic license island-violating who-questions through binding. If such a binding relation were fully responsible for licensing these dependencies, we might expect a stronger or more categorical improvement in acceptability than what is observed, rather than the relatively subtle shift in mean ratings. This subtle, gradient nature of the animacy–resumption interaction suggests that their ameliorating effect operates beyond the domain of categorical syntactic licensing.
However, this does not imply that the results are incompatible with syntactic explanations per se, but rather that they call for syntactic models that allow graded well-formedness. As noted in the Background (see Note 1), although island constraints are traditionally treated as categorical in generative syntax, an increasing body of work has argued that grammatical well-formedness itself may be gradient, even within the competence grammar (Haegeman & Jiménez-Fernández & Radford 2014; Villata & Sprouse & Tabor 2019; Villata & Tabor 2022).
Within this perspective, the present findings can be naturally interpreted under the Self-Organized Sentence Processing (SOSP) model (Villata et al. 2019; Villata & Tabor 2022). SOSP assumes a flexible structure-formation system in which the parser attempts to assemble an interpretable representation even under suboptimal conditions by coercing elements into roles that are not ideally licensed. When no sufficiently compatible grammatical structure is available, coercion fails and acceptability remains low. When a partially compatible structure exists, however, the system may converge on a strained but interpretable parse, yielding intermediate acceptability rather than absolute ungrammaticality. From this perspective, the intermediate ratings observed in the Animate–RP condition reflect a reduction in structural strain: although adjunct island constraints continue to block full dependency formation, the presence of a resumptive pronoun—particularly with an animate filler—provides a configuration that is similar enough to independently licensed resumptive structures in Arabic, such as relative clauses (Aoun et al. 2009), to support partial convergence.
A remaining question is why coercion succeeds in the Animate–RP condition but fails in the Inanimate–RP condition. Under the SOSP account, this asymmetry can be explained by considering the constraints governing the grammatical distribution of resumptive constructions in Arabic. Independently licensed resumptive dependencies in Arabic, most prominently in relative clauses, are systematically associated with referential, highly individuated fillers (Aoun et al. 2009). As a result, animate fillers paired with resumptive pronouns inside adjunct islands provide a closer match to an existing grammatical template, reducing structural strain and allowing partial convergence. In contrast, in the Inanimate–RP condition, no sufficiently similar grammatical template is available. Although a resumptive pronoun is present, inanimate fillers are poorly matched to the types of resumptive dependencies independently licensed by the grammar, preventing successful coercion. Consequently, structural strain remains high and acceptability remains low.
While the acceptability pattern in the Animate–RP condition is naturally captured by SOSP, the online processing data are not fully accounted for by this analysis. Under SOSP, coercion is expected to arise when an alternative, independently licensed structure is sufficiently compatible with the input. In such cases, increased reading times are interpreted as reflecting the additional processing effort involved in coercing the input into a partially compatible representation. However, increased reading times are also observed for animate fillers in gap conditions, where no such compatible structure is available and coercion is therefore not expected to occur.
3.2 Memory-based processing accounts
Memory-based accounts do not straightforwardly predict the observed pattern. These approaches attribute the unacceptability of island-violating sentences to working-memory limitations involved in maintaining and retrieving fillers during dependency formation (Kluender 1991; Kluender & Kutas 1993a; 1993b; Hofmeister & Sag 2010); on this view, animacy should increase the accessibility of the filler and thereby ease processing, since accessible fillers are expected to be easier to maintain and retrieve during dependency formation. Resumptive pronouns should further reduce processing demands by providing an overt integration site for the filler (McDaniel & Cowart 1999; Chacón 2019; Hammerly 2021).
Our results diverge from these predictions. If accessibility of filler phrases in memory were the driving factor, animate wh-fillers should speed processing at the integration site, regardless of tail type, yielding shorter RTs. Contrary to the predictions of memory-based accounts, our results reveal the opposite trend: animacy consistently slowed processing in both gap and RP conditions. This reversal is not just about the size of the effect—it’s about its direction. Rather than facilitating processing, animacy imposed a cost. Such a pattern is not predicted by the core assumption that animacy enhances retrieval by increasing accessibility. In fact, our findings suggest a misalignment between offline and online measures: animacy improves acceptability in the RP condition but slows processing across the board. This asymmetry is not straightforwardly predicted by memory-based accounts, which assume that higher accessibility should make retrieval easier, not harder.
Taken together, the self-paced reading results do not provide evidence for the facilitation predicted by memory-based accessibility accounts, as animacy was associated with longer reading times across conditions. Nevertheless, this pattern should be interpreted with caution, since task demands (the combination of self-paced reading and acceptability judgment) and the relatively high proportion of unacceptable sentences may have influenced readers’ processing strategies and made accessibility-based facilitation effects harder to detect in the self-paced reading data (see the Limitations section below for discussion).
An alternative interpretation of the self-paced reading results comes from similarity-based interference accounts of dependency formation, which predict increased processing cost when a dependency must be maintained or retrieved in the presence of featurally similar interveners (Jäger & Engelmann & Vasishth 2017). In the present materials, animate wh-fillers (e.g., meen ‘who’) introduce a discourse-prominent animate representation that must be maintained in the presence of other animate referents (e.g., second-person pronouns and lexical NPs in the adjunct clause such as the principal in 10), increasing similarity along features such as [+animate] and [+human].
- (10)
- meen
- who
- tafaʔʒaʔt-i
- got.surprised-you
- lamma
- when
- al-mudīra
- the-principal
- intaqadat-hu
- criticized-him
- fī
- in
- ijtimāʕ
- meeting
- al-idāra?
- the-administration
- (‘Who did you get surprised about when the principal criticized him in the administration meeting?’)
From this perspective, the observed slowdown for animate fillers across both gap and resumptive-pronoun conditions follows naturally: increased reading times reflect interference arising from maintaining a salient animate representation amid competing animate discourse entities, rather than difficulty associated with dependency resolution per se. Importantly, this account also predicts the absence of a processing advantage for resumptive pronouns. In Arabic, bare wh-fillers are not morphologically specified for gender or number, whereas resumptive pronouns are marked for these features but do not encode animacy. Consequently, resumptive pronouns do not provide retrieval cues that distinguish animate from inanimate fillers under similarity-based accounts and are therefore not expected to facilitate online processing selectively for either type. If resumptive pronouns exert any effect under similarity-based interference, it is more likely to increase interference when their morphosyntactic features overlap with those of intervening constituents, rather than to reduce it.8
At the same time, this explanation leaves the acceptability pattern unaccounted for. While similarity-based interference predicts increased processing cost for animate fillers regardless of tail type, it does not explain why acceptability improves selectively in the Animate–RP condition. The resulting dissociation between online processing difficulty and offline acceptability suggests that additional interpretive or discourse-level factors contribute to RP-specific amelioration.
Because the present study did not independently manipulate the animacy, gender, or number features of intervening constituents, feature-based interference effects cannot be isolated directly. Accordingly, while interference-based accounts can plausibly explain the animacy-related processing cost observed here, they do not by themselves account for the selective offline acceptability improvement observed only in the RP condition. A more direct test of the interference-based explanation would require future experiments that independently manipulate intervener similarity with the wh-filler (e.g., animacy) and with the resumptive pronoun (e.g., gender or number).
3.3 A multivariate account
To account for these findings, we propose a multivariate account.9 Animate fillers (like who) are naturally more prominent because they refer to specific, individuated entities that are easy to anchor and track in the discourse, unlike inanimate fillers like what, which tend to denote kinds or categories (Alexopoulou & Keller 2013). When a resumptive pronoun appears, it can link to this salient filler and form an interpretable anaphoric dependency (Erteschik-Shir 1992; Frazier & Clifton 2002). That is, the resumptive pronoun functions as an anaphoric pronoun that is linked back to a highly salient referent, resulting in a more interpretable sentence. This explains the higher acceptability ratings observed in the animate–RP condition.
At the same time, the process is not cost-free. Constructing a discourse-based dependency requires additional effort, as the parser shifts away from syntactic dependency resolution toward an anaphoric, discourse-based interpretation (Alexopoulou 2010). This additional interpretive work is reflected in the slower reading times observed for animate fillers across both gap and RP conditions. In other words, animacy improves interpretability but increases processing demands, a trade-off that captures the mismatch between acceptability judgments and reaction times in our data.
Specifically, the parser treats who as a strongly individuated discourse entity that can naturally serve as the antecedent of an anaphoric pronoun. Even though grammar blocks gap-search inside islands (Stowe 1986; Frazier & d’Arcais 1989; Traxler & Pickering 1996), parsers may still anticipate a pronoun as a possible non-grammatical resolution strategy (Keshev & Meltzer-Asscher 2017). When no resumptive pronoun appears (gap condition), the parser is left with an unfulfilled prediction, forcing it to sustain unresolved search pressure, which slows processing and contributes to degraded acceptability. In what-questions, by contrast, no such pronoun search is triggered, since inanimates are less natural antecedents for anaphoric pronouns in anaphoric dependencies. Thus, while both animate and inanimate gaps remain unresolved, only who-gap involves an unmet anaphoric expectation, which explains the longer reading times.
The lack of processing differences between the animate–gap and animate–RP conditions can be understood in this light. In the animate–RP condition, the parser appears to consume memory resources in search of an antecedent, which increases processing load—just as it does with gaps. But unlike gaps, resumptive pronouns offer a partial discourse-level repair: they enable the parser to construct an anaphoric link back to the salient animate filler. This seems to support a modest rise in acceptability, even if the extra processing demands remain. Put differently, both tails in the animate condition increase processing costs—gaps because they leave anaphoric predictions unmet, and resumptives because they require additional linking operations. Yet only the latter offers a path to a more coherent interpretation.
This explanation aligns well with the eclectic model of island effects proposed by Chaves & Putnam (2020), which treats syntax, processing, and discourse as jointly shaping both acceptability and real-time comprehension. A purely syntactic account struggles to explain why animacy modestly improves RP sentences, and a memory-based account falls short of explaining why that improvement comes with a processing penalty.
Our findings suggest that animacy here acts as a discourse cue rather than a processing facilitator: it enhances interpretability without easing integration effort. The lack of a processing benefit for RPs further underscores their role as interpretive, not syntactic or processing devices in Saudi Arabic. They function like discourse pronouns that improve the interpretability of island-violating dependencies, but only when the filler is semantically rich enough to support such an interpretation. This perspective aligns with an anaphoric view of resumption, in which RPs behave more like referential pronouns than structural devices (Erteschik-Shir 1992; Frazier & Clifton 2002).
More broadly, the multivariate account developed here fits naturally within recent approaches to locality phenomena that treat acceptability and processing as the outcome of interacting syntactic, discourse, and processing constraints. In a large-scale experimental investigation of bridge effects in English, (Huang et al. 2025) show that bridge effects emerge from the interaction of morphosyntactic licensing, information-structural constraints, and processing costs, with no single factor proving sufficient; their multivariate model substantially outperforms single-predictor accounts. The present findings extend this conclusion to adjunct island violations in Arabic. While animacy and resumption interact to improve interpretability at the level of judgment, they do not alleviate—and may even increase—online processing cost. This dissociation supports a view in which discourse-level factors modulate interpretability without eliminating structural constraints or processing difficulty, reinforcing the need for a multivariate account of island effects.
3.4 Implications
3.4.1 Resumptive pronouns and grammatical licensing across constructions
The present findings bear on the broader debate concerning the grammatical status of resumptive pronouns across languages and constructions (see Meltzer-Asscher 2021 for review). While resumptive pronouns are often argued to be fully grammatical in languages such as Hebrew and Arabic, this claim is well known to be construction-specific (Choueiri 2017). In Arabic experimental literature, resumptive pronouns are strongly preferred in relative clause constructions compared to wh-questions. For instance, in Jordanian Arabic, resumptives are rated substantially higher in relative clauses than in wh-dependencies (Al-Aqarbeh & Sprouse 2023). Similarly, in Baha Arabic, resumptive dependencies in ʔilli-containing relative clauses and ʔilli-containing cleft wh-questions are highly acceptable, with a likelihood of high acceptance ratings (5–7) of approximately 75% (Hayyas & De Cat 2023). Examples of ʔilli-containing constructions from these materials are shown in (11).
- (11)
- Licensed resumptive pronouns in ʔilli-constructions
- A.
- Non-island configuration
- šeft
- saw.1SG
- as-sāʕa
- the-watch
- ʔilli
- that
- qulti
- said.2SG.F
- lī
- to-me
- ʔinn
- that
- Ṣāliḥ
- Saleh
- ʔištarā-hā
- bought.3SG.M-it
- ‘I saw the watch that you told me that Saleh bought.’
- B.
- Adjunct-island configuration
- kabbai-t
- spill.PST-1SG
- al-ḥaleeb
- the-milk
- ʔilli
- that
- khaled
- Khaled
- meriḍ
- got.sick.3SG.M
- baʕd-ma
- after
- šerib-uh
- drank.3SG.M-it
- (‘I spilled the milk that Khaled became sick after he drank.’)
- (Baha Arabic, Hayyas & De Cat 2023)
The examples in (11) show that resumptive pronouns in ʔilli-constructions remain fully acceptable even when the dependency crosses an adjunct-island configuration, supporting their analysis as structurally licensed rather than repair-like resumptives. By contrast, the resumptive pronouns examined in the present study—restricted to bare wh-questions crossing adjunct island boundaries—receive markedly lower acceptability ratings. Even in the Animate–RP condition, mean judgments remain well below commonly assumed acceptability thresholds, with the likelihood of high acceptance ratings (5–7) ranging only between 15–18% (see Figure 2). Although this comparison is necessarily cross-experimental, the contrast is striking and points to a principled distinction between two types of resumptive pronouns: structurally licensed resumptives, which are base-generated and grammatically integrated into the syntax (e.g., in ʔilli-relatives and clefts), and last-resort or repair-like resumptives, which arise when movement dependencies fail and interpretation is recovered via binding or discourse-level mechanisms.
This distinction is directly relevant to ongoing debates on resumptive pronouns in English. A large body of earlier work argues that English resumptives are not syntactically licensed but instead function as processing-related repair strategies, particularly in complex dependencies and island environments (e.g., Kroch 1981; McDaniel & Cowart 1999; Alexopoulou 2010; Hofmeister & Norcliffe 2013; Chacón 2019; Hammerly 2021). By contrast, Ackerman & Frazier & Yoshida (2018) argue on the basis of forced-choice judgment tasks that English resumptives may be grammatical. Because forced-choice paradigms can obscure absolute acceptability contrasts, the apparent acceptability of English resumptives in such tasks may reflect relative preference under forced alternatives rather than full grammatical licensing. The present findings are more consistent with the repair-based view.
Taken together, the low to intermediate acceptability of Animate–RP sentences in the present study does not undermine the existence of grammatically licensed resumptive pronouns in Arabic more generally. Rather, it reinforces the need to distinguish carefully between resumptives that are licensed by the grammar in specific constructions and those that function as repair strategies when syntactic movement is blocked, yielding amelioration without full grammaticality.
3.4.2 Methodological implications for experimental studies of islands
These findings raise important considerations for experimental research. Animacy, often treated as a secondary lexical feature, turns out to play a central role in how island-violating dependencies are judged—especially when paired with resumptive pronouns. This challenges the common assumption that fillers like who and what are interchangeable. In reality, animate fillers tend to be more prominent in discourse: they refer to specific, identifiable individuals who are easier to track across a sentence. That added salience can influence how island violations are perceived and processed.
This has broader methodological implications. Much of the existing experimental literature—particularly work using superadditivity paradigms (e.g., Sprouse 2007; Al-Aqarbeh & Sprouse 2023)—treats who and what as functionally equivalent “bare wh-phrases.” But if animacy affects acceptability, then part of what appears to be a structural sensitivity might actually reflect properties of the filler itself. This makes it crucial for researchers to control or manipulate animacy explicitly; otherwise, lexical and discourse-level effects may be mistaken for grammatical constraints. Future work will need to determine whether animacy’s influence is additive, interactive, or orthogonal to structural factors—but it is no longer something that can be safely ignored.
This study also underscores the value of combining offline and online measures. While animacy boosted acceptability in RP contexts, it did not make those dependencies easier to process. That divergence would have gone unnoticed if we had relied on only one type of data. Together, the two measures provide a fuller picture—one that helps disentangle the contributions of syntax, discourse, and processing in shaping how island violations are understood.
3.5 Limitations
A growing body of work shows that reading behavior in self-paced reading tasks is sensitive to task demands, particularly to whether readers anticipate providing acceptability judgments or answering comprehension questions. Such task manipulations have been shown to influence processing strategies and grammaticality bias, thereby shaping online reading-time profiles (Hammerly & Staub & Dillon 2019; Laurinavichyute et al. 2023; Laurinavichyute & von der Malsburg 2024). In particular, Laurinavichyute & von der Malsburg (2024) argue that acceptability-based tasks reduce default assumptions of well-formedness and are associated with more conservative, analytic processing, in which readers remain sensitive to structural uncertainty during incremental parsing rather than relying on shallow heuristics that often suffice in comprehension-based paradigms.
Evidence for the impact of task configuration also comes from prior work using closely comparable materials in Baha Arabic. In that work, self-paced reading paired with comprehension questions yielded no significant effects of resumption, D-linking, or islandhood, whereas pairing the same paradigm with acceptability judgments revealed systematic and theoretically interpretable reading-time effects (Hayyas 2020). Together, these findings suggest that task demands can systematically modulate online processing strategies in island-violating dependencies.
A related limitation concerns the distribution of sentence types in the experimental materials. Participants were exposed to a relatively high proportion of unacceptable configurations, due both to the inclusion of island-violating test sentences and to the number of ungrammatical fillers. This distribution differs from that typically adopted in self-paced reading studies aimed at approximating naturalistic comprehension, where ungrammatical sentences are often kept below approximately 15–20%. Prior work shows that increasing the proportion of unacceptable material can alter participants’ expectations about sentence well-formedness and modulate online reading behavior (Hammerly et al. 2019; Laurinavichyute et al. 2023; Laurinavichyute & von der Malsburg 2024). Such input statistics may lead readers to adopt more cautious parsing strategies, yielding reading-time profiles that differ from those observed under more naturalistic distributions.
In light of these task-related and distributional considerations, the dissociation observed between online reading times and offline acceptability judgments in the experimental items suggests that reading times in the present study should not be treated as direct proxies for perceived sentence quality. Rather, they appear to index the processing effort involved in maintaining and interpreting island-violating wh-dependencies during incremental parsing. Importantly, this dissociation is not observed for filler items: reading times reliably distinguish grammatical from ungrammatical fillers and closely track acceptability judgments. The divergence in the experimental conditions therefore appears to reflect properties specific to island-violating dependencies, rather than a general limitation of the self-paced reading method under evaluative task demands.
Crucially, these task-related factors do not undermine the interpretation of the animacy-related slowdown reported here but instead help clarify how the task may have shaped the processing strategies through which this effect became visible. Because acceptability-based reading encourages readers to remain attentive to unresolved dependencies during incremental parsing, the longer reading times observed for animate fillers remain consistent with the interpretation that animacy increased the effort associated with maintaining and resolving an anaphoric dependency rather than facilitating retrieval.
Given that task demands and input distributions shape online processing strategies, future work employing alternative methods—such as eye-tracking, maze tasks, or forced-choice paradigms—would be valuable for assessing the robustness of animacy and resumption effects under conditions that more closely approximate incremental comprehension.
4 Conclusion
Our findings show that animate fillers improve the acceptability of island-violating dependencies in Saudi Arabic, but only when paired with a resumptive pronoun. This effect, while statistically reliable, is modest, highlighting a non-structural source. In fact, the improvement appears to come at a processing cost, which suggests that the underlying mechanism is interpretive rather than facilitatory. More specifically, these results indicate that animacy and resumption interact at the discourse level. Salient fillers may help the parser construct an anaphoric link to the resumptive pronoun, making the sentence more interpretable in offline acceptability tasks, even if it remains ungrammatical. Building this anaphoric dependency is costly in real time, as reflected in elevated reading times.
Taken together, the findings argue against single-factor explanations of island phenomena. A purely syntactic account cannot explain the modest acceptability boost, and a purely memory-based account cannot explain the persistent processing cost. Instead, the data point to a multivariate account in which syntax, processing, and discourse each impose their own constraints. Syntax rules out the movement, discourse linking can partially repair interpretation, and processing mechanisms bear the cost of that repair. These results underline the need for models of filler-gap dependencies that distinguish between interpretability benefits and integration costs, and that take filler properties, like animacy, seriously in both experimental design and theoretical accounts.
Abbreviations
2SG = second person singular; 3SG = third person singular; F = feminine; M = masculine; RP = resumptive pronoun.
Data availability
All data and analysis scripts for this study are available at the following OSF repository: https://doi.org/10.17605/OSF.IO/SGWZ5
Additional file
The additional file for this article can be found as follows:
Appendices. Appendices A to I. DOI: https://doi.org/10.16995/glossa.25777.s1
Ethics and consent
This study was conducted in accordance with institutional and academic ethical standards. Ethical approval was not required as per the policies of Al-Baha University, which does not mandate ethical review for linguistic studies that do not involve sensitive personal data or vulnerable populations.
Acknowledgements
I am grateful to Prof. Sol Lago and the anonymous reviewers for their careful reading and insightful comments, which substantially improved the clarity and coherence of this manuscript. I am especially grateful to Prof. Cécile De Cat for her guidance and support during my doctoral research, which helped shape the ideas developed in this paper. Any remaining errors are my own.
Competing interests
The author has no competing interests to declare.
Notes
- The syntactic traditional view that island constraints are categorical remains influential in generative syntax (e.g., Bošković 2018; Shafiei & Graf 2020). At the same time, a growing body of work has argued that island effects may exhibit gradient grammaticality within the competence grammar itself, rather than reflecting a strict grammatical–ungrammatical dichotomy. Such approaches include proposals based on weighted or soft syntactic constraints (e.g., Haegeman & Jiménez-Fernández & Radford 2014), as well as probabilistic models of structure formation such as the Self-Organized Sentence Processing (SOSP) framework (e.g., Villata & Sprouse & Tabor 2019; Villata & Tabor 2022). These accounts maintain a central role for syntactic configuration while allowing acceptability to vary as a function of structural compatibility and representational strain. The present study does not adjudicate between categorical and probabilistic grammatical frameworks; however, the relevance of gradient grammaticality approaches for interpreting the observed amelioration effects is discussed further in Section 3.1. [^]
- Although processing-based explanations of island effects often appeal to limitations in working memory, several experimental studies have found little or no relationship between individual working memory capacity and sensitivity to island violations (Sprouse & Wagers & Phillips 2012; Pañeda et al. 2020; Pham et al. 2020; Aldosari & Covey & Gabriele 2024). [^]
- Since this evidence comes from dependency-related diagnostics rather than independent morphosyntactic tests outside wh-dependencies, we adopt this analysis here as a working structural assumption and test whether the animate–inanimate contrast patterns in a way consistent with its predictions. [^]
- It is worth noting that the grammatical status of resumptive pronouns varies considerably across languages. In English, for instance, their acceptability has long been debated, with many analyses treating them as processing-related repair strategies rather than syntactically licensed elements (e.g., Kroch 1981; McDaniel & Cowart 1999; Alexopoulou 2010; Hofmeister & Norcliffe 2013; Chacón 2019; Hammerly 2021). However, some work, using specific experimental paradigms, argues that resumptive pronouns may be grammatically licensed (Ackerman & Frazier & Yoshida 2018). We return to this issue in the Discussion in connection with the interpretation of the present results. [^]
- Importantly, Aoun et al.’s (2009) argument is based on evidence from multiple varieties of Arabic (Lebanese Arabic, Moroccan Arabic, and Standard Arabic), suggesting that the relevant distinction reflects stable properties of Arabic wh-dependencies rather than dialect-specific idiosyncrasies. Although Saudi Arabic is not directly addressed in Aoun et al. (2009), the present study adopts their feature-based analysis as a working assumption, while remaining open to the possibility of micro-variation across dialects. [^]
- Including both grammatical and ungrammatical filler sentences is standard practice in acceptability judgment tasks, as it encourages participants to use the full rating scale. However, in the context of self-paced reading, a relatively high proportion of unacceptable sentences may influence participants’ expectations about well-formedness and modulate online reading behavior (Hammerly & Staub & Dillon 2019; Laurinavichyute et al. 2023; Laurinavichyute & von der Malsburg 2024). The potential consequences of this design choice for interpreting the reading-time data are discussed further in the Limitations section. [^]
- Notably, the region-by-region plots indicate that divergence between Animate–RP and Inanimate–RP may begin at the integration site itself, despite identical clitic morphology across these conditions. Although this region was not included in the statistical analysis, the descriptive pattern is consistent with our interpretation that animate fillers incur greater processing cost to resolve the dependency. [^]
- This interpretation is consistent with recent evidence that resumptive pronouns do not necessarily eliminate similarity-based interference. In particular, Saul & Keshev & Meltzer-Asscher (2025) show that fillers in Hebrew relative clauses remain susceptible to interference even in the presence of gender-marked resumptive pronouns. Although the present study does not directly manipulate intervener similarity, the observed pattern is compatible with the view that resumptive pronouns are not inherently interference-reducing devices. [^]
- Discourse-oriented approaches to island effects differ in important ways. Classical discourse-based accounts treat island sensitivity as a consequence of information-structural constraints encoded in the grammar itself, such as backgroundedness, focus structure, or discourse accessibility (e.g., Erteschik-Shir 1973; Ambridge & Goldberg 2008; Goldberg 2013; Abeillé et al. 2020). On such views, island violations arise because extraction targets elements that are discourse-inappropriate or insufficiently prominent, rather than because of structural or processing constraints per se. By contrast, multivariate approaches such as Chaves & Putnam (2020) and Huang & Almeida & Sprouse (2025) do not reduce island effects to discourse factors alone. Instead, they treat acceptability as emerging from the interaction of multiple constraints—including syntactic configuration, processing complexity, semantic compatibility, and discourse prominence—none of which is individually sufficient or necessary. The present analysis adopts this multivariate perspective: while discourse-related properties such as animacy and referential prominence play a role in modulating acceptability, they do so only in interaction with syntactic constraints and processing mechanisms, rather than as independent grammatical licensing conditions. [^]
References
Abeillé, Anne & Hemforth, Barbara & Winckel, Elodie & Gibson, Edward. 2020. Extraction from subjects: Differences in acceptability depend on the discourse function of the construction. Cognition 204. 104293. DOI: http://doi.org/10.1016/j.cognition.2020.104293
Ackerman, Lauren & Frazier, Michael & Yoshida, Masaya. 2018. Resumptive pronouns can ameliorate illicit island extractions. Linguistic Inquiry 49(4). 847–859. DOI: http://doi.org/10.1162/ling_a_00291
Al-Aqarbeh, Rania & Sprouse, Jon. 2023. Island effects and amelioration by resumption in Jordanian Arabic: An auditory acceptability-judgment study. Syntax 27. DOI: http://doi.org/10.1111/synt.12262
Aldosari, Saad & Covey, Lauren & Gabriele, Alison. 2024. Examining the source of island effects in native speakers and second language learners of English. Second Language Research 40. DOI: http://doi.org/10.1177/02676583221099243
Alexopoulou, Theodora. 2010. Truly intrusive: resumptive pronominals in questions and relative clauses. Lingua 120(3). 485–505. DOI: http://doi.org/10.1016/j.lingua.2008.10.009
Alexopoulou, Theodora & Keller, Frank. 2013. What vs. who and which: kind-denoting fillers and the complexity of whether-islands. In Sprouse, Jon & Hornstein, Norbert (eds.), Experimental Syntax and Island Effects, 310–340. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139035309.016
Almeida, Diogo. 2014. Subliminal wh-islands in Brazilian Portuguese and the consequences for syntactic theory. Revista Da ABRALIN 13. 55–91. DOI: http://doi.org/10.5380/rabl.v13i2.39611
Ambridge, Ben & Goldberg, Adele. 2008. The island status of clausal complements: Evidence in favor of an information structure explanation. Cognitive Linguistics 19(3). 357–389. DOI: http://doi.org/10.1515/COGL.2008.014
Aoun, Joseph & Benmamoun, Elabbas & Choueiri, Lina. 2009. The Syntax of Arabic. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9780511691775
Atkinson, Emily & Apple, Aaron & Rawlins, Kyle & Omaki, Akira. 2016. Similarity of wh-phrases and acceptability variation in wh-islands. Frontiers in Psychology 6. 2048. DOI: http://doi.org/10.3389/fpsyg.2015.02048
Baayen, Harald & Vasishth, Shravan & Kliegl, Reinhold & Bates, Douglas. 2016. The cave of shadows: Addressing the human factor with generalized additive mixed models. Journal of Memory and Language 94. 206–234. DOI: http://doi.org/10.1016/j.jml.2016.11.006
Barr, Dale J. & Levy, Roger & Scheepers, Christoph & Tily, Harry J. 2013. Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language 68(3). 255–278. DOI: http://doi.org/10.1016/j.jml.2012.11.001
Bates, Douglas & Mächler, Martin & Bolker, Ben & Walker, Steve. 2015. Fitting linear mixed-effects models using lme4. Journal of Statistical Software 67(1). 1–48. DOI: http://doi.org/10.18637/jss.v067.i01
Boeckx, Cedric. 2012. Syntactic Islands. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139022415
Bošković, Željko. 2018. On movement out of moved elements, labels, and phases. Linguistic Inquiry 49(2). 247–282. DOI: http://doi.org/10.1162/LING_a_00273
Box, G. E. P. & Cox, D. R. 1964. An analysis of transformations. Journal of the Royal Statistical Society: Series B (Methodological) 26(2). 211–243. DOI: http://doi.org/10.1111/j.2517-6161.1964.tb00553.x
Chacón, Dustin. 2019. Minding the gap?: Mechanisms underlying resumption in English. Glossa: A Journal of General Linguistics 4. 68. DOI: http://doi.org/10.5334/gjgl.839
Chaves, Rui P. & Putnam, Michael T. 2020. Unbounded dependency constructions: Theoretical and experimental perspectives. Oxford: Oxford University Press. DOI: http://doi.org/10.1093/oso/9780198784999.001.0001
Chomsky, Noam. 1973. Conditions on transformations. In Anderson, Stephen R. & Kiparsky, Paul (eds.), A festschrift for Morris Halle, 232–286. New York: Rinehart and Winston.
Chomsky, Noam. 1986. Barriers. MIT Press.
Choueiri, Lina. 2017. Resumption in varieties of Arabic. In Benmamoun, Elabbas & Bassiouney, Reem (eds.), The Routledge handbook of Arabic linguistics, 379–394. London: Routledge. DOI: http://doi.org/10.4324/9781315147062-8
Cinque, Guglielmo. 1990. Types of Ā-dependencies. Cambridge, MA: MIT Press.
Erteschik-Shir, Nomi. 1973. On the nature of island constraints. MIT dissertation.
Erteschik-Shir, Nomi. 1992. Resumptive pronouns in islands. In Goodluck, Helen & Rochemont, Michael (eds.), Island Constraints: Theory, acquisition and processing, 89–108. Dordrecht: Springer Netherlands. DOI: http://doi.org/10.1007/978-94-017-1980-3_4
Frazier, Lyn & Clifton, Charles Jr. 2002. Processing “d-linked” phrases. Journal of Psycholinguistic Research 31(6). 633–659. DOI: http://doi.org/10.1023/a:1021269122049
Frazier, Lyn & d’Arcais, Giovanni B. Flores. 1989. Filler driven parsing: A study of gap filling in Dutch. Journal of Memory and Language 28. 331–344. DOI: http://doi.org/10.1016/0749-596X(89)90037-5
Gibson, Edward & Hemforth, Barbara & Winckel, Elodie & Abeillé, Anne. 2021. Acceptability of extraction out of adjuncts depends on discourse factors. Presented at the 34th CUNY conference on human sentence processing.
Goldberg, Adele E. 2013. Backgrounded constituents cannot be “extracted.” In Sprouse, Jon & Hornstein, Norbert (eds.), Experimental syntax and island effects, 221–238. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139035309.012
Goodall, Grant. 2014. The D-linking effect on extraction from islands and non-islands. Frontiers in Psychology 5. DOI: http://doi.org/10.3389/fpsyg.2014.01493
Haegeman, Liliane & Jiménez-Fernández, Ángel L. & Radford, Andrew. 2014. Deconstructing the Subject Condition in terms of cumulative constraint violation. The Linguistic Review 31. 150–73. DOI: http://doi.org/10.1515/tlr-2013-0022
Hammerly, Christopher. 2021. The pronoun which comprehenders who process it in islands derive a benefit. Linguistic Inquiry 53. 823–835. DOI: http://doi.org/10.1162/ling_a_00422
Hammerly, Christopher & Staub, Adrian & Dillon, Brian. 2019. The grammaticality asymmetry in agreement attraction reflects response bias: Experimental and modeling evidence. Cognitive Psychology 110. 70–104. DOI: http://doi.org/10.1016/j.cogpsych.2019.01.001
Hayyas, Asmaa. 2020. Resumptive pronouns in Baha Arabic: An experimental study. Leeds: University of Leeds dissertation.
Hayyas, Asmaa & De Cat, Cecile. 2023. Resumptive pronouns in Baha Arabic wh-dependencies. DOI: http://doi.org/10.31219/osf.io/xpt3g
Hofmeister, Philip. 2011. Representational complexity and memory retrieval in language comprehension. Language and Cognitive Processes 26(3). 376–405. DOI: http://doi.org/10.1080/01690965.2010.492642
Hofmeister, Philip & Norcliffe, Elizabeth. 2013. Does resumption facilitate sentence comprehension? In Hofmeister, Philip & Norcliffe, Elizabeth (eds.), The core and the periphery: Data-driven perspectives on syntax inspired by Ian A. Sag, 225–246. Stanford, CA: CSLI Publications.
Hofmeister, Philip & Sag, Ivan A. 2010. Cognitive constraints and island effects. Language 86(2). 366–415. DOI: http://doi.org/10.1353/lan.0.0223
Huang, C. 1982. Move WH in a language without WH movement. The Linguistic Review 1(4). 369–416. DOI: http://doi.org/10.1515/tlir.1982.1.4.369
Huang, Nick & Almeida, Diogo & Sprouse, Jon. 2025. A nearly exhaustive experimental investigation of bridge effects in English: Supplementary material. Language 101. DOI: http://doi.org/10.1353/lan.2025.a955267
Jäger, Lena A. & Engelmann, Felix & Vasishth, Shravan. 2017. Similarity-based interference in sentence comprehension: Literature review and Bayesian meta-analysis. Journal of Memory and Language 94. 316–339. DOI: http://doi.org/10.1016/j.jml.2017.01.004
Keshev, Maayan & Meltzer-Asscher, Aya. 2017. Active dependency formation in islands: how grammatical resumption affects sentence processing. Language 93(3). 549–568. DOI: http://doi.org/10.1353/lan.2017.0036
Kluender, Robert. 1991. Cognitive constraints on variables in syntax. San Diego, CA: University of California dissertation.
Kluender, Robert & Kutas, Marta. 1993a. Bridging the gap: evidence from ERPs on the processing of unbounded dependencies. Journal of Cognitive Neuroscience 5(2). 196–214. DOI: http://doi.org/10.1162/jocn.1993.5.2.196
Kluender, Robert & Kutas, Marta. 1993b. Subjacency as a processing phenomenon. Language and Cognitive Processes 8(4). 573–633. DOI: http://doi.org/10.1080/01690969308407588
Kroch, Anthony. 1981. On the role of resumptive pronouns in amnestying island constraint violations. In Hendrick, Roberta A. & Masek, Carrie S. & Miller, Mary Frances (eds.), Papers from the seventeenth regional meeting of the Chicago Linguistic Society, 125–135. Chicago: Chicago Linguistic Society.
Kush, Dave & Lohndal, Terje & Sprouse, Jon. 2018. Investigating variation in island effects: A case study of Norwegian WH-extraction. Natural Language and Linguistic Theory 36(3). 743–779. DOI: http://doi.org/10.1007/s11049-017-9390-z
Laurinavichyute, Anna & von der Malsburg, Titus. 2024. Agreement attraction in grammatical sentences and the role of the task. Journal of Memory and Language. DOI: http://doi.org/10.31234/osf.io/n75vc
Laurinavichyute, Anna & Yadav, Himanshu & von der Malsburg, Titus & Vasishth, Shravan. 2023. The role of goal in sentence processing. In Architectures and Mechanisms for Language Processing (AMLaP). San Sebastian, Spain.
Liu, Yingtong & Winckel, Elodie & Abeillé, Anne & Hemforth, Barbara & Gibson, Edward. 2022. Structural, functional, and processing perspectives on linguistic island effects. Annual Review of Linguistics. DOI: http://doi.org/10.1146/annurev-linguistics-011619-030319
McDaniel, Dana & Cowart, Wayne. 1999. Experimental evidence for a minimalist account of English resumptive pronouns. Cognition 70(2). 15–24. DOI: http://doi.org/10.1016/s0010-0277(99)00006-2
Meltzer-Asscher, Aya. 2021. Resumptive pronouns in language comprehension and production. Annual Review of Linguistics 7(1). 177–194. DOI: http://doi.org/10.1146/annurev-linguistics-031320-012726
Open Science Tools Ltd. 2023. Pavlovia. Retrieved from https://pavlovia.org
Pañeda, Claudia & Lago, Sol & Vares, Elena & Veríssimo, João & Felser, Claudia. 2020. Island effects in Spanish comprehension. Glossa: A Journal of General Linguistics 5. 21. DOI: http://doi.org/10.5334/gjgl.1058
Peirce, Jonathan W. & Gray, Jeremy R. & Simpson, Sol & MacAskill, Michael & Höchenberger, Richard & Sogo, Hiroyuki & … Lindeløv, Jonas K. 2019. PsychoPy2: experiments in behavior made easy. Behavior Research Methods 51. 195–203. DOI: http://doi.org/10.3758/s13428-018-01193-y
Pham, Catherine & Covey, Lauren & Gabriele, Alison & Aldosari, Saad & Fiorentino, Robert. 2020. Investigating the relationship between individual differences and island sensitivity. Glossa: An International Journal of Linguistics 5(1). 1–17. DOI: http://doi.org/10.5334/gjgl.1199
Phillips, Colin. 2013. On the nature of island constraints I: Language processing and reductionist accounts. In Sprouse, Jon & Hornstein, Norbert (eds.), Experimental Syntax and Island Effects, 64–108. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139035309.005
Ross, John Robert. 1967. Constraints on variables in syntax. Cambridge, MA: MIT dissertation.
Sag, Ivan A. & Hofmeister, Philip & Casasanto, Laura. 2013. Islands in the grammar? Standards of evidence. In Sprouse, Jon & Hornstein, Norbert (eds.), Experimental Syntax and Island Effects, 42–63. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139035309.004
Saul, Niki & Keshev, Maayan & Meltzer-Asscher, Aya. 2025. Interference in the formation of filler-gap dependencies: Evidence from Hebrew relative clauses. Journal of Memory and Language 143. 104626. DOI: http://doi.org/10.1016/j.jml.2025.104626
Shafiei, Nazila & Graf, Thomas. 2020. The subregular complexity of syntactic islands. In Proceedings of the Society for Computation in Linguistics (SCiL), Vol. 3, 1–10.
Sprouse, Jon. 2007. A program for experimental syntax: Finding the relationship between acceptability and grammatical knowledge. University of Maryland dissertation.
Sprouse, Jon & Caponigro, Ivano & Greco, Ciro & Cecchetto, Carlo. 2016. Experimental syntax and the variation of island effects in English and Italian. Natural Language and Linguistic Theory 34(1). 307–344. DOI: http://doi.org/10.1007/s11049-015-9286-8
Sprouse, Jon & Wagers, Matthew & Phillips, Colin. 2012. A test of the relation between working-memory capacity and syntactic island effects. Language 88(1). 82–123. DOI: http://doi.org/10.1353/lan.2012.0004
Stowe, Laurie A. 1986. Parsing WH-constructions: Evidence for on-line gap location. Language and Cognitive Processes 1(3). 227–245. DOI: http://doi.org/10.1080/01690968608407062
Traxler, Matthew J. & Pickering, Martin J. 1996. Plausibility and the processing of unbounded dependencies: An eye-tracking study. Journal of Memory and Language 35(3). 454–475. DOI: http://doi.org/10.1006/jmla.1996.0025
Truswell, Robert. 2011. Events, phrases, and questions. Oxford University Press. DOI: http://doi.org/10.1093/acprof:oso/9780199577774.001.0001
Tucker, Matthew A. & Idrissi, Ali & Sprouse, Jon & Almeida, Diogo. 2019. Resumption ameliorates different islands differentially: Acceptability data from Modern Standard Arabic. In Benmamoun, Elabbas & Bassiouney, Reem (eds.), Perspectives on Arabic linguistics XXX: Papers from the annual symposia on Arabic linguistics, Stony Brook, New York, 2016 and Norman, Oklahoma, 2017, 213–232. Amsterdam: John Benjamins. DOI: http://doi.org/10.1075/sal.7.09tuc
Villata, Sandra & Franck, Julie. 2024. An empirical investigation of featural similarity in wh-islands. International Journal of Linguistics 16(1). 1. DOI: http://doi.org/10.5296/ijl.v16i1.21596
Villata, Sandra & Sprouse, Jon & Tabor, Whitney. 2019. Modeling ungrammaticality: A self-organizing model of islands. In Proceedings of the 41st annual meeting of the Cognitive Science Society, 1178–1184. Montreal, QC, Canada.
Villata, Sandra & Tabor, Whitney. 2022. A self-organized sentence processing theory of gradience: The case of islands. Cognition 222. 104943. DOI: http://doi.org/10.1016/j.cognition.2021.104943
Wood, Simon N. 2011. Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society: Series B 73(1). 3–36. DOI: http://doi.org/10.1111/j.1467-9868.2010.00749.x






