1 Introduction
Constructions that inhibit the formation of long-distance dependencies like wh-questions are known as islands. In many languages, subjects are islands: a wh-phrase can be extracted from an object, as in (1), but not a subject, as in (2).
- (1)
- What did you notice that she broke <a bottle of __> in the kitchen?
- (2)
- *What did you notice that <a bottle of __> appeared in the kitchen?
Some islands vary cross-linguistically. For instance, Spanish is reported to permit extraction from subjects selectively: Post-verbal subjects are not islands (3a), while pre-verbal subjects are (3b) (Uriagereka 1988; Gallego & Uriagereka 2007).
- (3)
- Subject islands in Spanish (Uriagereka 1988)
- a.
- ¿De qué conferenciantes te parece que me van a impresionar <las propuestas ___>?
- b.
- *¿De qué conferenciantes te parece que <las propuestas ___> me van a impresionar?
- ‘Which speakers does it seem to you that the proposals by will impress me?’
The contrast in Spanish aligns with a cross-linguistic generalization: Stepanov (2007) argues that what unites languages that do not exhibit subject islands is the ability to extract from a post-verbal position.
However, experimental studies of islands sometimes find patterns differing from reports in the literature, leading to calls for “systematically re-testing languages for island effects” (Sprouse & Villata 2021: 252). The need for such verification is especially acute for languages beyond English (Chacón 2021). Responding to this need, three experiments have examined subject islands in Spanish. Pañeda et al. (2020), López-Sancio (2015), and Stigliano et al. (2025) found evidence of island effects, but they tested only pre-verbal subjects, where effects were expected. What makes subject islands in Spanish noteworthy, though, is the claim that post-verbal subjects are not islands, which has not yet been tested experimentally.
We test the role of subject position using a formal acceptability judgment experiment to compare extraction from pre-verbal and post-verbal positions. In addition to the empirical contribution of verifying the contrast, our study addresses two key debates in experimental approaches to islands.
First, island effects are of theoretical interest because their source continues to be debated, with prominent approaches appealing to syntactic constraints, information structure clashes, or processor overload (see §2.3). We address this debate by holding all else equal as much as possible between subject positions to attempt to isolate the role of syntactic position from the role of information structure, and we control several properties of subjects and extractees not fully incorporated in previous experiments. We also explore processing effects by examining individual working memory variation.
Second, the last two decades have seen an experimental turn in studies of morphosyntax, yet open questions remain regarding how to interpret variation in judgment studies (see §2.4). We address this debate by examining our evidence from multiple angles and critically considering how we interpret statistical significance and effect sizes in rating tasks.
Previewing our results, we find statistically significant island effects in both positions, but pre-verbal subjects produce much larger effects than post-verbal subjects, with more individual variation. In post-verbal position, island violations are not rated substantially lower than non-violation sentences, and ratings cluster at the high end of the scale. We conclude our experiment finds a meaningful contrast by subject position, suggesting that pre-verbal subjects restrict extraction while post-verbal subjects do not. We interpret the results to largely support a structural view emphasizing the role of syntactic position.
2 Background & motivation
2.1 Subject islands in Spanish
Subject islands are historically regarded as strong islands, from which no extraction is allowed (Szabolcsi & Lohndal 2017).
- (4)
- Subject islands (Szabolcsi & Lohndal 2017)
- a.
- *Which book do you believe <the first chapter of ____> to be full of lies?
- b.
- *Which man does <every friend of ____> admire Lincoln?
Subject islands vary cross-linguistically. Some languages do not (always) display subject island effects. Stepanov (2007) argues extraction is related to post-verbal subjects, pointing to languages like Japanese (Omaki et al. 2020), Russian (King 1994), Hindi (Stepanov 2008), German (Haider 1993), and Palauan (Georgopoulos 1991); to these we can add Czech and Slovak (Starke 2001), Dutch (Broekhuis 2011), Italian (Bianchi & Chesi 2014), and Hungarian (Surányi & Turi 2018).
For Spanish, Uriagereka (1988) presented example (5), repeated from (3), to demonstrate that extraction from subjects is permitted selectively: Post-verbal subjects are not islands (5a), while pre-verbal subjects are (5b).
- (5)
- Subject island (Uriagereka 1988)
- a.
- ¿De qué conferenciantes te parece que me van a impresionar <las propuestas ___>?
- b.
- *¿De qué conferenciantes te parece que <las propuestas ___> me van a impresionar?
- ‘Which speakers does it seem to you that the proposals by will impress me?’
Subsequently, Gallego and Uriagereka (2007) and Gallego (2011) presented additional evidence from Spanish; for instance, a similar contrast by position emerges with intransitives, as in (6).
- (6)
- Subject island (Gallego 2011)
- a.
- ¿De qué coche parece que ya ha llegado <el conductor ___>?
- b.
- ??¿De qué coche parece que <el conductor ___> ya ha llegado?
- ‘Of which car does it seem that the driver has already arrived?
Position is not the only factor, though. For instance, Gallego notes subextraction “in a monoclausal environment” (7a) or from non-final position (7b) is degraded (Gallego 2011: 51–52).
- (7)
- Illicit extractions from post-verbal position (Gallego 2011)
- a.
- ??¿De qué coche ha ganado dos carreras <el piloto ___>?
- b.
- *¿De qué coche ha ganado <el piloto ___> dos carreras?
- ‘Of which car has the driver won two races?’
Haegeman et al. (2014) identified several further syntactic and semantic constraints. In addition to subject position, they highlighted properties of the subject and extractee that affect acceptability.
Building on previous work,1 Haegeman et al. pointed out that extraction from non-specific or non-referential subjects, as in (8a), is better than specific subjects (8b). Theta role also matters (9): extraction from higher in the thematic hierarchy, especially from agents (9a), is worse than lower in the hierarchy, like themes (9b).
- (8)
- Specificity (Haegeman et al. 2014)
- a.
- ?¿De qué cantante te parece que <varias fotos ___> me van a escandalizar?
- ‘Of which singer does it seem to you that several photos will shock me?’
- b.
- *¿De qué conferenciantes te parece que <las propuestas ___> me van a impresionar?
- ‘Which speakers does it seem to you that the proposals by will impress me?’
- (9)
- Theta role (Haegeman et al. 2014)
- a.
- ?*¿De qué electrodoméstico parece que causó <el inventor ___> tanta conmoción?
- ‘Of which electrical appliance does it seem that the inventor caused such a stir?’
- b.
- ??¿De qué electrodoméstico parece que causó <el invento ___> tanta conmoción?
- ‘Of which electrical appliance does it seem that the invention caused such a stir?’
Properties of the extractee also affect acceptability: extracting a D-linked (complex) wh-phrase (10a) is better than a simple wh-word (10b),2 and extracting arguments of the head noun, as in (11a), is better than adjuncts (11b).
- (10)
- D-linking
- a.
- ?¿De qué cantante te parece que <varias fotos ___> me van a escandalizar? (Haegeman et al. 2014)
- ‘Of which singer does it seem to you that several photos will shock me?’
- b.
- *¿De qué sabe María que <una botella ___> se cayó de la mesa? (Goodall 2011)
- ‘Of what does María know that a bottle fell from the table?’
- (11)
- Argument vs. Adjunct (Haegeman et al. 2014)
- a.
- ¿De qué político crees que han causado tanta conmoción <algunas propuestas ___>?
- ‘Of which politician do you think that some proposals caused a stir?’
- b.
- *¿Con cuántos rotos crees que han causado tanta conmoción <unos vaqueros ___>?
- ‘With how many holes do you think that some jeans have caused a stir?’
For Haegeman et al., violations of each of these proposed constraints—along with constraints on subject position—accumulate, such that a D-linked argument extracted from a non-specific, post-verbal theme subject of a passive should be completely well formed (12a), while a plain wh-word adjunct extracted from a specific pre-verbal agent subject of a transitive should be quite degraded (12b), with these options forming the ends of a continuum of gradient acceptability.
- (12)
- Acceptability extremes for subject extraction (modified from Haegeman et al. 2014)
- a.
- ¿De qué príncipe fueron publicadas <varias fotos comprometedoras ___>?
- ‘Of which prince were several compromising photos published?’
- b.
- *¿Con qué en la cara dijiste que <el hombre ___> asesinó a una señora anciana?
- ‘With what on his face did you say the man murdered an old lady?’
In summary, several semantic and syntactic factors condition extraction from subjects in Spanish, and descriptions concur that subject position is a crucial factor.
2.2 Experimental evidence
Although most formal experiments on islands have corroborated informal judgments, differences have emerged (Sprouse & Villata 2021), so experimentally confirming reported judgments is important (Chacón 2021). Indeed, previous experiments have found other islands in Spanish differ from the descriptions in the theoretical literature (e.g., Hoot & Ebert 2024, Pañeda et al. 2024).
Three experiments have examined subject islands in Spanish and found evidence of island effects.
López-Sancio (2015) carried out a seven-point rating task testing several islands in both wh- and relative clause dependencies with Peninsular Spanish speakers. For the subject islands, all subjects were pre-verbal and non-specific. He did not control for theta role, D-linking, or other properties of the extractee. His results revealed evidence of an island effect with wh-questions but not with relative clause dependencies.
Pañeda et al. (2020) conducted a timed yes/no task with Peninsular Spanish speakers to test several islands, including subjects. They expressly designed their items to provoke an island effect—extracting plain wh-words from specific, pre-verbal agentive subjects—and they found such sentences were starkly unacceptable, reaffirming that pre-verbal subjects in Spanish are islands, at least when the properties of subjects and extractees align with islandhood.
López-Sancio operationalized the island effect slightly differently from Pañeda et al.: instead of comparing extraction from the embedded clause to extraction in the matrix clause, his experiment compared extraction from embedded subjects to extraction from embedded objects (following Sprouse et al. 2016). The third study, Stigliano et al. (2025), asked whether these two ways of defining the design for subject islands might have affected the experiments’ outcomes. They obtained judgments on a seven-point scale from Argentine Spanish speakers and found both designs produced similar effects. Stigliano et al. also examined individual variation and pointed out that the island violation condition produced substantial inter-individual variation in the judgments, unlike the other conditions.
In sum, the existing experimental evidence suggests that pre-verbal subjects in Spanish provoke island effects. However, they all tested only the pre-verbal position, where such effects were predicted, and López-Sancio (2015) did not control properties of the extractee, while Pañeda et al. (2020) included properties designed to provoke an effect. Yet for Spanish, the key claim that post-verbal position allows extraction remains untested.
Beyond Spanish, experiments investigating subject islands in languages that generally allow post-verbal subjects have found a wide range of results: no subject/object asymmetry (Japanese: Omaki et al. 2020); no effect by subject position (Russian: Polinsky et al. 2013; Hungarian: Surányi & Turi 2018), potentially due to semantics (Italian: Bianchi & Chesi 2014); a weak but statistically significant effect in post-verbal position (Romanian: Schoenmakers & Stoica 2024); a cumulative effect combining movement and a subject/object contrast (German: Jurka 2010); or even higher acceptability in pre- than post-verbal position (Russian: Belova 2021). This range of results could be due in part to methodological differences. For example, some of these studies did not isolate subject position from other factors known to affect acceptability, e.g. specificity (Jurka 2010; Polinsky et al. 2013; Belova 2021), and most did not use the factorial design that helps isolate the island effect (Schoenmakers & Stoica 2024 is an exception). Since no consistent pattern has emerged, new evidence from post-verbal subjects in Spanish stands to contribute to understanding subject islands cross-linguistically.
2.3 Source of island effects
The mere presence of an effect in a judgment task does not distinguish its source, and Liu et al. (2022) identify three main families of explanations for island effects, which they call “structural,” “functional,” and “processing” accounts (see also Newmeyer 2016 for a similar typology). These theories have been applied to explain islands of all types; we focus especially on accounts of subject islands.
Structural accounts appeal to syntactic constraints represented in the speaker’s competence grammar. In syntactic theory, subject islands are explained by positing a grammatical restriction that prevents extraction from subjects once they move to Spec-TP. The long history of these formal mechanisms includes the Condition on Extraction Domains (Huang 1982), Criterial Freezing (Rizzi 2006), probe/goal theory (Chomsky 2000), and phases (Chomsky 2008). The two most prominent structural accounts of Spanish subject islands are Gallego’s phase-based approach and Haegeman et al.’s cumulative constraint approach. Gallego (2011) argued extraction is evaluated at the phase level. Pre-verbal subjects in Spec-TP are rendered inactive due to A-movement to Spec-TP and Case valuation, prohibiting extraction. Post-verbal, in situ subjects in Spec-v*P, without A-movement or Case valuation, remain active when evaluated at the v*P phase, permitting extraction. On the other hand, Haegeman et al. (2014) analyzed subject islands as responding to a set of cross-linguistic syntactic and semantic constraints, which are violable and cumulative, with unacceptability increasing as more constraints are violated. Comparing Spanish and English, they analyze the difference between the two languages as due to independent availability of post-verbal subjects in Spanish, whereas English generally requires movement to Spec-TP. These accounts both predict that pre-verbal subjects in Spanish should create an island effect while post-verbal subjects should not, after avoiding other constraints.
Functional accounts also appeal to the speaker’s competence but locate the problem in a clash of information structure or pragmatic/semantic anomaly. In these approaches, no syntactic restriction bars extraction from subjects; instead, the information-structural effects or semantics of question-formation or relativization clash with the subject’s discourse status, producing ill-formedness. For example, Abeillé et al. (2020) propose the Focus-Background Conflict constraint: “A focused element should not be part of an unfocused/backgrounded constituent.” Subjects are usually background; wh-words are focused. Thus, extracting a focused wh-word out of a backgrounded subject produces an information-structural clash.3 This is not a syntactic restriction on extraction from a particular position; it’s a discourse restriction on extraction from constituents with a particular information status. Indeed, experiments have found that ratings of subject islands are linked to the degree of topicality of the subject (Chaves & King 2019) and that information structure affects judgments of subject islands (Chaves & Dery 2019; Abeillé et al. 2020). These accounts predict that, all else being equal,4 if the relevant information structure properties are held constant between pre- and post-verbal subjects, their acceptability should be comparable.
Processing accounts differ from the other two in appealing to difficulty parsing island structures in real time. Since the seminal work of Kluender and Kutas (1993), island effects have been attributed to an overload on limited working memory due to their complexity (Hofmeister & Sag 2010) or to difficulty in retrieving the referent from memory upon encountering the gap (Keshev & Meltzer-Asscher 2019). In resource-limitation accounts, island structures are not grammatically restricted but instead cause processing breakdown due to limits on working memory or other cognitive capacities. Because cognitive capacity varies by individual, Sprouse et al. (2012) predicted that perceived unacceptability would covary with language processing abilities, such as working memory.5 They did not observe the expected correlation, and several studies have replicated their findings (Michel 2014; Aldosari 2015; Pañeda et al. 2020; Pham et al. 2020), although others have raised questions about the adequacy of the working memory tests employed (Hofmeister & Casasanto & Sag 2012a; 2012b), and work on the relationship of sentence processing to island effects remains an extremely active area of research (see Liu et al. 2022; Sprouse & Villata 2021).
We attempt to isolate structural factors from discourse and processing factors in two ways. First, we used contexts to hold the information structure constant across subject positions. The contexts make the same information given or new in both subject positions, with the aim of disentangling position and information structure. Second, we followed previous work in measuring working memory (with a task addressing some earlier critiques) and examining whether it correlated to individual island sensitivity, to identify a hypothesized relationship between sentence-processing ability and island effects.
2.4 Factorial designs and interpretation in experimental syntax
We adopt the factorial definition of islands popularized by Sprouse and colleagues (e.g., Sprouse et al. 2012; 2016). Concretely, we adopt Pañeda et al.’s (2020) 2 × 2 design for subject islands. Recall that previous studies have operationalized subject islands in two ways, but Stigliano et al. (2025) found both methods produced equivalent results.
Factorial designs isolate a predicted grammatical effect from other features that may impact acceptability by crossing two factors. For instance, since more complex sentences may receive lower ratings, perhaps the mere presence of a complex island structure makes a sentence less acceptable. In the experiment paradigm in (13), we operationalize this difference as the Island factor: the presence (13a) or absence (13b) of the complex subject. Similarly, all else being equal, long-distance dependencies are less acceptable than shorter dependencies (Goodall 2021), which we operationalize as the Gap factor: a short-distance dependency involving only the matrix clause (13a/b) versus a long-distance dependency into the embedded clause (13c/d).
- (13)
- Experimental paradigm
- a.
- Non-island, Matrix: Which friend ___ thinks that some stories delighted María?
- b.
- Island, Matrix: Which friend ___ thinks that <some stories of Cortázar> delighted María?
- c.
- Non-island, Embedded: Which stories do you think ___ delighted María?
- d.
- Island, Embedded: Of which author do you think <some stories ___> delighted María?
Sprouse and colleagues argue that observing an effect greater than the linear sum of each individual effect in the interaction condition (13d), which they call a super-additive effect, indicates the presence of a grammatical restriction specific to that condition. A super-additive effect can be observed as a significant statistical interaction and a diverging-lines pattern in an interaction plot (Figure 1).
While this approach to operationalizing island effects has become standard in experimental syntax, recent work has raised questions about interpreting these effects.
Stigliano et al. (2025) found a statistical interaction without finding the expected acceptability decrease for each factor alone, sometimes called a ‘non-monotonic’ pattern. The logic of the factorial design assumes each factor degrades acceptability. However, sometimes one factor does not evince the expected effect, which requires researchers to look beyond statistical significance to the details of the interaction.
More broadly, exclusive reliance on statistical significance is widely acknowledged as problematic. Nearly any putative effect can be statistically significant with sufficient measurement precision and a large enough sample, which has led many methodological authorities within linguistics (e.g., Plonsky 2015) and beyond (e.g., Ferguson 2009) to urge greater emphasis on effect sizes. Yet the field of experimental syntax has reached no consensus regarding effect size interpretation (Sprouse & Villata 2021).
Third, acceptability judgments, as behavioral measures, are generally variable and gradient, yet there is similarly no consensus on the interpretation of this variation (see Francis 2022; Schütze & Sprouse 2013 for discussion).
In summary, the factorial design for islands is a powerful, widely used tool to detect island effects, but, like all experimental designs, the results require interpretation. We address these issues—with the dual aims of contributing to ongoing methodological discussions in the field and strengthening the interpretation of our results—as we interpret our experiment, examining the details of the super-additive effects we find, reporting effect sizes, and examining gradience and variation.
2.5 Research questions
Our research questions are: (i) do Spanish speakers display a contrast in island effects between subject positions when other factors are held constant as much as possible, and (ii) do island effects on judgments vary by individual working memory capacity?
3 Methods
3.1 Materials
3.1.1 Psych verbs
To test the position effect while holding most other factors constant, our experiment leverages the features of a type of Spanish verb known as Class II psych(ological) verbs.6 Class II psych verbs include confundir ‘confuse,’ emocionar ‘move/thrill,’ and asustar ‘frighten’. The subject of these verbs always receives a theta role of theme (Belletti & Rizzi 1988; Halloran Gonzalez 2018) or cause(r) (Parodi-Lewin 1991; Pesetsky 1996); the other argument’s theta role is always experiencer. The experiencer can be realized in one of two syntactic roles—either an accusative direct object or a dative indirect object—and each option corresponds to a different canonical word order—SVO for accusative experiencers (14) and OVS for dative experiencers (15) (Parodi-Lewin 1991; Gutiérrez-Bravo 2006; Halloran Gonzalez 2018).
- (14)
- ACC-experiencer (SVO)
- Los
- the
- temblores
- earthquakes
- asustan
- frighten.3pl
- a
- acc
- Juan.
- Juan
- ‘Earthquakes frighten Juan.’
- (15)
- DAT-experiencer (OVS)
- A
- dat
- Juan
- Juan
- le
- cl.dat
- asustan
- frighten.3pl
- los
- the
- temblores.
- earthquakes
- ‘Earthquakes frighten Juan.’
The accusative/dative alternation is also marked by two other differences. First, the dative-experiencer version (15) doubles the indirect object with the dative clitic le, which is absent when the experiencer is accusative (14).7 Second, the accusative-experiencer realization has an eventive (Parodi-Lewin 1991) or change-of-state (Halloran Gonzalez 2018) reading, whereas the dative-experiencer has a stative reading.8 This means that a sentence like (16) corresponds to a reading in which there is some event that produces a change in Juan (e.g., an earthquake occurs and he becomes afraid), whereas (17) describes a state of being without a specified change (i.e., Juan is afraid of earthquakes).
To establish the [±change-of-state] interpretation, we embedded our sentences in contexts, and the accusative-experiencer verbs used preterite—the perfective past tense form indicating completed actions—while the dative-experiencer verbs used simple present to suggest an ongoing state. These two choices serve to support the appropriate semantics for either pre-verbal or post-verbal subjects, which helps make the subject position both unambiguous and unmarked/canonical. Table 1 and Table 2 present sample tokens and their contexts.
Table 1: Pre-verbal subject tokens.
| ISLAND | GAP | SAMPLE TOKEN | GLOSS |
| NON-ISLAND | Matrix | Context: Un colega vio que cuando María leía un libro de cuentos se emocionó a tal punto que se echó a llorar. ¿Qué colega ___ piensa que varios cuentos emocionaron a María? |
Context: A colleague saw that when María was reading a book of short stories she was moved to the point that she began to cry. Which colleague ___ thinks that several stories moved María? |
| Embedded | Context: Cuando María leía un libro de cuentos se emocionó a tal punto que se echó a llorar. ¿Qué cuentos piensas que ___ emocionaron a María? |
Context: When María was reading a book of short stories, she was moved to the point that she began to cry. Which stories do you think ___ moved María? |
|
| ISLAND | Matrix | Context: Un colega vio que cuando María leía algunos cuentos de Cortázar se emocionó a tal punto que se echó a llorar. ¿Qué colega ___ piensa que varios cuentos de Cortázar emocionaron a María? |
Context: A colleague saw that when María was reading some Cortázar short stories she was moved to the point that she began to cry. Which colleague ___ thinks that several stories by Cortázar moved María? |
| Embedded | Context: Cuando María leía algunos cuentos de su autor favorito se emocionó a tal punto que se echó a llorar. ¿De qué escritor piensas que varios cuentos ___ emocionaron a María? |
Context: When María was reading some stories by her favorite author, she was moved to the point that she began to cry. By which author do you think that several stories ___ moved María? |
Table 2: Post-verbal subject tokens.
| ISLAND | GAP | SAMPLE TOKEN | GLOSS |
| NON-ISLAND | Matrix | Context: María le mencionó a un colega del trabajo que lee y relee un libro específico de cuentos porque le emociona mucho. ¿Qué colega ___ piensa que a María le emocionan varios cuentos? |
Context: María mentioned to a work colleague that she reads and rereads a specific book of short stories because it really moves her. Which colleague ___ thinks that several stories move María? |
| Embedded | Context: María te mencionó que lee y relee un libro específico de cuentos porque le emociona mucho. ¿Qué cuentos piensas que a María le emocionan ___? |
Context: María mentioned to you that she reads and rereads a specific book of short stories because it really moves her. Which stories do you think ___ move María? |
|
| ISLAND | Matrix | Context: María le mencionó a un colega que lee y relee algunos cuentos de Cortázar porque le emocionan mucho. ¿Qué colega___ piensa que a María le emocionan varios cuentos de Cortázar? |
Context: María mentioned to a work colleague that she reads and rereads some Cortázar short stories because they really move her. Which colleague ___ thinks that several stories by Cortázar moved María? |
| Embedded | Context: María te mencionó que lee y relee algunos cuentos de su autor favorito porque le emocionan mucho. ¿De qué escritor piensas que a María le emocionan varios cuentos ___? |
Context: María mentioned to you that she reads and rereads some stories by her favorite author because they really move her. By which author do you think that several stories ___ move María? |
Our goal was to hold properties relevant to subject islands constant as much as possible, yet our pre- and post-verbal conditions were not perfect minimal pairs. We introduced two differences—the [±change-of-state] context and the verbal aspect to reinforce that reading—and one difference was inherent in them—the dative/accusative alternation. These choices served to make the pre- or post-verbal subject canonical, default, and unambiguous with as few differences as possible. We also controlled information structure and other features to make the items maximally comparable.
3.1.2 Controlling information structure
The subject NPs were identical for the pre- and post-verbal contexts, rendering the same information given, presupposed, or present in the background for both conditions, and holding semantic properties other than [±change-of-state] constant.
Schwarzschild (1999) defines discourse-given in terms of presupposition.9 A simple wh-phrase like Who saw Bill?, presupposes X saw Bill. Thus, “X saw Bill” is given, as is everything in the resulting sentence except what’s inserted for X: [John]F saw Bill. If the same holds for our questions, we see no reason to expect the presuppositions to be different between our pre- and post-verbal subjects. As shown in (16) and (17), the entire NP several stories by X author is part of the presupposition of both question types. That is, in each case, there is a predicate move, which takes two arguments, an individual María and the subject NP: move(s, María) & s = [several stories by X author]. This suggests that the focus in both sentences is the variable X, while all else is given.
- (16)
- ¿De qué escritor piensas que varios cuentos ___ emocionaron a María?
- ‘By which author do you think that several stories ___ moved María?’
- PRESUPPOSES:
- You think that [several stories by X author] moved María.
- (17)
- ¿De qué escritor piensas que a María le emocionan varios cuentos ___?
- ‘By which author do you think that several stories ___ move María?’
- PRESUPPOSES:
- You think that [several stories by X author] move María.
On the other hand, although the subjects in our sentences are discourse-given, they could also be interpreted as foci. For instance, because Spanish is a pro-drop language in which backgrounded subjects are generally omitted, Biezma (2014) argues that all overt subjects in Spanish with salient antecedents are foci. Neither of these definitions would predict a difference in discourse status by subject position.10
Nevertheless, there is a general tendency in Spanish for pre-verbal subjects to be topics (Casielles-Suárez 2004) and for focus to be sentence-final (Heidinger 2022). Context can override both defaults, however: numerous judgment and production studies show pre-verbal subjects are acceptable or preferred for Spanish subject focus (Heidinger 2022; Hoot 2025), and a corpus study found about one-third of narrowly focused subjects in Spanish are pre-verbal (Heidinger 2022). Furthermore, we created sentences in which both pre- and post-verbal subjects were in their canonical, unmarked position, making clear that no focus movement or topicalization has occurred. Instead, the context establishes the focal and background information.
An anonymous reviewer points out that our results suggest information structure was controlled effectively by the context. A strict syntax/discourse mapping would predict an overall difference in ratings by position. However, as we will see in §4, ratings of the matrix conditions with overt subjects are reasonably acceptable and do not differ by position, suggesting neither position is infelicitous overall and that the information status of both positions is the same.
Instead of focus and background, Chaves and Dery (2019: 509) argue that the essential discourse property determining the acceptability of subject islands is “the degree to which the subject-embedded referent is relevant for the assertion.” They apply a test for relevance to the “declarative counterparts” of subject island questions. Adopting this idea and taking the declarative counterparts of our sentences to be something like the presuppositions in (16) and (17), we contend that the identity of the author in question is relevant to the overall proposition in both contexts, which again suggests that discourse status is controlled across positions.
In summary, we attempted to control information structure in our test sentences by including contexts, which make the same information given and new in both pre- and post-verbal sentences while keeping the referent of the wh-word relevant to the proposition. Although the sentences’ complexity likely obscured diagnostics of the subjects’ discourse status, the crucial fact is that it was essentially the same between the two subject positions.
3.1.3 Controlling other features
To provide a fair test of the hypothesis that some subject positions permit extraction, we chose other features of the sentences to facilitate it. All our items were questions that extracted a D-linked (‘complex’) wh-phrase like qué amigo ‘which friend’ rather than simple wh-words like quién ‘who.’ The subject head noun was inanimate and explicitly non-specific (marked by indefinite quantifiers like ciertas ‘certain’ or varios ‘several’), and the extracted wh-phrase in the island violation sentences was an argument of the head noun, not an adjunct in the NP. Post-verbal subjects were final and part of a multiclausal sentence.
We created eight lexicalizations using the verbs sorprender ‘surprise,’ relajar ‘relax,’ confundir ‘confuse,’ asustar ‘frighten,’ inspirar ‘inspire,’ fascinar ‘fascinate,’ impresionar ‘impress/affect,’ and emocionar ‘move/thrill.’ All are very frequent (within the 5,000 most common words; Davies 2006). Verbs were plural to facilitate identifying the subject, which was always plural. The experiencer was a proper name. The wh-phrase was extracted to a matrix clause containing one of three declarative verbs (creer ‘believe,’ pensar ‘think,’ saber ‘know’),11 avoiding rogative verbs (like preguntar ‘ask’) that produce “quantitatively greater” island effects (Pañeda & Kush 2022: 496). When the gap was in the matrix clause, the matrix verb agreed with the wh-phrase (third-person singular); when the gap was in the subordinate clause, the matrix verb had second-person singular (tú ‘you’) agreement to avoid possible misinterpretation of the wh-word as the subject.12
Four native speakers of Mexican Spanish reviewed a representative sample of the items and contexts to maximize their naturalness, and we conducted a pilot study with 20 participants to identify any further problems, explicitly asking about plausibility in context. Feedback resulted in several improvements to contexts but not items.
3.1.4 Fillers and other materials
We created 16 fillers (a 2:1 ratio of fillers:target), adapted (by adding contexts) from a previous study. Fillers thus had known acceptability values, and we could ensure they ran the entire acceptability range and that the experiment contained a roughly equal proportion of acceptable to unacceptable sentences. Filler sentences were also questions.
The main judgment task was preceded by training materials that explained the task, emphasized the requested judgments were to be based on how much something “sounds like Spanish” and not on prescriptive criteria, and explicitly encouraged participants to use the entire scale. The training also included several anchoring items exemplifying judgments at the ends and middle of the scale (see Schütze & Sprouse 2013). After the training, participants completed ten announced practice items, then six unannounced practice items with a range of acceptability (two high, two low, two middle).
Participants also completed a short background questionnaire on language history and use.
Materials are available at https://osf.io/dq8mp.
3.1.5 Measuring working memory
We used a backward digit span task (BDS; Wechsler 1997) to measure working memory.13 Like the span tasks Pañeda et al. (2020) and Pham et al. (2020) used, this task was intended to address criticisms of some early work investigating working memory by requiring both recall and a transformation operation.
The BDS task presents single-digit numbers one by one and asks participants to recall them in the opposite order of presentation by entering the numbers on the keyboard. The task presents two sequences of numbers of length 2, then two sequences of length 3, and so on up to length 8, a total of 16 possible trials, thus producing a score from 0 to 16. If a participant incorrectly recalls both sets of a given length, the task ends.
3.2 Procedure
The pre- and post-verbal items were divided into separate surveys. Each participant thus only judged sentences with pre- or post-verbal subjects (a between-subjects design).
This decision was made primarily to avoid satiation effects. Satiation is a phenomenon in which unacceptable sentence types become more acceptable with repeated exposures (Snyder 2021). A meta-analysis of 21 studies on English found subject islands “reliably” exhibit satiation (Lu & Frank & Degen 2024: see especially Fig. 4), though Goodall (2011) did not find satiation with subject islands in Spanish. To avoid potential satiation effects, we chose to limit the experiment’s length and the number of sentences per cell of the design. Thus, within each survey, we created four lists, with tokens assigned via Latin square. Each list contained two tokens per cell, so participants judged eight target sentences (plus 16 fillers, for 24 total judgments).
This decision potentially reduces statistical power. Within-subjects comparisons have better power because individual participant-level variation is controlled. They increase reliability for the same reason. Fewer items per cell also could reduce reliability (Schoenmakers 2025). However, Likert-style numerical scales with the z-score transformation, as we employ here, have high reliability and good power in both within- and between-subjects contexts (see Langsford et al. 2018). Thus, we judged the tradeoff was worth it to avoid satiation and fatigue effects, and we conducted a power analysis (see §3.3) to ensure our samples would be large enough.
We examined the ratings of the fillers to verify that participants did not behave differently on the task overall by version (pre- or post-verbal); we observed no differences (Welch’s t = –1.56, p = .12). Although we believe both groups were drawn from the same population and we have no evidence for differences between the groups, we acknowledge that the between-subjects design could perhaps introduce some error due to characteristics of the samples that we did not measure.
Participants were recruited from Prolific and completed the experiment via PCIbex (Zehr & Schwarz 2018). Participants were randomly assigned to either the pre- or post-verbal subject condition, then to one of four lists. Token presentation order was pseudo-randomized such that two target items never appeared consecutively. Participants recorded judgments on a seven-point scale, with the ends labeled mal ‘(sounds) bad’ and bien ‘(sounds) good’ and intermediate values unlabeled.
3.3 Participants
We conducted a power analysis to determine the required sample size.14 We estimated groups of 40 would be sufficient to detect the relevant effects. To allow for participant exclusions, we recruited 60 people for each of the pre- and post-verbal conditions (120 total). Six participants had technical problems and did not complete the task. Of the remainder, 59 were sorted into the pre-verbal condition and 55 into post-verbal.
Participants were pre-screened on Prolific, and we subsequently checked background questionnaire responses to confirm all participants met inclusion criteria. We then examined the data for evidence of “non-cooperative” behavior, following Juzek and Häussler (2015).
First, randomly inserted among the items were three “instructional manipulation checks” (Prolific Team 2022), which verify participants’ attention by asking them to select a particular answer (e.g., “This is not a regular sentence; it is to verify attention. Please select 5.”). No participants were excluded for incorrectly answering two or more of these.
Second, we examined reaction times. Juzek (2016) recommends taking half the expected reading time for the shortest sentence in the experiment as the cutoff for inattentive behavior, calculating the expected reading time from the formula used by Bader and Häussler (2010) for a speeded judgment task. Under these assumptions (concretely, 225 milliseconds per word plus 25 milliseconds per character, plus 500 milliseconds to make a judgment), the shortest sentence in our experiment would take 3075 ms to read and judge (without reading the context), so we concluded that any judgment registered in less than 1500 ms was unlikely to have been made in good faith. To identify consistently inattentive participants, we chose to exclude anyone who made 20% of their judgments faster than our threshold and excluded one participant.
Third, we examined the fillers. Following Pañeda and Kush (2022), we excluded participants whose mean rating for ungrammatical fillers was at or above the scale midpoint, reasoning that participants who are unwilling to reject ungrammatical sentences do not provide useful data. Fifteen participants were excluded.15
Fourth, we identified three ungrammatical and two grammatical fillers, which previously received ratings at the ceiling or floor of the scale, to serve as “booby-trap items” (Juzek 2016) to verify attentiveness. Ten participants who ‘missed’ (rating grammatical items below 4 or ungrammatical items above 4) two or more of these items were excluded as inattentive.
After all planned exclusions, 87 participants remained in the sample. Of these, 43 judged items with pre-verbal subjects and 44 post-verbal subjects. Forty-three were female, 43 male, and one non-binary, with a mean age of 27.4 years (range: 21–47). No one reported cognitive or linguistic disabilities.
Participants all reported speaking a Mexican variety of Spanish and being born and living in Mexico. Most previous experiments on islands in Spanish have examined speakers of Peninsular Spanish. We chose Mexican Spanish to broaden the field’s empirical coverage to the country with the largest number of Spanish speakers. Although Mexico is home to a range of dialects, descriptions of syntactic variation in Mexican Spanish (e.g., Gutiérrez-Bravo 2020) give us no reason to expect variation with regard to islands, and restricting recruitment to one country allows for at least some level of uniformity. All acquired Spanish before age 5 in Mexico and were raised monolingually (without another home language in childhood), although they were not monolingual at the time of testing, as English is required to navigate Prolific.
3.4 Data analysis
We z-score transformed each sentence rating to control for scale-use variation and skew (Schütze & Sprouse 2013). We then examined by-item interaction plots and calculated influence statistics (Cook’s D) to identify whether any individual item contributed an outsize effect on the results; no items showed evidence of being outliers.
For the statistical analysis, most experimental syntax studies treat transformed judgment data as linear and use linear mixed-effects modeling to perform statistical tests. Yet some scholars contend that ordinal regression, which directly models the ordinal nature of rating data, should be preferred (Liddell & Kruschke 2018; Bürkner & Vuorre 2019). Following Kobzeva et al. (2022), we carried out both analyses, but, on the advice of an anonymous reviewer, we report only the linear model in the main text, since the results are the same and our power analysis assumed a linear model.
We conducted one test for pre-verbal position and one for post-verbal position, using lme4 (Bates et al. 2015) in R (v. 4.2.2; R Core Team 2022). For each, the fixed factors were Island (Yes/No), Gap (Matrix/Embedded), and the Island×Gap interaction. The interaction is always the fixed factor of interest. We also carried out an analysis with a three-way interaction among Subject Position (Pre/Post), Island, and Gap to determine if the effect was different by position. Because we carried out multiple tests, we adjusted the alpha level for each test using the Holm-Bonferroni procedure (Holm 1979) to control the familywise error rate. To explore significant interactions, we conducted post hoc pairwise comparisons using emmeans (Lenth & Piaskowski 2026), with the same adjustment for multiple comparisons.
We accounted for repeated measures by including random effects by participant and item. We used a top-down approach to find the maximal random effects structure (Barr et al. 2013), beginning with random intercepts and all random slopes, then removing the random slope accounting for the least variance and repeating this procedure as needed to achieve convergence.
In the text we present only key statistical results. The full output tables, including the random effects structures, estimates of fixed effects, and the ordinal regression analysis not reported here, are in a supplementary file available at https://osf.io/dq8mp.
As measures of effect size for each condition, we calculated difference-in-differences (DD) scores (Maxwell & Delaney 2004). Many methodological authorities urge increased emphasis on effect sizes in addition to (or instead of) null hypothesis significance testing (Cumming 2012; Lakens 2013; Plonsky 2015). The effect size measure to choose and its interpretation are field-specific (Plonsky & Oswald 2014; Larson-Hall & Plonsky 2015). DD scores are widely used in experimental syntax because they facilitate understanding the factorial design’s interaction and comparing with other studies, which is crucial for interpreting effect sizes (Larson-Hall & Plonsky 2015).16 We discuss effect sizes in detail in §5.2.
After conducting group-level analyses, we examined individual variation. First, we plotted the distribution of z-scores by condition as histograms and scatterplots (following Kush & Lohndal & Sprouse 2018; 2019; Pañeda & Kush 2022). Second, we calculated an individual DD score for each participant as a measure of sensitivity to the relevant island effect and carried out regression analyses with the lm function in R (R Core Team 2022) to determine the effect of working memory on island sensitivity by individual (as in Sprouse et al. 2012). The dependent variable was individual DD score, and the independent variable was mean-centered BDS score. We conducted each test twice: once for all participants and once with participants with negative DD scores removed, following Sprouse et al. (2012) and Pham et al. (2020). The rationale for conducting an analysis without negative DD scores is that no account predicts such ‘subadditive’ effects, and including them “might potentially mask the ability to observe a relationship between the individual difference measures and DD scores” (Pham et al. 2020: 8). We report both analyses and again adjusted alpha using the Holm-Bonferroni procedure to avoid inflated Type I error.17
4 Results
Figure 2 displays the group-level results. We observe the characteristic diverging-lines pattern of a super-additive effect for both pre-verbal and post-verbal subjects, yet the effect is much larger for pre-verbal than for post-verbal subjects, and the ratings of the violation condition for the post-verbal subjects are above the scale midpoint (z = 0, representing the mean rating across all items for each individual).
Table 3 presents the mean ratings by condition.
Table 3: Mean ratings and z-scores.
| SUBJECTS | ISLAND | GAP | Z-SCORE MEAN [95% CI] | RAW RATING MEAN [95% CI] |
| PRE-VERBAL | Non-island | Matrix | 0.08 [–0.07, 0.24] | 4.80 [4.37, 5.23] |
| Non-island | Embedded | 0.60 [0.45, 0.76] | 5.94 [5.59, 6.30] | |
| Island | Matrix | 0.27 [0.12, 0.41] | 5.22 [4.84, 5.61] | |
| Island | Embedded | –0.27 [–0.42, –0.11] | 4.09 [3.67, 4.51] | |
| POST-VERBAL | Non-island | Matrix | 0.38 [0.25, 0.51] | 5.82 [5.37, 6.06] |
| Non-island | Embedded | 0.63 [0.53, 0.73] | 6.31 [6.07, 6.55] | |
| Island | Matrix | 0.33 [0.19, 0.48] | 5.72 [5.37, 6.06] | |
| Island | Embedded | 0.16 [0.01, 0.30] | 5.30 [4.92, 5.67] | |
| FILLERS | All | –0.35 [–0.42, –0.29] | 4.02 [3.88, 4.17] | |
| Grammatical | 0.83 [0.76, 0.90] | 6.60 [6.45, 6.75] | ||
| Ungrammatical | –0.59 [–0.66, –0.52] | 3.51 [3.36, 3.66] |
We observe a significant Island×Gap interaction both for pre-verbal subjects (p < .001) and post-verbal subjects (p = .004). A significant three-way interaction (p = .008) suggests the Island×Gap interaction differed by subject position.
In pre-verbal position, post hoc pairwise comparisons indicate the violation condition is rated significantly lower than the others. The non-island embedded condition is also rated higher than the others. Examining the confidence intervals of the mean z-scores in Table 3 shows that their ranges do not overlap with those of the other conditions, suggesting their means are likely different.
In post-verbal position, pairwise comparisons did not find evidence that the island embedded condition was rated lower than the island matrix condition (p = .110), and their confidence intervals overlap. This result does not indicate that the mean values are equal, but we did not find evidence of a difference. The non-island embedded condition is again rated higher than the others.
The effect for pre-verbal subjects (DD = 1.05) is large and in the typical range for island effects (Kush et al. 2019), while the effect for post-verbal subjects (DD = 0.43) is small, within the range of scores that Kush et al. argue deserve “closer scrutiny” (Kush et al. 2019: 401). To better characterize these effects, we next examine ‘second-order acceptability effects,’ including variation across and within participants.
To examine variation, we visualize the distribution of ratings in each condition via histograms with overlaid density plots in Figure 3. For pre-verbal subjects, we observe unimodal clusters around the high end of the scale in all but one condition; the violation condition presents a relatively flat distribution, with scores distributed across the range. For the post-verbal subjects, scores in all conditions, including the violation condition, cluster at the high end of the scale.
Variation in scores could be inter-participant (some people give high ratings and some low) or intra-participant (the same people give inconsistent ratings). To distinguish these possibilities, we follow Pañeda and Kush (2022) in plotting individual consistency by comparing each participant’s higher rating against their lower rating in Figure 4. Because each person gave two judgments per condition, we can see whether both an individual’s judgments are above the scale midpoint of z = 0 (green quadrant, top right), both below the midpoint (red quadrant, bottom left), or one above and the other below (white quadrant, top left). Because AJTs with scales are necessarily variable and gradient, some inter- and intra-individual variation is expected, and all conditions have some variation, but we can draw conclusions from overall patterns. Most people rate both instances of most conditions above 0 (clustering in the green quadrant). The exception is the embedded island condition with pre-verbal subjects, where relatively few participants rated both tokens above 0 and the largest number rated them both below 0 (red quadrant).
Turning to individual variation by working memory, linear regressions for both pre- and post-verbal subjects reveal no significant effect of BDS score on individual island sensitivity, whether all DD scores are included or only those at or above 0. Figure 5 plots working memory against individual island effect size, with the solid line for all participants and the dashed line for DD ≥ 0 (see Sprouse et al. 2012; Pham et al. 2020). The relatively flat slope of all four lines suggests a limited effect of working memory on individual island sensitivity. We do not observe a strong positive or negative correlation between BDS score and individual DD score.
5 Discussion
Our study set out to make three main contributions: verify empirical claims from the theoretical literature experimentally, address debates about data in experimental syntax, and evaluate theories of potential sources of island effects. This section interprets our results in these three areas and acknowledges the limitations of our conclusions.
5.1 Empirical contribution
We observe super-additive interactions suggesting island effects in both pre- and post-verbal position.
Yet we also observe a contrast: with all else held equal as much as possible, the effect for pre-verbal subjects is much larger than for post-verbal subjects. Additionally, post hoc tests reveal a significantly lower violation condition, versus other conditions, only pre-verbally. This violation condition also displays the largest variation in scores, including the fewest individuals consistently accepting it and the most consistent rejectors. We take the sum of the evidence to suggest the island violation in pre-verbal position is degraded.
For post-verbal subjects, the effect is statistically significant but small. The violation condition does not present clear evidence of a decrease in acceptability compared to control conditions. Instead, the statistically significant effect is likely driven by the increased acceptability of the non-island embedded condition. Scores cluster at the high end for the violation condition, which most participants consistently accepted, and its mean score is above the scale midpoint. We take the sum of the evidence to suggest the post-verbal island violation is not degraded. Instead, the statistically significant effect is attributable to other factors.
Finally, we observe no effect of working memory in either position.
Overall, our results support the claim that pre-verbal subjects in Spanish provoke island effects while post-verbal do not. As discussed in §2.2, experimental corroboration of this claim was necessary given that previous experiments have found divergence from informal judgments, and no previous experiments had tested subject position in Spanish. Our results also support Stepanov’s generalization linking post-verbal position and lack of island effects.
5.2 Interpreting experimental evidence
Interpreting our results requires engaging with debates about data interpretation in experimental syntax. In this section, we discuss how we interpreted our results and the implications of these moves for questions raised in §2.4.
5.2.1 Non-monotonic super-additivity
Like Stigliano et al. (2025), we found non-monotonic super-additive effects: the gap position factor does not always produce a decrease in acceptability. Instead, within the non-island conditions, extraction from the embedded clause was more acceptable than extraction from the matrix clause, i.e., the longer dependency was more acceptable than the shorter dependency.
Several other studies have found non-monotonic results, including for subject islands (Sprouse et al. 2012; 2016; Pañeda et al. 2020) and other islands (Kush et al. 2018; Al-Aqarbeh & Sprouse 2024; Schoenmakers & Stoica 2024). The usual interpretation is to leave the divergence uncommented and simply note the interaction as evidence of an island effect, although some studies (e.g., Schoenmakers & Stoica 2024; Stigliano et al. 2025) examine the non-monotonic effect explicitly.
Because our post hoc tests did not find evidence of decreased acceptability for the violation condition with post-verbal subjects, we interpreted the significant interaction in post-verbal position as driven not by reduced acceptability for the island violation, but by increased acceptability for the non-island embedded condition. Why might this condition be rated high? One possibility is that it is also generally the shortest sentence of the foursome. Additionally, it is the only sentence with no overt embedded subject (since the subject has been extracted), which perhaps seems simpler. Importantly, the higher rating for the non-island embedded condition appears the same across subject positions, suggesting a consistent factor driving the increase for both.
We have therefore heeded Stigliano et al.’s call to look beyond statistical significance, leading us to interpret effect sizes and variation.
5.2.2 Effect sizes
As in other social and behavioral sciences (see Ferguson 2009), interpretation of effect sizes is a major unresolved question in experimental syntax (Sprouse & Villata 2021), and clear benchmarks are lacking. As with non-monotonic effects, the most common approach is to focus primarily on statistical significance, not effect size. For example, Kim and Goodall interpret DD = 0.28 as “a very clear wh-island effect” (Kim & Goodall 2016: 5). Kush et al. (2019: 401) exemplify an alternative position: “Although there is, in principle, no quantitative threshold for defining a ‘true’ island effect, … DD scores for island effects typically fall within the range of 0.75–1.25 … [so,] any intermediate-sized island effect bears closer scrutiny.” Kush and colleagues argue not for a simple cutoff point but rather careful examination of smaller effects to understand what they represent.
Meaningful interpretation of effect sizes includes comparing to similar previous studies (Larson-Hall & Plonsky 2015). To facilitate such interpretation, we compare our results to previous subject island experiments in Table 4, which is modeled on Sprouse and Villata (2021: Table 10.2) and includes all studies for which we could obtain or calculate DD scores. Other experimental studies that include subject islands but use a method (e.g., Matchin et al. 2024) or design (e.g., Chaves & King 2019) not permitting the calculation of effect sizes the same way are not included. Three studies reporting DD scores but on a different scale (e.g., raw ratings rather than z-scores) are included for completeness, although they are less directly comparable.
Table 4: Experimental studies of subject islands by language, dependency, and effect size.19
| Study | Language | Dependency | Effect size | Scale |
| Present study: Pre-verbal Present study: Post-verbal |
Spanish Spanish |
Complex wh Complex wh |
1.1 0.4 |
z-score z-score |
| Abeillé et al. 2020: Exp. 1 | English | Relative clause | –0.3 | z-score |
| Abeillé et al. 2020: Exp. 2 | English | Relative clause, pied-piping | 0.0 (n.s.) | z-score |
| Abeillé et al. 2020: Exp. 2 | English | Relative clause, p-stranding | 0.6 | z-score |
| Abeillé et al. 2020: Exp. 3 | English | Complex wh, pied-piping | 0.4 | z-score |
| Abeillé et al. 2020: Exp. 3 | English | Complex wh, p-stranding | 1.2 | z-score |
| Abeillé et al. 2020: Exp. 4 | French | Relative clause | 0.0 (n.s.) | z-score |
| Abeillé et al. 2020: Exp. 5 | French | Complex wh | 0.7 | z-score |
| Aldosari 2015 | English | Complex wh | 1.6 | z-score |
| Cartner et al. 2026: Exp. 1 | English | Complex wh | 0.8 | z-score |
| Cartner et al. 2026: Exp. 2 | English | Relative clause | 0.5 | z-score |
| Cartner et al. 2026: Exp. 3 | English | Topicalization | 0.3 | z-score |
| Kush et al. 2018 | Norwegian | Bare wh | 1.4 | z-score |
| Kush et al. 2018 | Norwegian | Complex wh | 1.4 | z-score |
| Kush et al. 2019 | Norwegian | Topicalization | 1.7 | z-score |
| López-Sancio 2015 | Spanish | Bare & Complex wh | 0.8 | z-score |
| López-Sancio 2015 | Spanish | Relative clause | 0.4 (n.s.) | z-score |
| Omaki et al. 2020 | Japanese | Scrambling | 0.1 (n.s.) | z-score |
| Schoenmakers & Stoica 2024 | Romanian | Bare wh | 0.4 | z-score |
| Sprouse et al. 2011 | English | Bare wh | 1.5 | z-score |
| Sprouse et al. 2011 | Japanese | Wh in situ | 0.2 (n.s.) | z-score |
| Sprouse et al. 2012: Exp. 1 | English | Bare wh | 0.8 | z-score |
| Sprouse et al. 2012: Exp. 2 | English | Bare wh | 1.3 | z-score |
| Sprouse et al. 2016 | English | Bare wh | 0.6 | z-score |
| Sprouse et al. 2016 | English | Complex wh | 0.5 | z-score |
| Sprouse et al. 2016 | English | Relative clause | 0.5 | z-score |
| Sprouse et al. 2016 | Italian | Bare wh | 1.4 | z-score |
| Sprouse et al. 2016 | Italian | Relative clause | –0.1 (n.s.) | z-score |
| Stepanov et al. 2018 | Slovenian | Bare wh | 0.6 | z-score |
| Pañeda et al. 2020 | Spanish | Bare wh | –6.0 | log odds |
| Pham et al. 2020 | English | Complex wh | 3.9 | raw (7-pt) |
| Stigliano et al. 2025: SOD18 | Spanish | Complex wh | 2.1 | raw (7-pt) |
| Stigliano et al. 2025: SCSD | Spanish | Complex wh | 2.2 | raw (7-pt) |
Our pre-verbal extraction effect size (DD = 1.05) is similar to others found for subject islands in English (Sprouse et al. 2011: bare wh; 2012: Exp. 2; Aldosari 2015), Italian (Sprouse et al. 2016), and Norwegian (Kush et al. 2018; 2019), and larger than López-Sancio’s (2015) Spanish wh-questions. In general, the pre-verbal extraction effect is comparable to subject island effects found cross-linguistically. In contrast, our post-verbal extraction effect size (DD = 0.43) is among the smallest of any significant interaction.
We visualize the distribution of effect sizes in Figure 6.
Although there is no simple cutoff for a meaningful effect size, comparing our effects with those of previous experiments suggests a difference by position.
5.2.3 Variation
As with all judgment tasks, we observed variation in all conditions, but we found more variation in the pre-verbal island violation condition than any other. Stigliano et al. found a similar pattern in Spanish pre-verbal subject islands with Argentine speakers, noting “considerable variation in the ratings of this structure, indicating substantial divergence in participants’ judgments compared to the other structures tested” (Stigliano et al. 2025: 19).
In contrast, the scores for post-verbal subjects cluster at the high end, resembling those reported for Norwegian adjunct and whether-islands (Kush et al. 2019: Figure 5) and for Spanish whether-islands (Pañeda & Kush 2022: Figure 5), which those authors conclude are not structural islands.
We interpreted this difference to suggest post-verbal subjects do not evince island effects, whereas pre-verbal subjects do, yet the question remains: Why are the ratings spread out in the pre-verbal island violation condition, rather than appearing as uniform rejection, such as that observed by Pañeda et al. (2020)?
Pañeda et al. designed their items to provoke island effects—extracting plain wh-words from specific, pre-verbal agentive subjects—whereas our study, like Stigliano et al.’s, was designed to facilitate extraction when the position licensed it—extracting complex wh-phrases from indefinite, theme/cause subjects. Our pre-verbal island violation sentences pattern more with Stigliano et al.’s than Pañeda et al.’s and share this design property, which we conjecture may explain some of the variation.
Further insight could come from the filler sentences. Our ungrammatical fillers ran the range of acceptability and were ungrammatical for different reasons. Some—with misplaced clitic pronouns—were uniformly rejected, with nearly all ratings at the low end. Others, however—gender agreement mismatches or incorrect complementizers—were unquestionably ungrammatical but nevertheless received middling scores with a wide range of ratings, much like the island violations. Thus, the island violation scores were different from the worst-rated fillers, but also different than grammatical sentences (which clustered at the high end). They pattern with ungrammatical but not starkly unacceptable sentences, like Stigliano et al.’s (2025) results.
Individual differences in the judgment process itself might contribute to the observed variation with ungrammatical sentences. Although there is no widely accepted model of the psychological process of judgment tasks, Schütze (2016) proposes a model in which part of judging an ill-formed string involves generating the nearest possible well-formed string and comparing them, with ill-formed strings that differ more significantly from the well-formed sentence receiving worse ratings. If individuals vary in their ability to imagine a well-formed version of the sentence—especially if that variation includes the degree of attention to different syntactic or semantic properties like definiteness or theta role—such variation could produce a range of ratings like we observe.
Ultimately, the role of individual differences in judgment tasks and how to interpret variation remain important open questions in experimental syntax. Variation and gradience in judgments stem from multiple sources (see Francis 2022 for discussion), and more research aimed at identifying, isolating, and interpreting variation would be valuable.
5.3 Source of island effects
5.3.1 Structural accounts
Taking our data to show a contrast by subject position broadly supports both prominent structural accounts of subject islands in Spanish (Gallego 2011; Haegeman et al. 2014). Because we designed our post-verbal items to avoid violating any of the constraints either approach proposed, they should be equally well-formed under both accounts, so our results do not distinguish between them in this regard.
We can make some distinctions, however. Haegeman et al. claim explicitly, “where the subject is non-specific, extraction is permitted from both pre- and post-verbal subjects” (Haegeman et al. 2014: 107; see also Jiménez-Fernández 2009 for similar claims). Our data does not support this claim: extraction from non-specific subjects in pre-verbal position is degraded, suggesting that position alone is enough to reduce acceptability. Conversely, Gallego’s claims are tied to position: pre-verbal Spec-TP is a freezing position, whereas post-verbal subjects in Spec-v*P permit extraction; our data is consistent with this claim.
Yet we also noted that our pre-verbal island violations, like Stigliano et al.’s, were not as starkly unacceptable as those of Pañeda et al., nor as bad as our worst ungrammatical fillers. A constraint-based account like that suggested by Haegeman et al. could incorporate Gallego’s positional claims in a system that also straightforwardly predicts the gradience we observe. Francis (2022: 236) argues that “compelling evidence is available in support of gradience within the grammar”, and Haegeman et al. suggest a weighted constraint approach like Linear Optimality Theory (Keller 2000) could account for cross-linguistic differences and the gradience observed for subject islands in our study and previous work.
Beyond these two specific two accounts, our findings support the view that what unites languages that do not exhibit subject islands is the ability to extract from post-verbal position, supporting Stepanov’s (2007) generalization. Our findings suggest Spanish can be added to the list of languages permitting extraction from post-verbal subjects, and they strengthen the claim, providing the best evidence we could muster to isolate subject position, suggesting that it is the extraction site that matters.
5.3.2 Functional accounts
We attempted to control the information structure in our experiment to isolate the effect of position. The contexts provided for the pre- and post-verbal versions varied only enough to support the verb’s [±change-of-state] semantics. What was established as focal, salient, and relevant (and thus, what is topical, given, or backgrounded) was held constant across subject positions. If our design succeeded and the information structure was the same for both positions—as suggested by the similar ratings for the non-violation conditions—a discourse-based account of the contrast becomes less plausible.
5.3.3 Processing accounts
We found no effect of working memory on island sensitivity, aligning with previous findings (see §2.3). Our results suggest the cognitive processes involved in making judgments of subject islands in Spanish do not covary with working memory to the extent that we were able to measure it. However, we recognize that merely not finding an effect is not evidence against a processing-based account (see Francis 2022: Chapter 4); we are limited to noting we did not find an effect with this measure.
Aside from the working memory test, could processing explain the contrast by position?
We controlled many possible sources of processing differences, matching items as closely as possible to isolate the effect of position, but position itself could obviously affect processability. Longer dependencies are harder to process—for instance, questions targeting objects tax the processor slightly more than those targeting subjects (Fanselow 2021)—and they reduce acceptability (Goodall 2021). We might thus expect a larger decrease in acceptability for extraction from post-verbal position, in which the gap and the filler are farther apart, which is the opposite of our results. Additionally, overall positional effects are distributed across conditions due to the factorial design (see Sprouse et al. 2016: 314). Finally, an anonymous reviewer points out that our results do not display evidence of increased processing costs with distance, given that the (long-distance) non-island embedded condition received higher ratings than both (short-distance) matrix conditions for both subject positions.
5.4 Limitations and future research
The present study has several limitations.
We chose a between-subjects design rather than a within-subjects design to reduce the probability of satiation effects, yet this choice might have reduced statistical power or introduced sampling error. A future within-subjects study could serve to corroborate our findings. Additionally, a future study including more items would be useful (see Schoenmakers 2025), especially if it explicitly tested satiation.
Our pre- and post-verbal sentences were not perfect minimal pairs, despite our attempt to make them as similar as possible. We introduced two differences between them—the [±change-of-state] context and the verbal aspect to reinforce that reading—with the purpose of making the pre- or post-verbal subject as independently natural as possible, in its canonical and default position. We cannot exclude the possibility that this manipulation produced confounds. In particular, we cannot eliminate the possibility of a pragmatic or semantic difference between the positions despite the contexts.
Although we used a working memory task designed to address criticisms of some previous tasks, it remains possible this task is too simple to measure working memory as it pertains to sentence processing.
Another potential source of limitations is uncontrolled variation in the experiment design. As previously noted, it is possible that certain verbs were perceived as pragmatically odd, that participants applied a reading other than the intended one to some items (see Schütze 2020), or that there is some other factor (like sentence length) we did not control.
We have also highlighted several areas in which future research could help refine experimental syntax methods. More work is needed to establish field-specific guidelines on characterizing and understanding individual variation, as well as interpreting effect sizes and non-monotonic judgment patterns.
6 Conclusions
Our experiment tested subject island effects in Spanish, providing the first comparison of pre- and post-verbal subjects while holding all else as constant as possible. We found super-additive effects in both positions but also a significant contrast between them. Taking the contrast to represent a meaningful difference, we suggested our findings broadly support the description in the theoretical literature claiming that pre-verbal subjects are islands in Spanish, while post-verbal subjects are not, which supports the cross-linguistic generalization that associates island effects with subject position.
Our experiment also addressed debates in experimental syntax. We suggest that multiple sources of converging evidence provide the best case for the presence of an island effect and the source of that effect. A statistically significant super-additive effect remains an important source of evidence, but it is not sufficient on its own. Any significant interaction needs to be contextualized by comparing the effect size to similar previous work, detailing the patterns observed in the data, and examining individual variation. Our work heeds the call of methodologists to engage in detailed analysis of experimental results, moving beyond statistical cutoff points to consider the full range of evidence.
Finally, our experiment addressed debates about the source of island effects. Although our conclusions must necessarily be cautious due to the limitations inherent to any one study, we argued that our findings support a structural view in which post-verbal position is sufficient to ameliorate subject island effects, beyond functional or processing factors.
Abbreviations
acc = accusative, cl = clitic, dat = dative, pl = plural
Data availability/Supplementary files
A supplementary file with full statistical output, as well as materials, anonymous data, and analysis scripts are available via OSF: https://osf.io/dq8mp.
Ethics and consent
This research was conducted in accordance with the ethical principles of the Declaration of Helsinki and approved by the Institutional Review Board (IRB) at DePaul University under approval number IRB-2023-1053. All participants provided informed consent before participation, and their data was anonymized to ensure confidentiality.
Acknowledgements
We are very grateful to Claudia Pañeda, Gregory Scontras, and Bryan Koronkiewicz for comments on an earlier draft of this work. We thank Claudia Pañeda and Irene Finestrat-Martínez for sharing previous experiment materials with us to adapt for this study. Thank you also to three anonymous reviewers, whose feedback helped strengthen the paper significantly, and to the participants in our experiment.
Funding information
This project received support from the DePaul University Research Council, DePaul University College of Liberal Arts and Social Sciences, and the University of Illinois Chicago’s School of Literatures, Cultural Studies and Linguistics.
Competing interests
The authors have no competing interests to declare.
Notes
- See Haegeman et al. for references and more details on each constraint. [^]
- Here we focus only on the empirical generalization that complex wh-phrases improve acceptability in subject islands, but this extends to other islands as well, and the distinct featural composition between bare and complex wh-phrases has played a central role within intervention-based accounts of islands, like featural Relativized Minimality (Villata & Rizzi & Franck 2016; Rizzi 2018). [^]
- See also Goldberg (2013) for a similar approach, as well as Bianchi and Chesi (2014), Chaves and King (2019), Erteschik-Shir and Lappin (1979), and Kuno (1987) for other functional and pragmatic accounts. [^]
- Of course, word order tends to covary with information structure, making it difficult to truly hold all else equal while varying the subject position; we discuss how we controlled information structure and other features in §3.1. [^]
- Keshev and Meltzer-Asscher (2019) suggest a processing account for “subliminal” islands that appeals to memory cue retrieval issues rather than resource limitations, and such an account may be less likely to vary according to individual differences in working memory. They note, however, that their account only applies to islands in which two fillers must be held in memory simultaneously and thus interfere with one another, which is not the case for subject islands. [^]
- Belletti and Rizzi (1988) first classified psych verbs into different types, and Parodi-Lewin (1991) adapted their analysis to Spanish. [^]
- This construction superficially resembles leísmo, the use of le as an accusative clitic. However, Parodi-Lewin (1991) and Halloran Gonzalez (2018) analyze (15) as a true dative, with several properties that distinguish it from the ACC-experiencer version (see also Masullo 1993; Montrul 1996). [^]
- Gutiérrez-Bravo (2006) offers an alternative account, arguing the difference is how agentive the subject is, drawing on Dowty’s (1991) argument proto-roles. One property of Dowty’s proto-agents is “causing an event or change of state in another participant” (Dowty 1991: 572), which aligns with Parodi-Lewin and Halloran Gonzalez’s accounts, and which we manipulated in our experiment. The other four properties of proto-agents were held constant in our experiment: subjects were inanimate, so they had neither volition nor sentience, none of the predicates involved movement, and subjects always existed independently of events. [^]
- See also Rochemont (1986) and Rooth (1992) for similar theories of focus based on presupposition. [^]
- Abeillé et al. (2020) present two tests to identify the discourse status of a constituent, which we attempted to apply to our sentences. One is the “speaking-of” test for whether the subjects are topics (topics are often but not necessarily background). However, according to Fábregas (2016), indefinite quantifiers cannot participate in topicalizing constructions, so we cannot apply this test to our items, which include quantifiers. The second is the sentential negation test to identify the focus (focus is necessarily not part of the background). Only foci can be the target of sentential negation, as shown for English here:
- (i)
- – [The football player liked]background [the color of the car]focus.
- – No, the size of the car.
- (ii)
- – [The football player]background [liked the color of the car]focus.
- – # No, the baseball player.
We obtained informal judgments from four linguists, native speakers of different varieties of Spanish, for this test on a representative sample of our sentences in their contexts. Perhaps because of the complexity of the sentences or the presence of indefinite quantifiers, their judgments were inconsistent: one found only pre-verbal subjects compatible with a focus reading, two rejected a focal reading in both positions, and one found subjects in both positions to be foci. [^]- (iii)
- – [The FOOTBALL PLAYER]focus [liked the color of the car]background.
- – No, the baseball player.
- An anonymous reviewer points out that creer ‘believe’ and pensar ‘think’ could sometimes be perceived as pragmatically odd. For instance, if my colleague saw María cry, it may be odd to say he thinks she was moved by Cortázar, rather than that he knows, since he observed it. We did attempt to identify potential lexically specific problems via native-speaker review and the pilot study, as noted. None of our consultants flagged this problem, but we acknowledge the limitation that we did not consider whether these verbs could be interpreted as insufficiently assertive for some contexts. [^]
- Pañeda (p.c.) pointed out a chance for misinterpretation in the island embedded cases, where a PP like de qué autor could be interpreted as about which author do you think…, that is, as a dependent of the matrix verb rather than extraction from the embedded clause. This is a general challenge also present in previous studies on Spanish. For our study, it is present to the same degree for both subject positions, so any effects of this confound should be distributed equally. [^]
- We thank Irene Finestrat-Martínez for generously sharing details of her implementation of this task. [^]
- We conducted the power analysis by simulation following Lane and Hennes (2018). The crucial decision in any power analysis is identifying the minimum meaningful effect size. However, as discussed in the main text, effect sizes remain an area of substantial debate in experimental syntax, with no widely accepted norm for a meaningful difference on a rating task. To determine the minimum effect we wanted our experiment to detect, we examined several experimental syntax studies of islands (including Aldosari 2015; López-Sancio 2015; Sprouse et al. 2016; Kush & Lohndal & Sprouse 2018; Kush et al. 2018; Pañeda et al. 2020; Pham et al. 2020) to identify typical effect sizes. Of these, Pham et al. (2020) was especially useful because their statistical models were reported in detail. We also noted that previous work on Spanish has found effects similar in magnitude to those found in other languages (Pañeda et al. 2020; 2024), so we could extrapolate from these studies of diverse languages to identify plausible minimum effect sizes. We have also worked with this population and task type before, finding smaller effects than some reported in the literature. Therefore, for the main effects we took the average of the two weakest main effects from Pham et al., rounded down, which are also roughly the magnitude of effects we observed in similar previous studies. The interaction is their smallest interaction rounded down. As a result of these decisions, we assumed regression coefficients of 0.8 for main effects and 0.6 for the Island*Gap interaction in our simulation, which we believe is a very conservative estimate of the minimal effects we wanted to be able to detect. [^]
- We also examined the grammatical fillers; no one gave them mean ratings at or below the scale midpoint. [^]
- Because we z-score transformed the ratings, our DD scores are standardized by participant, but the DD score is not a standardized effect size like Cohen’s d or η², with a group or pooled standardizer. This may limit DD scores’ comparability across studies, because each participant’s ratings are standardized separately, so sampling effects could render them slightly different. That said, even standardized effect sizes like Cohen’s d are often computed with different standardizers across studies (see Lakens 2013), and DD scores after the z-score transformation are expressed in SD units, making them more comparable than DD scores on raw ratings. [^]
- We considered the working memory tests a separate ‘family’ for determining the familywise error rate. See Bender and Lange (2001) for discussion. Adjusting the alpha level for all tests as a single family does not change any test’s outcome. [^]
- Stigliano et al. (2025) compared two designs to operationalize subject islands, which they called the subject/object design (SOD) and the simple/complex subject design (SCSD). [^]
- We calculated DD scores for Abeillé et al. (2020) and Stigliano et al. (2025) using data provided on OSF. For Omaki et al. (2020), Pham et al. (2020), and Sprouse et al. (2011), we estimated DD scores by measuring pixel distances relative to labeled axis marks in their figures. [^]
References
Abeillé, Anne & Hemforth, Barbara & Winckel, Elodie & Gibson, Edward. 2020. Extraction from subjects: Differences in acceptability depend on the discourse function of the construction. Cognition 204. 104293. DOI: http://doi.org/10.1016/j.cognition.2020.104293
Al-Aqarbeh, Rania & Sprouse, Jon. 2024. Island effects and amelioration by resumption in Jordanian Arabic: An auditory acceptability-judgment study. Syntax 27(1). 48–84. DOI: http://doi.org/10.1111/synt.12262
Aldosari, Saad. 2015. The role of individual differences in the acceptability of island violations in native and non-native speakers. University of Kansas dissertation.
Bader, Markus & Häussler, Jana. 2010. Toward a model of grammaticality judgments. Journal of Linguistics 46(2). 273–330. DOI: http://doi.org/10.1017/S0022226709990260
Barr, Dale J. & Levy, Roger & Scheepers, Christoph & Tily, Harry J. 2013. Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language 68(3). 255–278. DOI: http://doi.org/10.1016/j.jml.2012.11.001
Bates, Douglas & Mächler, Martin & Bolker, Ben & Walker, Steve. 2015. Fitting linear mixed-effects models using lme4. Journal of Statistical Software 67(1). 1–48. DOI: http://doi.org/10.18637/jss.v067.i01
Belletti, Adriana & Rizzi, Luigi. 1988. Psych-verbs and θ-theory. Natural Language & Linguistic Theory 6(3). 291–352. DOI: http://doi.org/10.1007/BF00133902
Belova, Daria. 2021. Островные свойства субъектов простой и зависимой клаузы в русском языке [Island properties of subjects in simple and dependent clauses in Russian]. Typology of Morphosyntactic Parameters 4(1). 11–29.
Bender, Ralf & Lange, Stefan. 2001. Adjusting for multiple testing—when and how? Journal of Clinical Epidemiology 54(4). 343–349. DOI: http://doi.org/10.1016/S0895-4356(00)00314-0
Bianchi, Valentina & Chesi, Cristiano. 2014. Subject islands, reconstruction, and the flow of the computation. Linguistic Inquiry 45(4). 525–569.
Biezma, María. 2014. Multiple focus strategies in pro-drop languages: Evidence from ellipsis in Spanish. Syntax 17(2). 91–131. DOI: http://doi.org/10.1111/synt.12014
Broekhuis, Hans. 2011. Extraction from subjects: Some remarks on Chomsky’s “On phases.” In Broekhuis, Hans & Corver, Norbert & Huybregts, Riny & Kleinhenz, Ursula & Koster, Jan (eds.), Organizing grammar: Linguistic studies in honor of Henk van Riemsdijk, 59–68. Berlin: De Gruyter. DOI: http://doi.org/10.1515/9783110892994
Bürkner, Paul-Christian & Vuorre, Matti. 2019. Ordinal regression models in psychology: A tutorial. Advances in Methods and Practices in Psychological Science 2(1). 77–101. DOI: http://doi.org/10.1177/2515245918823199
Cartner, Mandy & Kogan, Matthew & Webster, Nikolas & Wagers, Matthew & Sichel, Ivy. 2026. Subject islands do not reduce to construction-specific discourse function. Cognition 271. 106467. DOI: http://doi.org/10.1016/j.cognition.2026.106467
Casielles-Suárez, Eugenia. 2004. The syntax-information structure interface: Evidence from Spanish and English. New York: Routledge.
Chacón, Dustin A. 2021. Acceptability (and other) experiments for studying comparative syntax. In Goodall, Grant (ed.), The Cambridge handbook of experimental syntax, 181–208. Cambridge University Press. DOI: http://doi.org/10.1017/9781108569620.008
Chaves, Rui P. & Dery, Jeruen E. 2019. Frequency effects in subject islands. Journal of Linguistics 55(3). 475–521. DOI: http://doi.org/10.1017/S0022226718000294
Chaves, Rui P. & King, Adriana. 2019. A usage-based account of subextraction effects. Cognitive Linguistics 30(4). 719–750. DOI: http://doi.org/10.1515/cog-2018-0135
Chomsky, Noam. 2000. Minimalist inquiries: The framework. In Martin, Roger & Michaels, David & Uriagereka, Juan (eds.), Step by step: Essays on Minimalist syntax in honor of Howard Lasnik, 89–156. Cambridge, Mass.: MIT Press.
Chomsky, Noam. 2008. On phases. In Freidin, Robert & Otero, Carlos & Zubizarreta, Maria Luisa (eds.), Foundational issues in linguistic theory: Essays in honor of Jean-Roger Vergnaud, 134–166. Cambridge, Mass.: MIT Press.
Cumming, Geoff. 2012. Understanding the new statistics: Effect sizes, confidence intervals, and meta-analysis. London: Taylor & Francis.
Davies, Mark. 2006. A frequency dictionary of Spanish: Core vocabulary for learners. New York/London: Routledge.
Dowty, David. 1991. Thematic proto-roles and argument selection. Language 67(3). 547–619. DOI: http://doi.org/10.2307/415037
Erteschik-Shir, Nomi & Lappin, Shalom. 1979. Dominance and the functional explanation of island phenomena. Theoretical Linguistics 6(1–3). 41–86. DOI: http://doi.org/10.1515/thli.1979.6.1-3.41
Fábregas, Antonio. 2016. Information structure and its syntactic manifestation in Spanish: Facts and proposals. Borealis – An International Journal of Hispanic Linguistics 5(2). 1–109. DOI: http://doi.org/10.7557/1.5.2.3850
Fanselow, Gisbert. 2021. Acceptability, grammar, and processing. In Goodall, Grant (ed.), The Cambridge handbook of experimental syntax, 118–153. Cambridge University Press. DOI: http://doi.org/10.1017/9781108569620.006
Ferguson, Christopher J. 2009. An effect size primer: A guide for clinicians and researchers. Professional Psychology: Research and Practice 40(5). 532–538. DOI: http://doi.org/10.1037/a0015808
Francis, Elaine. 2022. Gradient acceptability and linguistic theory. New York: Oxford University Press.
Gallego, Ángel J. 2011. Successive cyclicity, phases, and CED effects. Studia Linguistica 65(1). 32–69. DOI: http://doi.org/10.1111/j.1467-9582.2010.01175.x
Gallego, Ángel J. & Uriagereka, Juan. 2007. Sub-extraction from subjects: A phase theory account. In Camacho, José & Flores-Ferrán, Nydia & Sánchez, Liliana & Déprez, Viviane & Cabrera, María José (eds.), Romance linguistics 2006: Selected papers from the 36th Linguistic Symposium on Romance Languages (LSRL), 149–162. Amsterdam: John Benjamins. DOI: http://doi.org/10.1075/cilt.287.12gal
Georgopoulos, Carol. 1991. Syntactic variables: Resumptive pronouns and A’ binding in Palauan. Dordrecht: Springer.
Goldberg, Adele E. 2013. Backgrounded constituents cannot be “extracted.” In Sprouse, Jon & Hornstein, Norbert (eds.), Experimental syntax and island effects, 221–238. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139035309.012
Goodall, Grant. 2011. Syntactic satiation and the inversion effect in English and Spanish wh-questions. Syntax 14(1). 29–47. DOI: http://doi.org/10.1111/j.1467-9612.2010.00148.x
Goodall, Grant. 2021. Sentence acceptability experiments: What, how, and why. In Goodall, Grant (ed.), The Cambridge handbook of experimental syntax, 7–38. Cambridge University Press. DOI: http://doi.org/10.1017/9781108569620.002
Gutiérrez-Bravo, Rodrigo. 2006. A reinterpretation of quirky subjects and related phenomena in Spanish. In Nishida, Chiyo & Montreuil, Jean-Pierre Y. (eds.), New perspectives on Romance linguistics: Vol. I: Morphology, syntax, semantics, and pragmatics. Selected papers from the 35th Linguistic Symposium on Romance Languages (LSRL), 127–142. Amsterdam: John Benjamins. DOI: http://doi.org/10.1075/cilt.275.11gut
Gutiérrez-Bravo, Rodrigo. 2020. La sintaxis del español de México: Un esbozo. Cuadernos de la ALFAL 12(2). 44–70.
Haegeman, Liliane & Jiménez-Fernández, Ángel L. & Radford, Andrew. 2014. Deconstructing the Subject Condition in terms of cumulative constraint violation. The Linguistic Review 31(1). 73–150. DOI: http://doi.org/10.1515/tlr-2013-0022
Haider, Hubert. 1993. Deutsche Syntax, generativ: Vorstudien zur Theorie einer projektiven Grammatik. Tübingen: Narr.
Halloran Gonzalez, Rebecca. 2018. A feature-based approach to the syntax of L2 Spanish psych verbs. Bloomington, Ind.: Indiana University dissertation.
Heidinger, Steffen. 2022. Corpus data and the position of information focus in Spanish. Studies in Hispanic and Lusophone Linguistics 15(1). 67–109. DOI: http://doi.org/10.1515/shll-2022-2056
Hofmeister, Philip & Casasanto, Laura Staum & Sag, Ivan A. 2012a. How do individual cognitive differences relate to acceptability judgments? A reply to Sprouse, Wagers, and Phillips. Language 88(2). 390–400.
Hofmeister, Philip & Casasanto, Laura Staum & Sag, Ivan A. 2012b. Misapplying working-memory tests: A reductio ad absurdum. Language 88(2). 408–409.
Hofmeister, Philip & Sag, Ivan A. 2010. Cognitive constraints and island effects. Language 86(2). 366–415. DOI: http://doi.org/10.1353/lan.0.0223
Holm, Sture. 1979. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 6(2). 65–70.
Hoot, Bradley. 2025. Focus in bilingual Spanish: A state of the science review. Isogloss: Open Journal of Romance Linguistics 11(4). 3 (1–41). DOI: http://doi.org/10.5565/rev/isogloss.509
Hoot, Bradley & Ebert, Shane. 2024. Syntactic island effects in Spanish: Experimental evidence. Natural Language & Linguistic Theory. DOI: http://doi.org/10.1007/s11049-024-09633-5
Huang, C. T. James. 1982. Move WH in a language without WH movement. The Linguistic Review 1(4). 369–416. DOI: http://doi.org/10.1515/tlir.1982.1.4.369
Jiménez-Fernández, Ángel. 2009. On the composite nature of subject islands: A phase-based approach. Finnish Journal of Linguistics 22. 91–138.
Jurka, Johannes. 2010. The importance of being a complement: CED effects revisited. College Park, MD: University of Maryland dissertation.
Juzek, Tom. 2016. Acceptability judgement tasks and grammatical theory. Oxford: University of Oxford dissertation.
Juzek, Tom & Häussler, Jana. 2015. Non-cooperative behaviour & acceptability judgement tasks. Paper presented at the Methods and Linguistic Theories Symposium, Bamberg, Germany.
Keller, Frank. 2000. Gradience in grammar: Experimental and computational aspects of degrees of grammaticality. University of Edinburgh dissertation. DOI: http://doi.org/10.7282/T3GQ6WMS
Keshev, Maayan & Meltzer-Asscher, Aya. 2019. A processing-based account of subliminal wh-island effects. Natural Language & Linguistic Theory 37(2). 621–657. DOI: http://doi.org/10.1007/s11049-018-9416-1
Kim, Boyoung & Goodall, Grant. 2016. Islands and non-islands in native and heritage Korean. Frontiers in Psychology 7.
King, Tracy Holloway. 1994. VP-internal subjects in Russian. In Avrutin, Sergey & Franks, Steven & Progovac, Ljiljana (eds.), Annual workshop on Formal Approaches to Slavic Linguistics: The MIT meeting 1993, 216–234. Ann Arbor: Michigan Slavic Publications.
Kluender, Robert & Kutas, Marta. 1993. Subjacency as a processing phenomenon. Language and Cognitive Processes 8(4). 573–633. DOI: http://doi.org/10.1080/01690969308407588
Kobzeva, Anastasia & Sant, Charlotte & Robbins, Parker T. & Vos, Myrte & Lohndal, Terje & Kush, Dave. 2022. Comparing island effects for different dependency types in Norwegian. Languages 7(3). 197. DOI: http://doi.org/10.3390/languages7030197
Kuno, Susumu. 1987. Functional syntax: Anaphora, discourse, and empathy. Chicago: University of Chicago Press.
Kush, Dave & Lohndal, Terje & Sprouse, Jon. 2018. Investigating variation in island effects: A case study of Norwegian wh-extraction. Natural Language & Linguistic Theory 36(3). 743–779. DOI: http://doi.org/10.1007/s11049-017-9390-z
Kush, Dave & Lohndal, Terje & Sprouse, Jon. 2019. On the island sensitivity of topicalization in Norwegian: An experimental investigation. Language 95(3). 393–420. DOI: http://doi.org/10.1353/lan.2019.0051
Lakens, Daniël. 2013. Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology 4. 863. DOI: http://doi.org/10.3389/fpsyg.2013.00863
Lane, Sean P. & Hennes, Erin P. 2018. Power struggles: Estimating sample size for multilevel relationships research. Journal of Social and Personal Relationships 35(1). 7–31. DOI: http://doi.org/10.1177/0265407517710342
Langsford, Steven & Perfors, Amy & Hendrickson, Andrew T. & Kennedy, Lauren A. & Navarro, Danielle J. 2018. Quantifying sentence acceptability measures: Reliability, bias, and variability. Glossa: A Journal of General Linguistics 3(1). DOI: http://doi.org/10.5334/gjgl.396
Larson-Hall, Jenifer & Plonsky, Luke. 2015. Reporting and interpreting quantitative research findings: What gets reported and recommendations for the field. Language Learning 65. 127–159. DOI: http://doi.org/10.1111/lang.12115
Lenth, Russell V. & Piaskowski, Julia. 2026. emmeans: Estimated marginal means, aka least-squares means (Version 2.0.2) [R package]. https://rvlenth.github.io/emmeans/
Liddell, Torrin M. & Kruschke, John K. 2018. Analyzing ordinal data with metric models: What could possibly go wrong? Journal of Experimental Social Psychology 79. 328–348. DOI: http://doi.org/10.1016/j.jesp.2018.08.009
Liu, Yingtong & Winckel, Elodie & Abeillé, Anne & Hemforth, Barbara & Gibson, Edward. 2022. Structural, functional, and processing perspectives on linguistic island effects. Annual Review of Linguistics 8(1). 495–525. DOI: http://doi.org/10.1146/annurev-linguistics-011619-030319
López-Sancio, Sergio. 2015. Testing syntactic islands in Spanish. Vitoria-Gasteiz: Universidad del País Vasco/Euskal Herriko Unibertsitatea MA thesis.
Lu, Jiayi & Frank, Michael & Degen, Judith. 2024. A meta-analysis of syntactic satiation in extraction from islands. Glossa Psycholinguistics 3(1). DOI: http://doi.org/10.5070/G60111425
Masullo, Pascual José. 1993. Two types of quirky subjects: Spanish versus Icelandic. In Schafer, Amy J. (ed.), Proceedings of the North East Linguistic Society 23, Vol. 2, 303–317. Amherst, Mass.: Graduate Linguistic Student Association, U. Mass. Amherst.
Matchin, William & Almeida, Diogo & Hickok, Gregory & Sprouse, Jon. 2024. An fMRI study of phrase structure and subject island violations. DOI: http://doi.org/10.1101/2024.05.05.592579
Maxwell, Scott E. & Delaney, Harold D. 2004. Designing experiments and analyzing data: A model comparison perspective Second edition. Mahwah, New Jersey: Lawrence Erlbaum.
Michel, Daniel. 2014. Individual cognitive measures and working memory accounts of syntactic island phenomena. San Diego, Calif.: University of California, San Diego dissertation.
Montrul, Silvina. 1996. Clitic-doubled dative “subjects” in Spanish. In Zagona, Karen (ed.), Grammatical theory and Romance languages: Selected papers from the 25th Linguistic Symposium on Romance Languages (LSRL XXV), 183–196. Amsterdam: John Benjamins. DOI: http://doi.org/10.1075/cilt.133.15mon
Newmeyer, Frederick J. 2016. Nonsyntactic explanations of island constraints. Annual Review of Linguistics 2(1). 187–210. DOI: http://doi.org/10.1146/annurev-linguistics-011415-040707
Omaki, Akira & Fukuda, Shin & Nakao, Chizuru & Polinsky, Maria. 2020. Subextraction in Japanese and subject-object symmetry. Natural Language & Linguistic Theory 38(2). 627–669. DOI: http://doi.org/10.1007/s11049-019-09449-8
Pañeda, Claudia & Kush, Dave. 2022. Spanish embedded question island effects revisited: an experimental study. Linguistics 60(2). 463–504. DOI: http://doi.org/10.1515/ling-2020-0110
Pañeda, Claudia & Lago, Sol & Vares, Elena & Veríssimo, João & Felser, Claudia. 2020. Island effects in Spanish comprehension. Glossa: A Journal of General Linguistics 5(1). 21. DOI: http://doi.org/10.5334/gjgl.1058
Pañeda, Claudia & Villata, Sandra & Kush, Dave & Sprouse, Jon. 2024. A translation-matched, experimental comparison of three types of wh-island effects in Spanish and English. Glossa: A Journal of General Linguistics 9(1). DOI: http://doi.org/10.16995/glossa.11164
Parodi-Lewin, Claudia. 1991. Aspect in the syntax of Spanish psych-verbs. Los Angeles: University of California Los Angeles dissertation.
Pesetsky, David. 1996. Zero syntax: Experiencers and cascades. Cambridge, Mass.: MIT Press.
Pham, Catherine & Covey, Lauren & Gabriele, Alison & Aldosari, Saad & Fiorentino, Robert. 2020. Investigating the relationship between individual differences and island sensitivity. Glossa: A Journal of General Linguistics 5(1). 94. DOI: http://doi.org/10.5334/gjgl.1199
Plonsky, Luke. 2015. Statistical power, p values, descriptive statistics, and effect sizes: A “back-to-basics” approach to advancing quantitative methods in L2 research. In Plonsky, Luke (ed.), Advancing quantitative methods in second language research, 23–45. New York: Routledge, Taylor & Francis Group.
Plonsky, Luke & Oswald, Frederick L. 2014. How big is ‘big’? Interpreting effect sizes in L2 research. Language Learning 64. 878–912.
Polinsky, Maria & Gallo, Carlos G. & Graff, Peter & Kravtchenko, Ekaterina & Milton Morgan, Adam & Sturgeon, Anne. 2013. Subject islands are different. In Sprouse, Jon & Hornstein, Norbert (eds.), Experimental syntax and island effects, 286–309. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/CBO9781139035309.015
Prolific Team. 2022. Prolific’s attention and comprehension check policy. (https://researcher-help.prolific.co/hc/en-gb/articles/360009223553-Prolific-s-Attention-and-Comprehension-Check-Policy) (Accessed 2022-5-9)
R Core Team. 2022. R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing. https://www.R-project.org/
Rizzi, Luigi. 2006. On the form of chains: Criterial positions and ECP effects. In Cheng, Lisa Lai Shen & Corver, Norbert (eds.), Wh-movement: Moving on, 97–134. Cambridge, Mass.: MIT Press.
Rizzi, Luigi. 2018. Intervention effects in grammar and language acquisition. Probus 30(2). 339–367. DOI: http://doi.org/10.1515/probus-2018-0006
Rochemont, Michael S. 1986. Focus in generative grammar. Amsterdam: John Benjamins.
Rooth, Mats. 1992. A theory of focus interpretation. Natural Language Semantics 1(1). 75–116.
Schoenmakers, Gert-Jan Thomas. 2025. How a simple increase in the number of items can enhance the reliability of linguistic judgments: The case of island experiments. Languages 10(11). 277. DOI: http://doi.org/10.3390/languages10110277
Schoenmakers, Gert-Jan Thomas & Stoica, Irina. 2024. An experimental investigation of wh-dependencies in four island types in Romanian. Glossa: A Journal of General Linguistics 9(1). DOI: http://doi.org/10.16995/glossa.15193
Schütze, Carson T. 2016. The empirical base of linguistics: Grammaticality judgments and linguistic methodology. Berlin: Language Science Press.
Schütze, Carson T. 2020. Acceptability ratings cannot be taken at face value. In Schindler, Samuel & Drożdżowicz, Anna & Brøcker, Karen (eds.), Linguistic intuitions: Evidence and method, 189–214. Oxford: Oxford University Press. DOI: http://doi.org/10.1093/oso/9780198840558.003.0011
Schütze, Carson T. & Sprouse, Jon. 2013. Judgment data. In Podesva, Robert J. & Sharma, Devyani (eds.), Research methods in linguistics, 27–50. Cambridge: Cambridge University Press.
Schwarzschild, Roger. 1999. Givenness, AvoidF and other constraints on the placement of accent. Natural Language Semantics 7(2). 141–177.
Snyder, William. 2021. Satiation. In Goodall, Grant (ed.), The Cambridge handbook of experimental syntax, 154–180. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/9781108569620.007
Sprouse, Jon & Caponigro, Ivano & Greco, Ciro & Cecchetto, Carlo. 2016. Experimental syntax and the variation of island effects in English and Italian. Natural Language & Linguistic Theory 34(1). 307–344. DOI: http://doi.org/10.1007/s11049-015-9286-8
Sprouse, Jon & Fukuda, Shin & Ono, Hajime & Kluender, Robert. 2011. Reverse island effects and the backward search for a licensor in multiple wh-questions. Syntax 14(2). 179–203. DOI: http://doi.org/10.1111/j.1467-9612.2011.00153.x
Sprouse, Jon & Villata, Sandra. 2021. Island effects. In Goodall, Grant (ed.), The Cambridge handbook of experimental syntax, 227–257. Cambridge: Cambridge University Press. DOI: http://doi.org/10.1017/9781108569620.010
Sprouse, Jon & Wagers, Matt & Phillips, Colin. 2012. A test of the relation between working-memory capacity and syntactic island effects. Language 88(1). 82–123. DOI: http://doi.org/10.1353/lan.2012.0004
Starke, Michal. 2001. Move reduces to merge: A theory of locality. LingBuzz. https://ling.auf.net/lingbuzz/000002
Stepanov, Arthur. 2007. The end of CED? Minimalism and extraction domains. Syntax 10(1). 80–126. DOI: http://doi.org/10.1111/j.1467-9612.2007.00094.x
Stepanov, Arthur. 2008. Ergativity, Case and the Minimal Link Condition. In Stepanov, Arthur & Fanselow, Gisbert & Vogel, Ralf (eds.), Minimality effects in syntax, 367–400. De Gruyter Mouton.
Stepanov, Arthur & Mušič, Manca & Stateva, Penka. 2018. Two (non-)islands in Slovenian: A study in experimental syntax. Linguistics 56(3). 435–476. DOI: http://doi.org/10.1515/ling-2018-0002
Stigliano, Laura & Verdecchia, Matias & Murujosa, Marisol. 2025. Comparing two experimental designs for the study of subject islands in Spanish. Isogloss: Open Journal of Romance Linguistics 11(5). 1–28. DOI: http://doi.org/10.5565/rev/isogloss.510
Surányi, Balázs & Turi, Gergő. 2018. Freezing, topic opacity and phase-based cyclicity in subject islands: Evidence from Hungarian. In Hartmann, Jutta & Jäger, Marion & Kehl, Andreas & Konietzko, Andreas & Winkler, Susanne (eds.), Freezing: Theoretical approaches and empirical domains, 317–350. Berlin: De Gruyter Mouton.
Szabolcsi, Anna & Lohndal, Terje. 2017. Strong vs. weak islands. In Everaert, Martin & van Riemsdijk, Henk C. (eds.), The Wiley Blackwell companion to syntax, 2nd ed., 1–51. Hoboken, NJ: Wiley. DOI: http://doi.org/10.1002/9781118358733.wbsyncom008
Uriagereka, Juan. 1988. On government. Storrs, Conn.: University of Connecticut dissertation.
Villata, Sandra & Rizzi, Luigi & Franck, Julie. 2016. Intervention effects and Relativized Minimality: New experimental evidence from graded judgments. Lingua 179. 76–96. DOI: http://doi.org/10.1016/j.lingua.2016.03.004
Wechsler, David. 1997. Wechsler memory scale, 3rd ed. San Antonio, Tex.: Psychological Corporation.
Zehr, Jérémy & Schwarz, Florian. 2018. PennController for Internet Based Experiments (IBEX). DOI: http://doi.org/10.17605/OSF.IO/MD832





