1 Introduction

Establishing unequivocal neural evidence for latent, abstract mental speech sound representations is a formidable challenge (Gershman & Niv 2010; Monahan 2018), largely because phonological properties often correlate with auditory cues. That is, there are generally lawful relationships between auditory cues and the resulting induced phonological representations. This is especially important for the substance-free phonology (SFP) program, whose principal aim is to identify how much abstraction from phonetic detail is present in human phonological representations and computations (Reiss 2017; Chabot 2024). Beyond the phonology literature, debates persist regarding the fundamental perceptual units of speech (Kazanina et al. 2018; Samuel 2020), and formulating plausible linking hypotheses between language and brain function is far from a trivial task (Embick & Poeppel 2015; Poeppel & Embick 2017; Poeppel & Idsardi 2022). Nonetheless, phonology may be the most tractable domain for establishing such linking hypotheses with cognitive neuroscience, given its close ties to the sensory-motor system and our detailed understanding of the physical properties of speech, the mammalian ascending auditory pathway, and auditory cortex (Poeppel et al. 2008).

Even invasive electrocorticographic (ECoG) research, which highlights the necessity of labels for characterizing auditory cortical encoding (Mai et al. 2024), falls short of providing clear evidence for abstract phonological representations. Somewhat unexpectedly, non-invasive methods may offer better insights into such abstractions. Mismatch negativity (MMN) paradigms (Näätänen et al. 2019) suggest that the brain encodes phonetic features (Phillips et al. 2000; Kazanina et al. 2006; Hestvik & Durvasula 2016) and can extract latent, abstract, and consistent features from category-varying standards (Fu & Monahan 2021). The strongest support for abstract features comes from MMN findings showing that the brain can integrate disparate auditory cues into a unified phonological percept. To date, the clearest example involves English obstruent “voicing”, where a temporal cue, that is, plosive voice onset time, is combined with a spectral cue, that is, fricative low-frequency periodicity (Monahan et al. 2022; Politzer-Ahles & Jap 2024). Still, the number of test cases is limited, and it remains unclear whether such cue integration across manner classes reflects a general mechanism for perceiving unified, abstract, latent features.

The problem is not intractable; however, the remaining challenges are considerable. To answer whether features are phonetically grounded or substance-free with neurobiological measures requires a) methodologies that permit precision sampling of the temporal and spatial neurophysiological dynamics in tandem, b) an established set of paradigms that faithfully tap into specific levels of linguistic representations (Idsardi & Poeppel 2011; Monahan 2018) and c) established linking hypotheses between the brain and linguistic representation (Embick & Poeppel 2015; Poeppel & Embick 2017). It is important to remember that imaging methods are nascent (<100 years old), and extant techniques that do offer simultaneous temporal and spatial precision require surgery. Furthermore, many paradigms fail to differentiate whether phonetic or phonological representations are accessed, rendering underdetermined whether distinctive features are grounded or substance-free. These methodological and paradigm issues motivate our extensive methodological review. Finally, the absence of established linking hypotheses often renders neural evidence difficult to interpret. Putatively temporal differences are sometimes observed as spatial contrasts and vice versa (Howard & Poeppel 2009; Fox et al. 2020). As an example, linking hypotheses for how distinctive features are neurally encoded are largely unknown. Vowel height has been shown to be both spatially (Obleser et al. 2006; Scharinger et al. 2011; Mesgarani et al. 2014; Manca et al. 2019; Oganian et al. 2023) and temporally encoded (Roberts et al. 2004; Monahan & Idsardi 2010), complicating monothetic interpretations and implying a spatial-temporal coding of unknown character. Table 1 presents an assessment of the current state-of-knowledge and outstanding challenges.

Table 1: Current knowledge and challenges in our understanding of phonological abstraction and substance-free phonology in the context of the neurobiology.

Five things we know:
  1. Ascending auditory pathway performs increasingly abstract analyses of sound signals generally. This is also true in phonology, entailing at least some substance-free representations and computations.

  2. Predictive brain mechanisms allow for change detection experiments, providing a paradigm that allows for testing different levels of linguistic representation.

  3. Phonetic categories are spatially and temporally coded in superior regions of temporal cortex. More abstract phonological categories appear to localize to areas posterior to phonetic category areas.

  4. Naturalistic, large datasets can be built and tested using forced alignment between the phonetic signal and temporal brain responses.

  5. Brain disjunctively combines distinct phonetic cues into unitary abstract phonological features.

Five significant challenges:
  1. No adequate linking hypotheses exist between ontological primitives in linguistics and neuroscience.

  2. Evaluative criteria for assessing linking hypotheses are also lacking. Progress in language lacks considerably behind other sensory domains (e.g., olfaction).

  3. Need novel methods with enhanced spatial and temporal precision, as well as paradigms that can selectively tap distinct levels of speech sound representations.

  4. Need comprehensive multi-language datasets with shared experimental designs to facilitate meta-analyses to reveal language-particular abstract entities.

  5. Clear statements of the relationship between auditory observables and abstract entities in substance-free phonology proposals need to be proposed and refined. This will facilitate the construction of linking hypotheses between phonology and the brain and allow for testing more sophisticated questions germane to substance-free phonology.

Here, we update our earlier overview of neurological, brain imaging, and neurophysiological methodologies (Idsardi & Poeppel 2011; Monahan et al. 2013). We then review empirical findings from the past decade that suggest that the brain does support abstract phonological features, a conclusion integral for the suitability of substance-free phonology approaches. Next, we outline key unresolved challenges and conclude by identifying promising directions for future research. The current review is limited to the auditory channel, leaving aside questions of the interface between the motor-system and phonology.

2 Methodologies

The development of both descriptive and generative phonological grammars has traditionally relied on acceptability judgments and fieldwork techniques as primary modes of data collection. In contrast, behavioural psycholinguistics has introduced a wide array of experimental tasks designed to probe the cognitive mechanisms underlying speech perception (Grosjean & Frauenfelder 1996). Common dependent measures for behavioral tasks include reaction times, accuracy rates, and standardized discriminability scores (e.g., d’; see Macmillan & Creelman 2004 for a review). Cognitive neuroscience methods remained generally underutilized in phonological research; as such, we previously outlined key technologies relevant to the field (Monahan et al. 2013). As emphasized there, no single methodology is inherently superior; each has unique advantages and limitations. The choice of method is to be guided by the specific research question. Here, we present an updated overview of common, available methodologies.

From the initial diagnoses of Broca and Wernicke aphasias (Broca 1865; Wernicke 1874) until the final quarter of the twentieth century, patient research was primary to understanding the brain basis of language. Internal vascular hemorrhages, traumatic brain injuries or slow, progressive degenerative brain diseases that affect specific left hemisphere regions can result in communication deficits (Dronkers & Baldo 2009). A range of assessments and tasks are used to assess the type and degree of linguistic deficit experienced by the individual (Salter et al. 2006). Similar performance on tasks that are assumed to be distinct suggests that the same brain region may be involved in both, even before specific regions are identified. In contrast, performance differences between two tasks indicate that each task likely engages separate brain areas.

Careful patient work resulted in detailed neuroanatomical models of the language network, including a meticulous classification of aphasia types and the concomitant linguistic deficits with specific lesion sites (Lichtheim 1885; Geschwind 1970). As we noted in 2013, a key advantage of neuropsychological studies is their ability to support limited causal inferences. In contrast, the absence of activity in a brain region observed through modern neuroimaging may simply reflect the limited sensitivity of the technique, rather than a lack of involvement of a particular brain region; conversely, activity observed through modern neuroimaging may merely reflect neurophysiological processes that correlate with the cognitive function of interest.

Before the advent of modern neuroimaging techniques, affected brain regions could only be identified through post-mortem autopsy. Now, structural magnetic resonance imaging (MRI) and computed tomography (CT) permit lesion localization in vivo. Voxel-based lesion-mapping allows for the establishment of a relationship between tissue damage and behaviour (Bates et al. 2003; Dronkers et al. 2004; Karnath et al. 2018), and convolutional neural networks segment and classify anatomical regions and lesion sites (Bernal et al. 2019). This approach, however, rests on the assumption of a relatively modular anatomical architecture, despite evidence that cognitive functions are supported by distributed neural networks (Hickok & Poeppel 2007; Lau et al. 2008; Binder et al. 2009; Hickok 2009; Scott 2019). Additionally, brain lesions are frequently diffuse and not confined to discrete functional regions, which themselves exhibit considerable inter-individual anatomical variability (Gentry et al. 1988). Finally, perilesional tissue is often subject to functional reorganization, further obscuring the relationship between localized damage and cognitive outcomes (Rorden & Karnath 2004). Furthermore, effective use of the lesion method requires a detailed understanding of the knowledge and processes recruited by patients to execute the task: Misunderstandings of the task have led to erroneous conclusions about the neural substrates of phonological processing (Hickok & Poeppel 2004). Limited sample sizes, combined with the heterogeneity of brain lesions and the variability in patient deficits, pose significant challenges to drawing generalizable conclusions about the broader population.

Over the past five decades, there has been remarkable progress in the real-time imaging and measurement of the human brain in healthy individuals. The late 1960s and early 1970s marked the advent of key neuroimaging technologies, including positron emission tomography (PET), magnetic resonance imaging (MRI), and magnetoencephalography (MEG). These innovations significantly broadened both the scope of testable populations and range of research questions that can be empirically addressed. Inquiries into the spatial and temporal dynamics of neurocognitive functioning can now be examined in vivo, across diverse linguistic and demographic groups. Contemporary methodologies in cognitive neuroscience are generally categorized by what they measure: blood flow (hemodynamic) or electrical impulses (electromagnetic). This classification reflects the distinct biological processes each technique measures. Hemodynamic approaches (e.g., functional MRI, some forms of PET) assess changes in blood flow (tracking increased oxygen or glucose) throughout the brain, providing insights into where cognitive processes localize in the brain. In contrast, electromagnetic techniques (e.g., MEG, ECoG, EEG) capture fluctuations in the brain’s electromagnetic fields, offering a measure of when a cognitive process occurs. Hemodynamic methods typically offer superior spatial resolution, enabling precise localization of brain regions engaged during specific cognitive tasks. Electromagnetic methods, by comparison, excel in temporal resolution, allowing researchers to track the timing of neural processes with millisecond precision. That said, some electromagnetic techniques offer superb (e.g., ECoG) or very good spatial resolution (e.g., MEG), and advances in sampling the hemodynamic response have improved the temporal resolution of fMRI (Yang & Lewis 2021). Despite such advances, researchers still need to assess whether their research questions can be answered in a primarily spatial or primarily temporal fashion. Or, if they have access to a limited variety of methods, to try to align their research questions with the available spatial and temporal resolutions of their experimental equipment.

2.1 Electromagnetic methodologies

As noted above, electromagnetic techniques measure changes in electrical voltage or their associated magnetic fields. In 1925–1926, Hans Berger, the German psychiatrist, developed the first electroencephalogram, measured faint cortical oscillations using clay electrodes placed on the scalp, connected to a string galvanometer on photographic paper (Berger 1929; Millett 2001). This was the first time that brain function was measured in real-time from outside the human head, unlocking the potential for the technological advances witnessed throughout the rest of the twentieth century.

Electroencephalography (EEG) is a safe, non-invasive, comparatively inexpensive method that measures changes in electrical voltages from the scalp. Protocols typically require placing an elastic cap that holds electrodes on the participant’s head, although alternative EEG configurations exist (e.g., geodesic EEG systems). These electrodes are often made of a metal (e.g., tin, platinum, silver-silver chloride), whose physical connection between to the scalp is facilitated by the application of an electrolyte gel or aqueous solution. EEG does not measure momentary or extremely local neural events (e.g., action potentials). Instead, the biological origin lies in the secondary or volume currents generated by post-synaptic potentials of pyramidal neurons, primarily located in the grey matter. The synchronized activity of several tens of thousands of neurons generates a post-synaptic potential large enough to be observable from outside the scalp. The voltage difference between each scalp electrode and a reference electrode (e.g., mastoid, earlobe) is recorded. The signal-to-noise ratio, like most cognitive neuroscience methods, is poor, requiring the presentation of dozens to hundreds of exogenous stimuli (e.g., auditory, visual) to obtain robust and interpretable EEG responses. In the cognitive neuroscience of language literature, the most common analysis involves averaging the continuous EEG time-locked to the onset of the stimulus, obtaining an evoked or event-related potential (ERP). The ERP is the weighted sum of underlying electrical sources of brain activity, with different spatial and temporal dynamics. As such, directly inferring underlying brain activity from scalp-level responses is not straightforward (Luck 2014). Peaks that recur in the ERP waveform under similar testing conditions are identified, labeled and routinely assigned functional significance (e.g., the N400 reflects semantic integration (Lau et al. 2008)). The amplitude and latency of these ERP components are common dependent variables that are compared across experimental conditions. Only neurophysiological activity phase locked to stimulus onset survives in an ERP analysis procedure. Unaligned information is then attenuated by the averaging process. An alternative approach involves calculating the oscillatory power of the EEG signal. When neural ensembles synchronize their firing patterns, they generate large-scale oscillatory activity which can be detected at the scalp. These oscillations often arise from feedback interactions among neurons and may reflect ensemble-level dynamics that are not evident at the level of individual neurons. To estimate oscillatory power, a time-frequency decomposition (i.e. a neural spectrogram) is performed on each trial and then averaged across trials within each experimental condition. Unlike ERP analyses, this method preserves induced activity—EEG signals that are not phase-locked to stimulus onset—providing a richer view of neural dynamics. Power at different frequencies is labeled (i.e., δ-band: 1–4 Hz; θ-band: 4–7 Hz; α-band: 8–12 Hz; β-band: 13–30 Hz; γ-band: 30–70 Hz) and routinely ascribed functional significance (e.g. θ-band tends to track the speech amplitude envelope). More advanced analysis methods of induced EEG activity include inter-trial phase coherence, inter-channel phase synchrony (Morales & Bowers 2022) and various connectivity analyses in the frequency domain (Chiarion et al. 2023). More recent analytic techniques include fitting temporal response functions to EEG (Gillis et al. 2021; Lindboom et al. 2023) and MEG (Brodbeck et al. 2018; Gwilliams et al. 2018; 2022; 2024) data while participants listen to continuous speech.

EEG is widely used for good reason: It is non-invasive, relatively inexpensive, and boasts superb temporal precision, an important feature for measuring the complex, fast dynamics of human speech and language. EEG analysis methods are generally less computationally intensive than those used in fMRI and MEG and tend to be more standardized than those used in MEG. This, along with the lower cost, lowers the barrier to entry for new researchers. With adequate statistical power, it is possible to detect temporal differences on the order of a few tens of milliseconds. On the downside, the application of electrolyte gel to the scalp is time- and labor-intensive; however, advancements in active electrode and dry-electrode technologies have significantly reduced preparation time. Moreover, several commercially available, research-grade EEG systems are now portable, enabling data collection outside the laboratory and facilitating research with populations that are otherwise difficult to access, being able to record without needing magnetically shielded enclosures makes EEG the only current method viable for use in the field. A well-documented limitation of EEG is its poor spatial resolution, which restricts the ability to reliably localize changes in neural activity. The conductive properties of the skull and surrounding tissues (e.g., cerebral spinal fluid, scalp) distort the signals generated by the brain. High-density electrode arrays can enhance spatial resolution (Michel & Brunet 2019); however, when the research question centers on the precise localization of cognitive processes within the brain, other neuroimaging modalities offer superior spatial resolution.

Electrocorticography (ECoG) is a type of intracranial electroencephalography. Electrode grids are placed directly onto the cortical surface after a craniotomy. ECoG is highly invasive and limited to presurgical patients only, typically individuals being treated for severe epilepsy symptoms which have resisted other treatments. ECoG dates to the 1950s, when Wilder Penfield and Herbert Jasper developed the surgical technique to localize the sources of epileptic activity for later surgical removal. Like EEG, the physiological basis of the ECoG signal is synchronized post-synaptic potentials in pyramidal neurons located in the outer layers of cortex. As such, ECoG offers an extremely high temporal resolution. Unlike EEG, however, ECoG better captures higher frequency oscillatory activity (e.g., high γ-band: >70 Hz) and offers excellent spatial information due to the electrode grids being placed directly on the surface of the brain, down to 1–100 µm in resolution (Fallegger et al. 2021). There are two principal disadvantages to highlight. First, due to the invasive nature of the technique, the available population is limited and neurologically divergent, and experiments are typically conducted on small samples (e.g., <10 participants). Second, while ECoG offers superior spatial resolution to EEG, this resolution is largely limited to cortical regions whereupon the electrode grid has been placed (Dubey & Ray 2019), which depends on each patient’s neurological diagnosis; EEG, on the other hand, measures activity from most of cortex. In addition to these disadvantages, it is also important to note the myriad of ethical considerations involved in conducting ECoG for research purposes, given the sensitive nature of the population and methodology (Chiong et al. 2018).

The physiological bases of EEG and magnetoencephalography (MEG) are closely intertwined. EEG captures voltage changes resulting from secondary currents, whereas MEG detects magnetic fields produced by primary currents generated by post-synaptic potentials in pyramidal neurons, predominantly located in the superficial layers of the cortex. All electrical currents produce magnetic fields that rotate perpendicularly to the direction of current flow, following the right-hand rule, also known as Oersted’s or Ampère’s Law. When the primary current flows tangentially to the scalp, such as within a cortical sulcus, the resulting magnetic field can be measured outside the head using superconducting quantum interference devices (SQUIDs), which operate within a cryogenic environment maintained by liquid helium. Radially oriented dipoles, those originating in cortical gyri, produce magnetic fields that MEG cannot detect. Consequently, MEG is considered blind to radial dipoles and is sensitive only to tangentially oriented dipoles. In contrast, EEG can detect dipoles regardless of their orientation; however, EEG is more vulnerable to signal cancellation when sources with opposing orientations are simultaneously active, resulting in a net zero signal at the scalp. Similar experimental paradigms are often employed in MEG and EEG studies. The evoked response field (ERF) is the magnetic equivalent of the ERP, and MEG components are often denoted with an “m” (e.g., M100, instead of the N1; MMNm instead of the MMN). In addition to the range of evoked and induced analysis techniques that are possible with both EEG and MEG signals, MEG also allows for more accurate estimations of the spatial characteristics of the neurophysiological activity. The principal advantage of MEG relative to EEG is that these magnetic fields are impervious to the intervening biological tissue between the electrical activity in the brain and the sensors. Two broad classes of spatial estimation techniques, single-source estimations (e.g., equivalent current dipole (ECD)) and distributed source estimations (e.g., minimum-norm estimates (MNE)).

Like EEG, MEG is silent, safe and non-invasive. Compared to EEG, MEG participant preparation is typically faster and more comfortable. As noted above, the SQUID sensors must be thermally insulated in a liquid helium-filled dewar to take on superconducting qualities. This means that the sensors are not directly attached to the participant’s scalp, as they are in EEG. Thus, tracking the participant’s head movement throughout a testing session is necessary. Although source localization in MEG is generally more accurate than in EEG, it remains a computationally ill-posed problem with too many possible solutions. Current methods are still poorly understood in scenarios involving multiple simultaneously active sources. Accurate localization also requires a separate structural MRI, adding both cost and complexity to MEG studies. The primary drawback of MEG is the start-up and operating expenses. A new, complete MEG system can cost between $4,000,000 and $6,000,000, multiple orders of magnitude more expensive than a new EEG system. A considerable part of this cost is the magnetically shielded room. SQUID sensors are extremely sensitive to environmental magnetic noise, and as such, require shielding from electromagnetic sources that are present in the surrounding environment. Moreover, liquid helium evaporates and must be replenished. In the past, this entailed the weekly delivery of liquid helium to the laboratory, an additional expense of approximately $200,000 per year at $40/Liter. Natural helium, however, is extraordinarily scarce in the environment and subject to persistent supply-chain disruptions (Anderson 2018). In response, there has been development of helium-recycling and recovery systems that reclaim the evaporated helium gas, as well as optically-pumped magnetometers (OPM-MEG), that do not rely on cryogenics and offer other advantages (Brookes et al. 2022) but still require magnetic shielding. We anticipate increased use of OPM-MEG systems as the technology matures, although general adoption is probably a decade or more away.

While EEG and MEG passively measure electromagnetic activity in the brain, other non-invasive electrophysiological methods actively modulate neural activity. Transcranial magnetic stimulation (TMS) uses magnetic fields, while transcranial direct current stimulation (tDCS) applies weak electrical currents. These techniques use coils or electrodes, respectively, to induce electrical currents in the brain. These methods are used clinically to treat depression and other psychiatric disorders (Gershon et al. 2003). These technologies have not been as widely applied to questions of phonological representations; instead, they are typically used to assess the role of motor cortex in speech perception (D’Ausilio et al. 2009; Möttönen & Watkins 2012; Murakami et al. 2013; Adank et al. 2017).

2.2 Hemodynamic methodologies

The foundations of modern brain imaging date to the late nineteenth century. In the 1880s, Italian physiologist Angelo Mosso developed the human circulation balance, a device that would allow him to show that cognitively demanding tasks increased cerebral blood flow (Mosso 2014; Sandrone et al. 2014). By 1890, it was established that brain activity was linked to localized changes in cerebral blood flow (Roy & Sherrington 1890). Positron emission tomography (PET) scanners were developed in the 1970s and measure metabolic activity in the brain via the injection of radioactive tracers into the bloodstream. While PET was employed in some of the earliest imaging tests of speech perception (Petersen et al. 1989; Sergent et al. 1992; Zatorre et al. 1992; see Poeppel 1996 for a critique), the relative ubiquity of MRI machines and moderate invasive nature of PET caused it fall out of favour. Although early water-cooled PET systems did have a distinct advantage in that the acquisition of the signal was nearly silent (Talavage et al. 2014), modern systems are air-cooled and incorporate computerized tomography (CT) equipment, making the background noise much louder (Speck et al. 2021).

fMRI has become the leading method for investigating where cognitive functions occur in the brain. Different tissues and biological substances have distinct magnetic properties, which magnetic resonance imaging (MRI) can detect to produce detailed images of brain anatomy. The most common dependent measure in fMRI is the blood-oxygen-level-dependent (BOLD) signal. When a brain region becomes metabolically active, it undergoes changes in blood oxygenation. Specifically, there is an increase in oxygen-rich blood (oxyhemoglobin) and a corresponding decrease in oxygen-poor blood (deoxyhemoglobin). This reduction in deoxyhemoglobin decreases magnetic field distortions, resulting in an increase in the BOLD signal. Unlike PET, which requires the injection of radioactive tracers that circulate through the brain, fMRI is entirely non-invasive and considered safe for human use, at least within typical magnetic field strengths. Extremely high field strengths (>3T) are rare in human studies.

Unlike EEG and MEG, fMRI provides a direct and spatially precise measure of magnetic susceptibility across the brain. EEG and MEG rely on models for spatial localization, limiting their accuracy. Despite fMRI’s superior spatial resolution, it has notable drawbacks. First, the BOLD signal reflects blood flow, which is slow, peaking 3–6 seconds after stimulus onset. This limits fMRI’s ability to capture fast, dynamic processes like language comprehension. Additionally, the link between BOLD signals and neural activity is complex and not fully understood. Factors such as nearby blood vessels and the balance of excitatory versus inhibitory activity can complicate interpretation (Sotero & Trujillo-Barreto 2007; Aksenov et al. 2019; Moon et al. 2021). Language studies face another challenge: the shifting magnetic gradients that are required for imaging create loud noise that can interfere with auditory stimuli delivery. Practically, fMRI requires expensive equipment and strict safety protocols, so scanners are typically housed in large institutions, and scanning costs can be high (~$500/hour). Finally, fMRI data analysis is computationally demanding and less intuitive than ERP analysis, but it is more standardized and better documented than MEG analysis (Poldrack et al. 2011).

The final method is functional near-infrared spectroscopy (fNIRS), a hemodynamic technique that uses infrared light to detect changes in blood oxygenation. It shares a similar biophysical basis with fMRI, and its signal has been shown to correlate with the BOLD response (Buxton et al. 1998; Wijeakumar et al. 2017). While fNIRS event-related designs with adults is relatively rare (Defenderfer et al. 2017), it is quite suitable for work in the developmental sciences (Wilcox & Biondi 2015) and has been shown to be an effective method to study language-specific phonological processing in infants potentially due to their relatively thinner skulls (Minagawa-Kawai et al. 2007).

3 Testing theories

Testing linguistic theory using behavioral and cognitive neuroscience methods presents significant challenges. Often, linguistic theories fail to yield hypotheses that are directly testable with these paradigms. In other cases, the necessary methodological and analytical precision is lacking. While neuroimaging has seen remarkable innovation over the past several decades, and computational models of auditory functions have become increasingly sophisticated (e.g., Grothe 2003), theoretical advances in phonology have largely developed in parallel: Phonological theory and cognitive neuroscience are pursued almost entirely exclusively of one another. Meaningful interdisciplinary integration requires the formulation of linking hypotheses, grounded in clearly defined ontological primitives of both phonological theory and neuronal function (Embick & Poeppel 2015; Poeppel & Embick 2017).

As noted in the Introduction, there are no linking hypotheses at present to sketch. Progress requires establishing the computational theory, the relevant representations and algorithms, as well as the hardware implementation (Marr 1982). We can look toward models of the olfactory system wherein both the neural circuitry and ontological sensory primitives are defined and linked (Murthy 2011; Giessel & Datta 2014); alternatively, closer to the current domain, the neural coding of interaural time delays in mammalian and aviary auditory systems to solve clearly defined computational theory (Grothe 2003). Our current understanding is that phonetic categories are both spatially and temporally coded, providing headway into questions of hardware implementation and potentially, the representations, which as we outline below, are currently better assessed using high-temporal, non-invasive methodologies.

3.1 Spatial localization

The neurophysiological architecture of speech processing is bilateral and distributed, integrating sensory and motor systems with core linguistic computation and can support abstract linguistic representations (Kazanina et al. 2006; Okada & Hickok 2006; Obleser & Eisner 2009; Okada et al. 2010; Poeppel et al. 2012; Gwilliams et al. 2025). Superior Temporal Sulcus and STG are posited to be the locus for encoding phonetic categories. In Hickok & Poeppel (2007), phonological processing localized to bilateral STS, mediating bidirectional connections with low-level acoustic phonetic processing in STG and lexical processing in MTG, along the ventral stream. Syllable and word-related phonetic and phonological codes have been posited in middle pSTG and dorsal STS (Hickok 2025), although other models place the locus of phonetic and phonological encoding anterior to primary auditory cortex (DeWitt & Rauschecker 2012). That said, left inferior frontal regions and not superior temporal regions have been implicated in phonetic category invariance (Myers et al. 2009). Further, and perhaps contrary to the ECoG and MEG literature, left MTG and left angular gyrus (i.e., region just posterior to Wernicke’s Area) have been reported in fMRI to underlie the mapping between low-level acoustics and phonetic categories (Blumstein et al. 2005). Auditory speech processing is largely bilateral until contact with the lexicon is made (i.e., MTG), at which point, neurophysiological activity is largely left-lateralized. See Figure 1C for sketch of the left hemisphere regions typically implicated in auditory, phonetic, phonological and lexical processing and representation. Moreover, the MMN was initially localized to regions adjacent to primary auditory cortex (Aulanko et al. 1993), although more recent designs that aim to capture more abstract phonological representation ultimately elicit MMN responses that reach maximum amplitude later than is typically observed (Fu & Monahan 2021; Monahan et al. 2022; Politzer-Ahles & Jap 2024); this potentially suggests that the locus of phonological computations may be more distributed and further from primary auditory cortex.

Figure 1: (A) Schematic of a mismatch negativity (MMN) auditory oddball paradigm. A series of standard tokens is auditorily presented and interrupted by the occasional deviant stimulus that differs from the standard along some physical, perceptual, or representational dimension. The evoked brain responses to standard (blue) and deviant (red) trials are extracted from the continuous electroencephalography (EEG) signal (frontocentral electrode Fz is presented) and averaged together on a condition-by-condition, channel-by-channel basis. (B) Channel locations from a standard 32-channel EEG electrode array and a sample standard and deviant response in an MMN oddball paradigm. The event-related potential (ERP) response to the deviant (red line) and standard (blue line) stimuli at electrode Fz is provided. The response to the deviant is more negative in the 150–250-ms time window; the difference between the standard and deviant responses is shaded in gray. (C) Some principal areas implicated in the neurobiological network for speech perception and spoken word recognition. Figure and caption reprinted from Monahan (2018).

fMRI studies have revealed that broader phonetic contrasts (i.e., voicing, manner, and place) show distributed activation patterns in bilateral superior temporal cortex, with left perisylvian regions selectively coding place and voicing, and right posterior lateral fissure coding manner (Arsenault & Buchsbaum 2015; see their Figure 7 for a visualization of the overlapping and non-overlapping regions that encode these broad class features), and place of articulation and voicing activate distinct bilateral regions in the superior and medial temporal lobes (Lawyer & Corina 2014). Cortical regions along the dorsal pathway, including auditory, sensorimotor, and motor areas, support generalization of place discrimination from stops to fricatives, implicating motor involvement in speech perception (Correia et al. 2015). Phonetic categories appear to be spatially coded. Using non-invasive methods, more anterior vowels localize to more anterior regions of STG (Obleser et al. 2003; 2004), although other reports indicate that five vowel systems spatially organize in accordance with the vowel trapezoid in auditory cortex (Manca et al. 2019). More complex mappings, for example 2 × 2 × 2 vowel systems (e.g., Turkish), spatially organize along orthogonal vowel maps in auditory cortex (Scharinger et al. 2011).

Since 2014, there has been considerable excitement over the potential of ECoG to elucidate the neurophysiological encoding of speech sounds (see Yi et al. 2019 for a programmatic review). Mesgarani et al., (2014) observed that human left STG spatially codes distinct phonetic classes, largely based on manner of articulation. Moreover, neuronal populations in left STG are tuned to the first (F1) and second (F2) vowel formants (Oganian et al. 2023), voice onset time (VOT; Fox et al. 2020), and tone contrasts (Li et al. 2021). While impressive, this work does not isolate phonological levels of representation and instead often confounds with phonetic cues. Moreover, it is unclear how the ECoG results reconcile with the spatial maps obtained via non-invasive techniques described in the preceding paragraph. The principal challenge is an accurate spatial map of superior temporal cortex functions as each method has its limitations: MEG lacks adequate spatial precision, fMRI lacks adequate temporal resolution, and ECoG requires aggregation across multiple patient datasets.

To attempt to disentangle the acoustic from phonological encoding in cortex, Mai et al (2024) conducted an ECoG study that tested English flapping, and identified various sites around the brain that were ostensibly selectively sensitive to the underlying phonological structure than the superficial acoustic structure. That is, evoked ECoG activity at certain electrode sites, as well as time-frequency responses—across all frequency bands tested—to the [ɾ] was more like other allophones of /t/ as compared to the flap [ɾ] derived from an underlying /d/ phoneme. Other electrode sites, however, appeared to encode the surface phonetic relationship, where the two [ɾ] allophones—one derived from underlying /t/ and one derived from underlying /d/—patterned together to the exclusion of the [t] phone. Moreover, low-frequency oscillatory activity (i.e., δ-band, θ-band) was best fit when phonemic category labels were included in the model in addition to spectral information. These results minimally suggest that both surface and underlying phonological structure are encoded in various neurophysiological responses.

One major tenet of SFP is that significant abstraction exists within the phonological system, and we believe that there is potential for substantiating this claim. The theoretical challenge lies in distilling the core principles of a substance-free framework and formulating testable predictions. Even researchers with a strong phonetic orientation must acknowledge a degree of abstraction, particularly when positing sound classes, such as nasal consonants, that lack a simple acoustic/auditory definition (Stevens 1998). Mielke (2008: 45) cites EEG and MEG evidence at the time that supported abstract phonological features in auditory cortex; however, it was noted that whether these features were innate or learned was still an open question. In Mielke’s preferred model, speech segments are initially processed as holistic units and later receive language particular featural interpretations. MMN experiments are relatively easy to design and run, and as such, we provide a more extended discussion.

3.2 Mismatch responses

The MMN is an ERP component observed in EEG/MEG signals that reflects automatic change detection in auditory processing (Näätänen 2001; Näätänen et al. 2007; 2019). In a typical oddball paradigm, participants hear a series of frequent standard stimuli interrupted by infrequent deviants that differ along physical, perceptual, or representational dimensions (see Figure 1A). Initially thought to reflect comparison against a stored memory trace, the MMN is now often interpreted through a predictive coding lens, where auditory cortex encodes regularities and the MMN arises from a mismatch between the deviant and predicted input (Winkler 2007). The MMN component itself typically peaks 150–350 milliseconds post-deviant and is largest over fronto-central EEG electrodes (see Figure 1B) and temporal MEG sensors. When auditory cortex encodes a distinction, an MMN is elicited, localizing to superior temporal regions (Sams et al. 1991; Csépe et al. 1992; Hari et al. 1992; Näätänen et al. 2007). Crucially, the MMN operates pre-attentively, that is, participants need not attend to the stimuli for it to be observed (Atienza et al. 2002; Vanhaudenhuyse et al. 2008). In speech, the MMN is sensitive to language-specific distinctions, including vowel categories (Näätänen et al. 1997; Winkler et al. 1999) and voice onset time (VOT) contrasts in stops (Sharma & Dorman 1999; Sharma et al. 2000), suggesting that linguistic knowledge may shape what is encoded by the MMN; however, at that point, it was unclear whether those auditory memory representations involved reflected acoustic, phonetic, or phonological levels of processing.

While we certainly advocate for a methodologically pluralist approach, that is, one should be careful not to prioritize neurophysiological or neurological evidence over other forms of evidence (e.g., perception, production, behaviour, eye-tracking), to date, evidence for abstraction in phonology in behavioural findings is not entirely clear (though see Caplan et al. (2021) for clever behavioural evidence for intermediate speech-sound categories not based in subphonemic acoustic detail); instead, results from MMN investigations have made significant contributions to these questions (Monahan 2018).

To this end, Phillips et al. (2000; see Figure 2) introduced intra-category acoustic variation in the standard and deviant stimuli (e.g., [da08 da00 da16 da24 da08 da16 da24 ta64 da16 da00…]; subscripts refer to VOT values in milliseconds), an innovation on previous designs (Sharma & Dorman 1999; Sharma et al. 2000). Despite this acoustic and auditory variation, an MMN was observed when the distribution of acoustic tokens aligned with the VOT distribution observed in English stop consonants, while no such MMN was observed when the same acoustic variation no longer aligned with English phonetic category boundaries. To our knowledge, this is the first published MMN study—at least in the domain of speech—wherein acoustic variability was included in the repeated standards. If auditory cortex was unable to abstract over the variation in VOT values, there would be no many-to-one relationship and as such, no MMN would have been predicted. That an MMN was observed, however, indicates that listeners did, indeed, abstract over the acoustic variation in VOT values and encoded each token as a phonetic category (e.g., [t], [d]). Similar results were found comparing phonemic versus allophonic contrasts in a cross-linguistic context (Kazanina et al. 2006), again highlighting the possibility that auditory cortex supports abstract category representations.

Figure 2: Schematic of the many-to-one oddball paradigm developed in Phillips et al. (2000). In the Phonological Experiment, participants are exposed to a series of auditory CV stimuli that repeat at the category-level (i.e., [dɑ]) but not at the acoustic level, as voice onset time (VOT) duration changes from stimulus to stimulus. These standard tokens are interrupted by an infrequent, deviant stimulus drawn from the other phonetic category (i.e., [tɑ]). In the Acoustic Experiment, 20 ms of VOT is added to each token, removing the many-to-one relationship between the tokens and their category membership. A mismatch negativity (MMN) is only reported in the Phonological Experiment.

Since then, numerous MMN speech studies have observed asymmetric responses: Some standard-deviant pairings elicit larger, asymmetric MMN responses depending on which category is the standard and which category is the deviant. Specifically, asymmetric MMNs have been observed for vowels (Obleser et al. 2006; Cornell et al. 2011; Scharinger et al. 2012; Scharinger & Monahan & et al. 2016; de Rue et al. 2021; Yu & Shafer 2021), consonants (Cornell et al. 2013; Hestvik & Durvasula 2016; Schluter et al. 2016; 2017; Højlund et al. 2019; Hestvik et al. 2020; Meng et al. 2021b; Rhodes et al. 2022), lexical tone (Politzer-Ahles et al. 2016) and local assimilation contexts (Meng et al. 2021a). Similar asymmetries have been observed outside of the oddball MMN paradigm: Specifically, increases in BOLD signal change using fMRI are observed in bilateral STS and STG when the first word of the word pair contains a specified vowel and the second word contains an underspecified vowel relative to when the first word contains an underspecified vowel and the second word contains a specified vowel (Scharinger & Domahs & et al. 2016). Asymmetries between coronal and non-coronal mispronunciations are observed in at 24-months of age, while no such asymmetries are observed at 18-months of age, indicating that such underspecified representations begin to emerge by the second year of life (Althaus et al. 2024). Moreover, at least in the MMN paradigm, however, no such asymmetries are observed when participants hear a single, repeating acoustic token (cf. Phillips et al., 2000), suggesting that the MMN has the power to index a variety of speech sound representational levels (Hestvik & Durvasula 2016). These findings are routinely interpreted in terms of underspecification theory (Archangeli 1988) or a Featurally Underspecified Lexicon (FUL; Lahiri & Reetz 2002; 2010; Lahiri 2018); however, a common shortcoming of MMN studies is that the tested phonological contrasts are strongly correlated with phonetic or even auditory properties, for example testing for voicing, could end up confounded with the auditory detection of periodicity (pitch) in a portion of the signal. One the surface, it appears that the Phillips et al. (2000) findings provided compelling evidence for abstract phonological representations (see Hestvik & Durvasula (2016) for a comparison of variation and no variation in the standard stimuli); however, only one acoustic variable was manipulated, rendering the findings consistent with phonetic category access. Moreover, in the analysis of the Acoustic Experiment, the stimuli were analyzed against a “long” versus “short” VOT continuum, obfuscating the category labels of the tokens and their relation to the many-to-one feature of the paradigm. An alternative analysis would include recoding the tokens based on the expected phonetic categorization.

Designing an experiment that isolates abstract, phonological representations is extraordinarily difficult, yet identifying the appropriate contrast set is paramount to moving forward and demands careful attention (Phillips 2001; Monahan 2018; Mai et al. 2024). One avenue forward is to force listeners to group across manners of articulation, as phonological classes that span distinct manners are less likely to share a common phonetic base, permitting a more decisive attribution of observed effects to the phonology. Using MEG, Flagg et al. (2006) showed, despite the cue to nasality being distinct between vowels and consonants, nasal vowels predicted nasal consonants. To account for these findings, one could postulate an abstract [nasal] feature, implying a certain amount of abstraction: nasal predicts nasal. On the other hand, one could provide a purely auditory account, postulating that nasal vowels predict nasal consonants in English. Most MMN studies do not employ an overt task, and as such, the researcher must ensure that listeners are grouping the standards and deviants in the intended. This is a relatively difficult task when working with two, minimally contrasting categories, but it is becoming clear that more complex designs are needed to tap into abstract, phonological representations.

Another possibility is to test a hallmark prediction of features. Namely, features organize individual sound categories into larger classes based on shared featural representations. Previous MMN studies had shown that auditory cortex is able to generalize across multiple varying acoustic parameters as long as one acoustic feature is constant (Gomes et al. 1995). One could include natural classes of phonological categories into the repeating standard stimuli; that is, instead of a single repeating phonetic category, participants are presented with distinct phonetic categories all belonging to the same natural class of segments. An observation of an MMN would indicate that the brain could ignore non-overlapping phonetic cues between standard tokens and store a memory trace of the one consistent feature, rendering a purely acoustic-phonetic analysis difficult to maintain. To determine whether the brain can group individual sound categories on the basis of such shared representations, Fu and Monahan (2021) tested Mandarin retroflex consonants by innovating an MMN paradigm to include inter-category variation in the standards (e.g., [ʂɤ tʂɤ ɻɤ tʂɤʰ …]). There, an MMN emerged only when retroflex consonants were the standard and nonretroflex the deviant, suggesting sensitivity to [retroflex], although it was unclear whether this reflects abstract phonological processing or acoustic cue detection (e.g., F3; Hussain et al. 2017). That is, this is another instance where it is difficult to dissociate phonological factors from shared acoustic properties.

To obviate this issue, Monahan et al. (2022) tested English voicing using stops and fricatives (e.g., [pʰɑ sɑ kʰɑ …]), which rely on distinct, temporal (i.e., VOT) and spectral (i.e., low-frequency energy) phonetic cues, respectively. We observed an MMN when voiceless obstruents were the standard, aligning with prior findings (Cornell et al. 2013; Schluter et al. 2017) and linguistic analyses of English laryngeal features (Avery & Idsardi 2001). The elicitation of an MMN indicated that listeners disjunctively coded these distinct phonetic cues into an abstract, phonological representation. This is the first finding to our knowledge that the brain can create an integrated percept based on the phonological structure of the language and that is not confounded with an acoustic-phonetic parameter. That is, there is no single phonetic property that can be used to group the voiceless stops and voiceless fricatives into a single category. In short, listeners form an abstract representation of a phonological class that is believed to require two distinct auditory features and bind them into a single, phonological feature. The MMN paradigm, however, is not without limitations. One data point (e.g., deviant presentation) requires approximately ten stimuli (i.e., standard presentation). This renders such experiments slow and requiring orders of magnitude of reduction from the raw data and stimuli to draw conclusions of a single contrast.

4 Substantiating Substance Free Phonology experimentally

SFP has several variants, but one common, shared attribute is a commitment to the abstractness of phonological representations—that they are (relatively) unmoored from their articulatory and auditory phonetic anchors. SFP is not alone in this, even textbook presentations of phonetics ignore certain distinctions in pronunciation. Contrast-based phonology (Dresher 2009) aims to find a minimal decision tree of phonological differences in each language, and Berent (2013; 2026) offers a broad defense of abstract phonological representations, even across different modalities (Berent et al. 2021). This broad adoption makes abstractness of speech sound representations an appealing candidate for neurophysiological investigation as any results will be useful not only to SFP but to other approaches as well. It is also true that SFP thus far has been primarily concerned with the analysis of speech sounds and features in phonological processes, often using set theory (Bale & Reiss 2018; Bale et al. 2019; Reiss 2022; Gorman & Reiss 2026). A simple but useful illustration of abstractness is the quartet of English speech sounds /p b f v/ as they occur in word-initial position before stressed vowels, that is, #_V, a common environment for taking phonetic measurements and an important position for spoken word recognition (Cutler 2012; Sun & Poeppel 2023). Phonetically, these sounds differ in myriad ways (what Avery & Idsardi 2001 term phonetic over-differentiation), such that the four sounds occupy the potential phonetic space quite sparsely, see Table 2. We use ordinary descriptive labels for speech sounds and attributes here to not prejudge the content of the mental representations for speech sounds; that is, we attempt to assess the “topology” or “shape” of the mental relationships.

Table 2: Some phonetic properties of English labial obstruents /p b f v/.

Place Bilabial Labiodental
Constriction Stop Fricative Stop Fricative
Closure voicing Periodic v
Aperiodic f
VOT lag Short b
Long p

Many of the empty cells in Table 2 can be filled with phonetic instantiations in other languages or in other word positions in English. Ewe distinguishes between labial and labio-dental fricatives (Maddieson 2005), and German /pf/ can be pronounced with a labiodental closure (Kehrein 2013). Korean “plain” /s/ is aspirated, that is, it has a substantial VOT duration after the end of sibilant frication (Martin 1992). The “voicing” contrasts are displayed as minimal in Table 2, in the sense that a single attribute is used to distinguish between the two fricatives (periodic vocal fold vibrations during closure) and between the two stops (VOT lag), but we note again that the distinguishing attribute is not the same across the two pairs (see previous section). Moreover, even this description is a simplification, especially when other word positions are considered (Lisker 1977; Smith 1997).

That said, in phonological analyses, the sub-distinctions for place and for laryngeal postures and timing are thought not to be relevant to the storage of wordforms in long-term memory (e.g., Berg 1989; but see Mitterer et al. 2013), and the quartet is analyzed as a 2 × 2 system of contrasts, as shown in Table 3 (again with descriptive labels and symbols):

Table 3: 2 × 2 phonological reduction of Table 1.

Stop Fricative
Voiceless p f
Voiced b v

Table 3 may seem to be the obvious phonological reduction of Table 2, but other arrangements are possible. One topological equivalent would replace manner (stop/fricative) with place (bilabial/labiodental), a less obvious one would use closure voicing and VOT lag as the primary factors. Likewise, voiced/voiceless could instead be analyzed in terms of glottal tension or glottal width (Avery & Idsardi 2001), and various possible underspecification analyses are possible for such 2 × 2 arrangements, with concomitantly different predictions regarding MMN asymmetries. That is, if the alternative, phonetically substantive analysis in Table 2 is followed instead, then /b/ and /v/ do not share any attribute in common, and so in an MMN paradigm could not be combined into a standard representation against which deviants are judged. But, as noted in the previous section, Monahan et al. (2022) were able to do exactly this in their “voiced standards block”, and their results were replicated by Politzer-Ahles & Jap (2024).

Additionally, how the work in accounting for both speech and long-term memory representations is divided between phonological representations and phonetic implementation (Keating 1996) is critically important for all theories incorporating abstract representations, including both SFP and contrast-based phonology. For instance, one relevant observation in the present context is the pronunciation of bilabial or labiodental nasals in /mp/ and /mf, nf/ clusters (“camping”, “camphor”, “infamous”), seen either as the sharing of place information in the phonology, or as phonetic coarticulation, or as a combination of both (see Flynn 2025).

Trampling over these important subtleties and borrowing the four-part (proportional) analogy from historical linguistics, then what phonologists generally want is an analysis for English /p b f v/ where both p:b::f:v and p:f::b:v obtain, i.e., the 2 × 2 arrangement. Is it possible to test the abstract structure of the distinctions (2 × 2) with neuro-physiological methods without worrying about the exact sonic implementations? No, in the sense that the development of experimental materials will have to attend to details of pronunciation to create suitable stimuli; but maybe, because the mismatch response does display variability in amplitude and timing, and in some cases, seems to provide at least a rough distance metric between sounds. For example, Garrido et al. (2013) examined the mismatch response to tones of different frequencies and found that the amplitude of the mismatch field varied with perceptual distance in frequency between the tones as measured on a logarithmic scale of octaves (semitones; or a quasi-logarithmic scale such as Bark; see also Bergelson et al. 2013). Similarly, although the amplitude of the MMN is sensitive to sub-phonemic (or even sub-allophonic) differences in some cases (Han et al. 2026), the inclusion of phonemic differences tends to diminish (or overwash) this effect, and in other cases (e.g., Kazanina et al. 2006), any allophonic effects were too small to detect, consistent with the common observation that MMN amplitudes are larger for phonemic contrasts relative to allophonic ones. Overall, then, we can attempt to look at relative MMN amplitudes to reveal some of the nature of the differences between standards and deviants.

Chabot et al. (2026) examined this quartet of English sounds in two mismatch designs, one with a common standard /p/ and roving (varying) deviants /b f v/, and one with roving standards /p f b/ and a common deviant /v/. The second design with a common deviant has the additional advantage of more nearly equalizing the number of presentations of each of the four sounds in the experimental block. The main finding, at least to a first approximation, accords with the 2 × 2 topology: Differences of a single attribute (p-f, p-b; b-v, f-v) yield smaller responses, whereas differences of two attributes (p-v) yield a significantly larger, approximately additive response. In contrast, the near equality of the single attribute responses and the additivity effect would be serendipitous coincidences in the substantive phonetic analysis of Table 2. While such a coincidence is not impossible, it would become less likely if these effects are replicated in future studies. Other experiments have also shown MMN additivity effects, with some caveats (Paavilainen et al. 2001; Wolff & Schröger 2001; Jacobsen et al. 2013).

If this initial additivity finding can be replicated across other languages and conditions, this would have consequences for the linking theory between brain responses and phonological representations, which for segments and features have been formulated using set theory in SFP. In set-theoretic terms, it suggests that in the MMN design, the comparison between English /p/ and /v/ yields something like the symmetric set difference, AΔB = (A\B) ∪ (B\A) (or, alternatively AΔB = (A∪B) \ (A∩B), or AΔB = {x : (x∊A) XOR (x∊B)}). For /p/ and /v/ this difference would descriptively be {voiced, voiceless, stop, fricative}. That is, as descriptively desired, the symmetric set difference returns something akin to the mismatches between the two sounds. This formulation is distinct from the SFP idea that sets of phonological features impose conditions of feature-value consistency—voiced and voiceless are inconsistent—and that phonological sets are combined using unification or priority union (Reiss 2022). Under that view, the mismatch response could be due to unification failure between the standards and the deviant. But when unification succeeds, it returns the union of consistent features; when unification fails, it does not return a set of mismatching features but instead a Boolean value of False. Importantly, this is not an incoherent linking theory for phonology and MMN responses, it simply predicts that the MMN responses should not be additive, as by definition, there are no degrees of being False. Of course, other creative amendments for set combination in MMN analysis are possible, unification failure could instead return the cardinality of the symmetric set difference (four in this case, Reiss p.c.). Such a set cardinality linking theory would have the implicit claim that MMN amplitude for single attribute mismatches should be approximately equal rather than what was observed in Chabot et al. (2026). As a practical experimental matter, it is probably not possible to measure more than a few degrees of mismatching as MMN amplitude analysis is bounded below by the experimental power, and above by saturation effects. Overall, the moral here accords with Hornstein (2026): If you want sets in your theory, then there is no escape from set theory, and your task is to find the best set-theoretic device to match the experimental behavior.

5 Summary

Neuro-physiological experiments have yielded important information about phonetic and phonological representations in human auditory cortex. Designing and interpreting such experiments remains quite challenging because auditory, phonetic and phonological properties are often correlated so no single experiment is likely to disentangle the confounds despite the creativity and subtlety of many of the experimental designs and neural measures. Therefore, experimental evidence regarding abstract phonological representations is likely to accrue slowly across a large set of experiments investigating different phonological contrasts across a variety of languages which we hope will eventually lead to a convergent, consensus answer consistent with findings from non-neural methods, such as typological analysis, acquisition patterns and computational considerations. This article has offered a survey of many of the relevant results to date and some speculation on how dissecting the MMN response into subcomponents might offer a rich vein of information for further experiments to mine.

Funding

This work was supported in part by the Natural Sciences and Engineering Research Council (NSERC) of Canada, grant number: RGPIN-2025-06584.

Competing Interests

The authors have no competing interests to declare.

References

Adank, Patti & Nuttall, Helen E. & Kennedy-Higgins, Dan. 2017. Transcranial magnetic stimulation and motor evoked potentials in speech perception research. Language, Cognition and Neuroscience 32(7). 900–909. DOI:  http://doi.org/10.1080/23273798.2016.1257816

Aksenov, Daniil P. & Li, Limin & Miller, Michael J. & Wyrwicz, Alice M. 2019. Role of the inhibitory system in shaping the BOLD fMRI response. NeuroImage 201. 116034. DOI:  http://doi.org/10.1016/j.neuroimage.2019.116034

Althaus, Nadja & Lahiri, Aditi & Plunkett, Kim. 2024. Coronal underspecification as an emerging property in the development of speech processing. Journal of Experimental Psychology: Learning, Memory, and Cognition 50(12). 1932–1953. DOI:  http://doi.org/10.1037/xlm0001367

Anderson, Steven T. 2018. Economics, helium, and the U.S. federal helium reserve: Summary and outlook. Natural Resources Research 27(4). 455–477. DOI:  http://doi.org/10.1007/s11053-017-9359-y

Archangeli, Diana. 1988. Aspects of underspecification theory. Phonology 5(2). 183–207. DOI:  http://doi.org/10.1017/S0952675700002268

Arsenault, Jessica S. & Buchsbaum, Bradley R. 2015. Distributed neural representations of phonological features during speech perception. The Journal of Neuroscience 35(2). 634–642. DOI:  http://doi.org/10.1523/JNEUROSCI.2454-14.2015

Atienza, Mercedes & Cantero, Jose L & Dominguez-Marin, Elena. 2002. Mismatch negativity (MMN): an objective measure of sensory memory and long-lasting memories during sleep. International Journal of Psychophysiology 46(3). 215–225. DOI:  http://doi.org/10.1016/S0167-8760(02)00113-7

Aulanko, Reijo & Hari, Riitta & Lounasmaa, Olli V. & Näätänen, Risto & Sams, Mikko. 1993. Phonetic invariance in the human auditory cortex. NeuroReport 4(12). 1356–1358.

Avery, Peter & Idsardi, William J. 2001. Laryngeal dimensions, completion and enhancement. In Hall, T. Alan (ed.), Distinctive Feature Theory, 41–70. Berlin, Germany: Walter de Gruyter.

Bale, Alan & Reiss, Charles. 2018. Phonology: A formal introduction. Cambridge, MA: MIT Press.

Bale, Alan & Reiss, Charles & Ta-Chun Shen, David. 2019. Sets, rules and natural classes: [ ] vs. { }. Loquens 6(2). e065. DOI:  http://doi.org/10.3989/loquens.2019.065

Bates, Elizabeth & Wilson, Stephen M. & Saygin, Ayse Pinar & Dick, Frederic & Sereno, Martin I. & Knight, Robert T. & Dronkers, Nina F. 2003. Voxel-based lesion–symptom mapping. Nature Neuroscience 6(5). 448–450. DOI:  http://doi.org/10.1038/nn1050

Berent, Iris. 2013. The phonological mind. Cambridge, UK: Cambridge University Press.

Berent, Iris. 2026. Three arguments for abstraction in phonology. Glossa: A Journal of General Linguistics 11(1). DOI:  http://doi.org/10.16995/glossa.27277

Berent, Iris & Bat-El, Outi & Brentari, Diane & Andan, Qatherine & Vaknin-Nusbaum, Vered. 2021. Amodal phonology. Journal of Linguistics 57(3). 499–529. DOI:  http://doi.org/10.1017/S0022226720000298

Berg, Thomas. 1989. How phonetic is a phonological feature representation? The case of labiodental fricatives. Speech Communication 8(4). 329–345. DOI:  http://doi.org/10.1016/0167-6393(89)90015-0

Bergelson, Elika & Shvartsman, Michael & Idsardi, William J. 2013. Differences in mismatch responses to vowels and musical intervals: MEG Evidence. PLoS ONE 8(10). e76758. DOI:  http://doi.org/10.1371/journal.pone.0076758

Berger, Hans. 1929. Über das Elektrenkephalogramm des Menschen. Archiv Für Psychiatrie Und Nervenkrankheiten 87(1). 527–570. DOI:  http://doi.org/10.1007/BF01797193

Bernal, Jose & Kushibar, Kaisar & Asfaw, Daniel S. & Valverde, Sergi & Oliver, Arnau & Martí, Robert & Lladó, Xavier. 2019. Deep convolutional neural networks for brain image analysis on magnetic resonance imaging: A review. Artificial Intelligence in Medicine 95. 64–81. DOI:  http://doi.org/10.1016/j.artmed.2018.08.008

Binder, Jeffrey R. & Desai, Rutvik H. & Graves, William W. & Conant, Lisa L. 2009. Where is the semantic system? A critical review and meta-analysis of 120 functional neuroimaging studies. Cerebral Cortex 19(12). 2767–2796. DOI:  http://doi.org/10.1093/cercor/bhp055

Blumstein, Sheila E. & Myers, Emily B. & Rissman, Jesse. 2005. The perception of voice onset time: An fMRI investigation of phonetic category structure. Journal of Cognitive Neuroscience 17(9). 1353–1366. DOI:  http://doi.org/10.1162/0898929054985473

Broca, Paul. 1865. Sur le siège de la faculté du langage articulé. Bulletins et Mémoires de la Société d’Anthropologie de Paris 6(1). 377–393. DOI:  http://doi.org/10.3406/bmsap.1865.9495

Brodbeck, Christian & Hong, L. Elliot & Simon, Jonathan Z. 2018. Rapid transformation from auditory to linguistic representations of continuous speech. Current Biology 28(24). 3976–3983.e5. DOI:  http://doi.org/10.1016/j.cub.2018.10.042

Brookes, Matthew J. & Leggett, James & Rea, Molly & Hill, Ryan M. & Holmes, Niall & Boto, Elena & Bowtell, Richard. 2022. Magnetoencephalography with optically pumped magnetometers (OPM-MEG): The next generation of functional neuroimaging. Trends in Neurosciences 45(8). 621–634. DOI:  http://doi.org/10.1016/j.tins.2022.05.008

Buxton, Richard B. & Wong, Eric C. & Frank, Lawrence R. 1998. Dynamics of blood flow and oxygenation changes during brain activation: The balloon model. Magnetic Resonance in Medicine 39(6). 855–864. DOI:  http://doi.org/10.1002/mrm.1910390602

Caplan, Spencer & Hafri, Alon & Trueswell, John C. 2021. Now you hear me, later you don’t: The immediacy of linguistic computation and the representation of speech. Psychological Science 32(3). 410–423. DOI:  http://doi.org/10.1177/0956797620968787

Chabot, Alexander Marx. 2024. Substance-free approaches to phonology. In Nasukawa, Kuniya & Samuels, Bridget & Schwartz, Geoff & Törkenczy, Miklós (eds.), Wiley-Blackwell Companion to Phonology, 42. Blackwell Publishing.

Chabot, Alexander Marx & Lau, Ellen F. & Monahan, Philip J. & Idsardi, William J. 2026. Integrated versus independent processing of auditory features in speech sounds. Language, Cognition and Neuroscience 41(4). 470–486. DOI:  http://doi.org/10.1080/23273798.2026.2630749

Chiarion, Giovanni & Sparacino, Laura & Antonacci, Yuri & Faes, Luca & Mesin, Luca. 2023. Connectivity analysis in EEG data: A tutorial review of the state of the art and emerging trends. Bioengineering 10(3). 372. DOI:  http://doi.org/10.3390/bioengineering10030372

Chiong, Winston & Leonard, Matthew K. & Chang, Edward F. 2018. Neurosurgical patients as human research subjects: Ethical considerations in intracranial electrophysiology research. Neurosurgery 83(1). 29–37. DOI:  http://doi.org/10.1093/neuros/nyx361

Cornell, Sonia A. & Lahiri, Aditi & Eulitz, Carsten. 2011. “What you encode is not necessarily what you store”: Evidence for sparse feature representations from mismatch negativity. Brain Research 1394. 79–89. DOI:  http://doi.org/10.1016/j.brainres.2011.04.001

Cornell, Sonia A. & Lahiri, Aditi & Eulitz, Carsten. 2013. Inequality across consonantal contrasts in speech perception: Evidence from mismatch negativity. Journal of Experimental Psychology: Human Perception and Performance 39(3). 757–772. DOI:  http://doi.org/10.1037/a0030862

Correia, Joao M. & Jansma, Bernadette M. B. & Bonte, Milene. 2015. Decoding articulatory features from fMRI responses in dorsal speech regions. Journal of Neuroscience 35(45). 15015–15025. DOI:  http://doi.org/10.1523/JNEUROSCI.0977-15.2015

Csépe, Valéria & Pantev, Christo & Hoke, Manfried & Hampson, Scott & Ross, Bernhard. 1992. Evoked magnetic responses of the human auditory cortex to minor pitch changes: Localization of the mismatch field. Electroencephalography and Clinical Neurophysiology/Evoked Potentials Section 84(6). 538–548. DOI:  http://doi.org/10.1016/0168-5597(92)90043-B

Cutler, Anne. 2012. Native listening: Language experience and the recognition of spoken words. The MIT Press. DOI:  http://doi.org/10.7551/mitpress/9012.001.0001

D’Ausilio, Alessandro & Pulvermüller, Friedemann & Salmas, Paola & Bufalari, Ilaria & Begliomini, Chiara & Fadiga, Luciano. 2009. The motor somatotopy of speech perception. Current Biology 19(5). 381–385. DOI:  http://doi.org/10.1016/j.cub.2009.01.017

de Rue, Nadine P. W. D. & Snijders, Tineke M. & Fikkert, Paula. 2021. Contrast and conflict in Dutch vowels. Frontiers in Human Neuroscience 15. 629648. DOI:  http://doi.org/10.3389/fnhum.2021.629648

Defenderfer, Jessica & Kerr-German, Anastasia & Hedrick, Mark & Buss, Aaron T. 2017. Investigating the role of temporal lobe activation in speech perception accuracy with normal hearing adults: An event-related fNIRS study. Neuropsychologia 106. 31–41. DOI:  http://doi.org/10.1016/j.neuropsychologia.2017.09.004

DeWitt, Iain & Rauschecker, Josef P. 2012. Phoneme and word recognition in the auditory ventral stream. Proceedings of the National Academy of Sciences 109(8). E505–E514. DOI:  http://doi.org/10.1073/pnas.1113427109

Dresher, Bezalel E. 2009. The contrastive hierarchy in phonology. Cambridge, UK: Cambridge University Press.

Dronkers, Nina F. & Baldo, Juliana V. 2009. Language: Aphasia. In Encyclopedia of neuroscience, 343–348. Elsevier. DOI:  http://doi.org/10.1016/B978-008045046-9.01876-3

Dronkers, Nina F. & Wilkins, David P. & Van Valin, Robert D. & Redfern, Brenda B. & Jaeger, Jeri J. 2004. Lesion analysis of the brain areas involved in language comprehension. Cognition 92(1). 145–177. DOI:  http://doi.org/10.1016/j.cognition.2003.11.002

Dubey, Agrita & Ray, Supratim. 2019. Cortical electrocorticogram (ECoG) is a local signal. Journal of Neuroscience 39(22). 4299–4311. DOI:  http://doi.org/10.1523/JNEUROSCI.2917-18.2019

Embick, David & Poeppel, David. 2015. Towards a computational(ist) neurobiology of language: correlational, integrated and explanatory neurolinguistics. Language, Cognition and Neuroscience 30(4). 357–366. DOI:  http://doi.org/10.1080/23273798.2014.980750

Fallegger, Florian & Schiavone, Giuseppe & Pirondini, Elvira & Wagner, Fabien B. & Vachicouras, Nicolas & Serex, Ludovic & Zegarek, Gregory & May, Adrien & Constanthin, Paul & Palma, Marie & Khoshnevis, Mehrdad & Van Roost, Dirk & Yvert, Blaise & Courtine, Grégoire & Schaller, Karl & Bloch, Jocelyne & Lacour, Stéphanie P. 2021. MRI-compatible and conformal electrocorticography grids for translational research. Advanced Science 8(9). 2003761. DOI:  http://doi.org/10.1002/advs.202003761

Flagg, Elissa J. & Oram Cardy, Janis E. & Roberts, Timothy P. L. 2006. MEG detects neural consequences of anomalous nasalization in vowel–consonant pairs. Neuroscience Letters 397(3). 263–268. DOI:  http://doi.org/10.1016/j.neulet.2005.12.034

Flynn, Darin. 2025. What the f? F-elements behind labiodentals and crazy rules. In DougSchrift: A collection of squibs and puzzles presented to Doug Pulleyblank, 24 pp. Vancouver, BC: University of British Columbia. Retrieved from https://arts-pulleyblank-2024.sites.olt.ubc.ca/

Fox, Neal P. & Leonard, Matthew & Sjerps, Matthias J. & Chang, Edward F. 2020. Transformation of a temporal speech cue to a spatial neural code in human auditory cortex. Elife 9. e53051.

Fu, Zhanao & Monahan, Philip J. 2021. Extracting phonetic features from natural classes: A mismatch negativity study of Mandarin Chinese retroflex consonants. Frontiers in Human Neuroscience 15. 609898. DOI:  http://doi.org/10.3389/fnhum.2021.609898

Garrido, Marta I. & Sahani, Maneesh & Dolan, Raymond J. 2013. Outlier responses reflect sensitivity to statistical structure in the human brain. PLOS Computational Biology 9(3). e1002999. DOI:  http://doi.org/10.1371/journal.pcbi.1002999

Gentry, Lindell R. & Godersky, John C. & Thompson, Brad. 1988. MR imaging of head trauma: Review of the distribution and radiopathologic features of traumatic lesions. American Journal of Neuroradiology 9(1). 101–110.

Gershman, Samuel J. & Niv, Yael. 2010. Learning latent structure: Carving nature at its joints. Current Opinion in Neurobiology 20(2). 251–256. DOI:  http://doi.org/10.1016/j.conb.2010.02.008

Gershon, Ari A. & Dannon, Pinhas N. & Grunhaus, Leon. 2003. Transcranial magnetic stimulation in the treatment of depression. American Journal of Psychiatry 160(5). 835–845. DOI:  http://doi.org/10.1176/appi.ajp.160.5.835

Geschwind, Norman. 1970. The organization of language and the brain: Language disorders after brain damage help in elucidating the neural basis of verbal behavior. Science 170(3961). 940–944. DOI:  http://doi.org/10.1126/science.170.3961.940

Giessel, Andrew J. & Datta, Sandeep Robert. 2014. Olfactory maps, circuits and computations. Current Opinion in Neurobiology 24. 120–132. DOI:  http://doi.org/10.1016/j.conb.2013.09.010

Gillis, Marlies & Vanthornhout, Jonas & Simon, Jonathan Z. & Francart, Tom & Brodbeck, Christian. 2021. Neural markers of speech comprehension: Measuring EEG tracking of linguistic speech representations, controlling the speech acoustics. Journal of Neuroscience 41(50). 10316–10329. DOI:  http://doi.org/10.1523/JNEUROSCI.0812-21.2021

Gomes, Hilary & Ritter, Walter & Vaughan, Herbert G. 1995. The nature of preattentive storage in the auditory system. Journal of Cognitive Neuroscience 7(1). 81–94. DOI:  http://doi.org/10.1162/jocn.1995.7.1.81

Gorman, Kyle & Reiss, Charles. 2026. Natural class reasoning in segment deletion rules. In Proceedings of the 56th annual meeting of the North East Linguistics Society. University of Massachusetts Amherst: GLSA.

Grosjean, Francois & Frauenfelder, Uli H. 1996. A guide to spoken word recognition paradigms: introduction. Language and Cognitive Processes 11(6). 553–558. DOI:  http://doi.org/10.1080/016909696386935

Grothe, Benedikt. 2003. New roles for synaptic inhibition in sound localization. Nature Reviews Neuroscience 4(7). 540–550. DOI:  http://doi.org/10.1038/nrn1136

Gwilliams, Laura & Bhaya-Grossman, Ilina & Zhang, Yizhen & Scott, Terri & Harper, Sarah & Levy, Deborah. 2025. Computational architecture of speech comprehension in the human brain. Annual Review of Linguistics 11. 209–226. DOI:  http://doi.org/10.1146/annurev-linguistics-031120-111245

Gwilliams, Laura & King, Jean-Remi & Marantz, Alec & Poeppel, David. 2022. Neural dynamics of phoneme sequences reveal position-invariant code for content and order. Nature Communications 13(1). 6606. DOI:  http://doi.org/10.1038/s41467-022-34326-1

Gwilliams, Laura & Linzen, Tal & Poeppel, David & Marantz, Alec. 2018. In spoken word recognition the future predicts the past. The Journal of Neuroscience 38(35). 7585–7599. DOI:  http://doi.org/10.1523/JNEUROSCI.0065-18.2018

Gwilliams, Laura & Marantz, Alec & Poeppel, David & King, Jean-Remi. 2024. Top-down information shapes lexical processing when listening to continuous speech. Language, Cognition and Neuroscience 39. 1045–1058. DOI:  http://doi.org/10.1080/23273798.2023.2171072

Han, Chao & Hestvik, Arild & Idsardi, William J. 2026. Within-category mismatch responses to parametrically varying speech sounds. European Journal of Neuroscience 63(8). e70482. DOI:  http://doi.org/10.1111/ejn.70482

Hari, Riitta & Rif, Josi & Tiihonen, Jari & Sams, Mikko. 1992. Neuromagnetic mismatch fields to single and paired tones. Electroencephalography and Clinical Neurophysiology 82(2). 152–154. DOI:  http://doi.org/10.1016/0013-4694(92)90159-F

Hestvik, Arild & Durvasula, Karthik. 2016. Neurobiological evidence for voicing underspecification in English. Brain and Language 152. 28–43. DOI:  http://doi.org/10.1016/j.bandl.2015.10.007

Hestvik, Arild & Shinohara, Yasuaki & Durvasula, Karthik & Verdonschot, Rinus G. & Sakai, Hiromu. 2020. Abstractness of human speech sound representations. Brain Research. 146664. DOI:  http://doi.org/10.1016/j.brainres.2020.146664

Hickok, Gregory. 2009. The functional neuroanatomy of language. Physics of Life Reviews 6(3). 121–143. DOI:  http://doi.org/10.1016/j.plrev.2009.06.001

Hickok, Gregory. 2025. Wired for words: The neural architecture of language. Cambridge, Mass: MIT Press. DOI:  http://doi.org/10.7551/mitpress/9000.001.0001

Hickok, Gregory & Poeppel, David. 2004. Dorsal and ventral streams: A framework for understanding aspects of the functional anatomy of language. Cognition 92(1–2). 67–99. DOI:  http://doi.org/10.1016/j.cognition.2003.10.011

Hickok, Gregory & Poeppel, David. 2007. The cortical organization of speech processing. Nature Reviews Neuroscience 8(5). 393–402. DOI:  http://doi.org/10.1038/nrn2113

Højlund, Andreas & Gebauer, Line & McGregor, William B. & Wallentin, Mikkel. 2019. Context and perceptual asymmetry effects on the mismatch negativity (MMNm) to speech sounds: An MEG study. Language, Cognition and Neuroscience 34(5). 545–560. DOI:  http://doi.org/10.1080/23273798.2019.1572204

Hornstein, Norbert. 2026. Are phrase markers sets? University of Maryland.

Howard, Mary F. & Poeppel, David. 2009. Hemispheric asymmetry in mid and long latency neuromagnetic responses to single clicks. Hearing Research 257(1). 41–52. DOI:  http://doi.org/10.1016/j.heares.2009.07.010

Hussain, Qandeel & Proctor, Michael & Harvey, Mark & Demuth, Katherine. 2017. Acoustic characteristics of Punjabi retroflex and dental stops. The Journal of the Acoustical Society of America 141(6). 4522–4542. DOI:  http://doi.org/10.1121/1.4984595

Idsardi, William & Poeppel, David. 2011. Neurophysiological techniques in laboratory phonology. In Cohn, Abigail C. & Fougeron, Cécile & Huffman, Marie K. (eds.), The Oxford handbook of laboratory phonology, 593–605. Oxford University Press. DOI:  http://doi.org/10.1093/oxfordhb/9780199575039.013.0020

Jacobsen, Thomas Konstantin & Steinberg, Johanna & Truckenbrodt, Hubert & Jacobsen, Thomas. 2013. Mismatch Negativity (MMN) to successive deviants within one hierarchically structured auditory object. International Journal of Psychophysiology 87(1). 1–7. DOI:  http://doi.org/10.1016/j.ijpsycho.2012.09.012

Karnath, Hans-Otto & Sperber, Christoph & Rorden, Christopher. 2018. Mapping human brain lesions and their functional consequences. NeuroImage 165. 180–189. DOI:  http://doi.org/10.1016/j.neuroimage.2017.10.028

Kazanina, Nina & Bowers, Jeffrey S. & Idsardi, William J. 2018. Phonemes: Lexical access and beyond. Psychonomic Bulletin & Review 25(2). 560–585. DOI:  http://doi.org/10.3758/s13423-017-1362-0

Kazanina, Nina & Phillips, Colin & Idsardi, William J. 2006. The influence of meaning on the perception of speech sounds. Proceedings of the National Academy of Sciences 103(30). 11381–11386. DOI:  http://doi.org/10.1073/pnas.0604821103

Keating, Patricia A. 1996. The phonetics-phonology interface. In UCLA Working Papers in Phonetics, 45–60.

Kehrein, Wolfgang. 2013. Phonological representation and phonetic phasing: Affricates and laryngeals. Berlin: De Gruyter. DOI:  http://doi.org/10.1515/9783110911633

Lahiri, Aditi. 2018. Predicting universal phonological contrasts. In Hyman, Larry M. & Plank, Frans (eds.), Phonological typology, 229–272. Berlin: De Gruyter Mouton. Retrieved from https://www.degruyterbrill.com/document/doi/10.1515/9783110451931-007/html

Lahiri, Aditi & Reetz, Henning. 2002. Underspecified recognition. In Gussenhoven, Carlos & Warner, Natasha (eds.), Laboratory phonology, Vol. 7, 637–675. Berlin: Mouton de Gruyter.

Lahiri, Aditi & Reetz, Henning. 2010. Distinctive features: Phonological underspecification in representation and processing. Journal of Phonetics 38(1). 44–59. DOI:  http://doi.org/10.1016/j.wocn.2010.01.002

Lau, Ellen F. & Phillips, Colin & Poeppel, David. 2008. A cortical network for semantics: (De)constructing the N400. Nature Reviews Neuroscience 9(12). 920–933. DOI:  http://doi.org/10.1038/nrn2532

Lawyer, Laurel & Corina, David. 2014. An investigation of place and voice features using fMRI-adaptation. Journal of Neurolinguistics 27(1). 18–30. DOI:  http://doi.org/10.1016/j.jneuroling.2013.07.001

Li, Yuanning & Tang, Claire & Lu, Junfeng & Wu, Jinsong & Chang, Edward F. 2021. Human cortical encoding of pitch in tonal and non-tonal languages. Nature Communications 12(1). 1161. DOI:  http://doi.org/10.1038/s41467-021-21430-x

Lichtheim, Ludwig. 1885. On Aphasia. Brain 7(4). 433–484. DOI:  http://doi.org/10.1093/brain/7.4.433

Lindboom, Elsa & Nidiffer, Aaron & Carney, Laurel H. & Lalor, Edmund C. 2023. Incorporating models of subcortical processing improves the ability to predict EEG responses to natural speech. Hearing Research 433. 108767. DOI:  http://doi.org/10.1016/j.heares.2023.108767

Lisker, Leigh. 1977. Rapid versus rabid: A catalogue of acoustic features that may cue the distinction. The Journal of the Acoustical Society of America 62(S1). S77–S78. DOI:  http://doi.org/10.1121/1.2016377

Luck, Steven J. 2014. An introduction to the event-related potential technique, 2nd Edition. Cambridge, Massachusetts: MIT Press.

Macmillan, Neil A. & Creelman, C. Douglas. 2004. Detection theory: A user’s guide. Mahwah, NJ: Lawrence Erlbaum Associates, Inc.

Maddieson, Ian. 2005. Bilabial and labio-dental fricatives in Ewe. UC Berkeley Phonology Lab Annual Reports 1. DOI:  http://doi.org/10.5070/P74R49G6QX

Mai, Anna & Riès, Stephanie & Ben-Haim, Sharona & Shih, Jerry J. & Gentner, Timothy Q. 2024. Acoustic and language-specific sources for phonemic abstraction from speech. Nature Communications 15(1). 677. DOI:  http://doi.org/10.1038/s41467-024-44844-9

Manca, Anna Dora & Di Russo, Francesco & Sigona, Francesco & Grimaldi, Mirko. 2019. Electrophysiological evidence of phonemotopic representations of vowels in the primary and secondary auditory cortex. Cortex 121. 385–398. DOI:  http://doi.org/10.1016/j.cortex.2019.09.016

Marr, David. 1982. Vision: A computational investigation into the human representation and processing of visual information. San Francisco: W.H. Freeman.

Martin, Samuel E. 1992. A reference grammar of Korean: A complete guide to the grammar and history of the Korean language. Rutland, VT: C.E. Tuttle.

Meng, Yaxuan & Kotzor, Sandra & Xu, Chenzi & Wynne, Hilary S.Z. & Lahiri, Aditi. 2021a. Asymmetric influence of vocalic context on Mandarin sibilants: Evidence from ERP studies. Frontiers in Human Neuroscience 15. 617318. DOI:  http://doi.org/10.3389/fnhum.2021.617318

Meng, Yaxuan & Kotzor, Sandra & Xu, Chenzi & Wynne, Hilary S.Z. & Lahiri, Aditi. 2021b. Mismatch negativity (MMN) as an index of asymmetric processing of consonant duration in fake Mandarin geminates. Neuropsychologia 163. 108063. DOI:  http://doi.org/10.1016/j.neuropsychologia.2021.108063

Mesgarani, Nima & Cheung, Connie & Johnson, Keith & Chang, Edward F. 2014. Phonetic feature encoding in human superior temporal gyrus. Science 343(6174). 1006–1010. DOI:  http://doi.org/10.1126/science.1245994

Michel, Christoph M. & Brunet, Denis. 2019. EEG source imaging: A practical review of the analysis steps. Frontiers in Neurology 10. DOI:  http://doi.org/10.3389/fneur.2019.00325

Mielke, Jeff. 2008. The emergence of distinctive features. Oxford: Oxford University Press.

Millett, David. 2001. Hans Berger: From psychic energy to the EEG. Perspectives in Biology and Medicine 44(4). 522–542. DOI:  http://doi.org/10.1353/pbm.2001.0070

Minagawa-Kawai, Yasuyo & Mori, Koichi & Naoi, Nozomi & Kojima, Shozo. 2007. Neural attunement processes in infants during the acquisition of a language-specific phonemic contrast. Journal of Neuroscience 27(2). 315–321. DOI:  http://doi.org/10.1523/JNEUROSCI.1984-06.2007

Mitterer, Holger & Scharenborg, Odette & McQueen, James M. 2013. Phonological abstraction without phonemes in speech perception. Cognition 129(2). 356–361. DOI:  http://doi.org/10.1016/j.cognition.2013.07.011

Monahan, Philip J. 2018. Phonological knowledge and speech comprehension. Annual Review of Linguistics 4. 21–47. DOI:  http://doi.org/10.1146/annurev-linguistics-011817-045537

Monahan, Philip J. & Idsardi, William J. 2010. Auditory sensitivity to formant ratios: Toward an account of vowel normalisation. Language and Cognitive Processes 25(6). 808–839. DOI:  http://doi.org/10.1080/01690965.2010.490047

Monahan, Philip J. & Lau, Ellen F. & Idsardi, William J. 2013. Computational primitives in phonology and their neural correlates. In Boeckx, Cedric & Grohmann, Kleanthes K. (eds.), The Cambridge Handbook of Biolinguistics, 233–256. Cambridge: Cambridge University Press. DOI:  http://doi.org/10.1017/CBO9780511980435.015

Monahan, Philip J. & Schertz, Jessamyn & Fu, Zhanao & Pérez, Alejandro. 2022. Unified coding of spectral and temporal phonetic cues: Electrophysiological evidence for abstract phonological features. Journal of Cognitive Neuroscience 34(4). 618–638. DOI:  http://doi.org/10.1162/jocn_a_01817

Moon, Hyun Seok & Jiang, Haiyan & Vo, Thanh Tan & Jung, Won Beom & Vazquez, Alberto L & Kim, Seong-Gi. 2021. Contribution of excitatory and inhibitory neuronal activity to BOLD fMRI. Cerebral Cortex 31(9). 4053–4067. DOI:  http://doi.org/10.1093/cercor/bhab068

Morales, Santiago & Bowers, Maureen E. 2022. Time-frequency analysis methods and their application in developmental EEG data. Developmental Cognitive Neuroscience 54. 101067. DOI:  http://doi.org/10.1016/j.dcn.2022.101067

Mosso, Angelo. 2014. Angelo Mosso’s circulation of blood in the human brain. New York, NY: Oxford University Press.

Möttönen, Riikka & Watkins, Kate E. 2012. Using TMS to study the role of the articulatory motor system in speech perception. Aphasiology 26(9). 1103–1118. DOI:  http://doi.org/10.1080/02687038.2011.619515

Murakami, Takenobu & Ugawa, Yoshikazu & Ziemann, Ulf. 2013. Utility of TMS to understand the neurobiology of speech. Frontiers in Psychology 4. DOI:  http://doi.org/10.3389/fpsyg.2013.00446

Murthy, Venkatesh N. 2011. Olfactory maps in the brain. Annual Review of Neuroscience 34(Volume 34, 2011). 233–258. DOI:  http://doi.org/10.1146/annurev-neuro-061010-113738

Myers, Emily B. & Blumstein, Sheila E. & Walsh, Edward & Eliassen, James. 2009. Inferior frontal regions underlie the perception of phonetic category invariance. Psychological Science 20(7). 895–903. DOI:  http://doi.org/10.1111/j.1467-9280.2009.02380.x

Näätänen, Risto. 2001. The perception of speech sounds by the human brain as reflected by the mismatch negativity (MMN) and its magnetic equivalent (MMNm). Psychophysiology 38(1). 1–21. DOI:  http://doi.org/10.1111/1469-8986.3810001

Näätänen, Risto & Kujala, Teija & Light, Gregory. 2019. Mismatch negativity: A window to the brain. Oxford, UK: Oxford University Press.

Näätänen, Risto & Lehtokoski, Anne & Lennes, Mietta & Cheour, Marie & Huotilainen, Minna & Iivonen, Antti & Vainio, Martti & Alku, Paavo & Ilmoniemi, Risto J. & Luuk, Aavo & Allik, Jüri & Sinkkonen, Janne & Alho, Kimmo. 1997. Language-specific phoneme representations revealed by electric and magnetic brain responses. Nature 385(6615). 432–434. DOI:  http://doi.org/10.1038/385432a0

Näätänen, Risto & Paavilainen, Petri & Rinne, Teemu & Alho, Kimmo. 2007. The mismatch negativity (MMN) in basic research of central auditory processing: A review. Clinical Neurophysiology 118(12). 2544–2590. DOI:  http://doi.org/10.1016/j.clinph.2007.04.026

Obleser, Jonas & Boecker, Henning & Drzezga, Alexander & Haslinger, Bernhard & Hennenlotter, Andreas & Roettinger, Michael & Eulitz, Carsten & Rauschecker, Josef P. 2006. Vowel sound extraction in anterior superior temporal cortex. Human Brain Mapping 27(7). 562–571. DOI:  http://doi.org/10.1002/hbm.20201

Obleser, Jonas & Eisner, Frank. 2009. Pre-lexical abstraction of speech in the auditory cortex. Trends in Cognitive Sciences 13(1). 14–19. DOI:  http://doi.org/10.1016/j.tics.2008.09.005

Obleser, Jonas & Lahiri, Aditi & Eulitz, Carsten. 2003. Auditory-evoked magnetic field codes place of articulation in timing and topography around 100 milliseconds post syllable onset. NeuroImage 20(3). 1839–1847. DOI:  http://doi.org/10.1016/j.neuroimage.2003.07.019

Obleser, Jonas & Lahiri, Aditi & Eulitz, Carsten. 2004. Magnetic brain response mirrors extraction of phonological features from spoken vowels. Journal of Cognitive Neuroscience 16(1). 31–39. DOI:  http://doi.org/10.1162/089892904322755539

Oganian, Yulia & Bhaya-Grossman, Ilina & Johnson, Keith & Chang, Edward F. 2023. Vowel and formant representation in the human auditory speech cortex. Neuron 111(13). 2105–2118.e4. DOI:  http://doi.org/10.1016/j.neuron.2023.04.004

Okada, Kayoko & Hickok, Gregory. 2006. Identification of lexical-phonological networks in the superior temporal sulcus using functional magnetic resonance imaging. NeuroReport 17(12). 1293–1296.

Okada, Kayoko & Rong, Feng & Venezia, Jon & Matchin, William & Hsieh, I-Hui & Saberi, Kourosh & Serences, John T. & Hickok, Gregory. 2010. Hierarchical organization of human auditory cortex: Evidence from acoustic invariance in the response to intelligible speech. Cerebral Cortex 20(10). 2486–2495. DOI:  http://doi.org/10.1093/cercor/bhp318

Paavilainen, Petri & Valppu, Sanna & Näätänen, Risto. 2001. The additivity of the auditory feature analysis in the human brain as indexed by the mismatch negativity: 1+1≈2 but 1+1+1<3. Neuroscience Letters 301(3). 179–182. DOI:  http://doi.org/10.1016/S0304-3940(01)01635-4

Petersen, Steven E. & Fox, Peter T. & Posner, Michael I. & Mintun, Mark & Raichle, Marcus E. 1989. Positron emission tomographic studies of the processing of singe words. Journal of Cognitive Neuroscience 1(2). 153–170. DOI:  http://doi.org/10.1162/jocn.1989.1.2.153

Phillips, Colin. 2001. Levels of representation in the electrophysiology of speech perception. Cognitive Science 25. 711–731.

Phillips, Colin & Pellathy, Thomas & Marantz, Alec & Yellin, Elron & Wexler, Kenneth & Poeppel, David & McGinnis, Martha & Roberts, Timothy. 2000. Auditory cortex accesses phonological categories: An MEG mismatch study. Journal of Cognitive Neuroscience 12(6). 1038–1055. DOI:  http://doi.org/10.1162/08989290051137567

Poeppel, David. 1996. A critical review of PET studies of phonological processing. Brain and Language 55(3). 317–351. DOI:  http://doi.org/10.1006/brln.1996.0108

Poeppel, David & Embick, David. 2017. Defining the relation between linguistics and neuroscience. In Cutler, Anne (ed.), Twenty-first century psycholinguistics, 103–118. Routledge. DOI:  http://doi.org/10.4324/9781315084503-8

Poeppel, David & Emmorey, Karen & Hickok, Gregory & Pylkkänen, Liina. 2012. Towards a new neurobiology of language. Journal of Neuroscience 32(41). 14125–14131. DOI:  http://doi.org/10.1523/JNEUROSCI.3244-12.2012

Poeppel, David & Idsardi, William. 2022. We don’t know how the brain stores anything, let alone words. Trends in Cognitive Sciences 26(12). 1054–1055. DOI:  http://doi.org/10.1016/j.tics.2022.08.010

Poeppel, David & Idsardi, William J. & van Wassenhove, Virginie. 2008. Speech perception at the interface of neurobiology and linguistics. Philosophical Transactions of the Royal Society B: Biological Sciences 363(1493). 1071–1086. DOI:  http://doi.org/10.1098/rstb.2007.2160

Poldrack, Russell A. & Mumford, Jeanette A. & Nichols, Thomas E. 2011. Handbook of functional MRI data analysis. Cambridge University Press.

Politzer-Ahles, Stephen & Jap, Bernard A. J. 2024. Can the mismatch negativity really be elicited by abstract linguistic contrasts? Neurobiology of Language 5(4). 818–843. DOI:  http://doi.org/10.1162/nol_a_00147

Politzer-Ahles, Stephen & Schluter, Kevin T. & Wu, Kefei & Almeida, Diogo. 2016. Asymmetries in the perception of Mandarin tones: Evidence from mismatch negativity. Journal of Experimental Psychology: Human Perception and Performance 42(10). 1547–1570.

Reiss, Charles. 2017. Substance free phonology. In Hannahs, S. J. & Bosch, Anna R. K. (eds.), The Routledge handbook of phonological theory, 1st ed., 425–452. London: Routledge. DOI:  http://doi.org/10.4324/9781315675428-15

Reiss, Charles. 2022. Priority union and feature logic in phonology. Linguistic Inquiry 53(1). 199–209. DOI:  http://doi.org/10.1162/ling_a_00400

Rhodes, Ryan & Avcu, Enes & Han, Chao & Hestvik, Arild. 2022. Auditory predictions are phonological when phonetic information is variable. Language, Cognition and Neuroscience 37(9). 1099–1114. DOI:  http://doi.org/10.1080/23273798.2022.2043395

Roberts, Timothy P. L. & Flagg, Elissa J. & Gage, Nicole M. 2004. Vowel categorization induces departure of M100 latency from acoustic prediction. NeuroReport 15(10). 1679–1682. DOI:  http://doi.org/10.1097/01.wnr.0000134928.96937.10

Rorden, Chris & Karnath, Hans-Otto. 2004. Using human brain lesions to infer function: A relic from a past era in the fMRI age? Nature Reviews Neuroscience 5(10). 812–819. DOI:  http://doi.org/10.1038/nrn1521

Roy, Charles S. & Sherrington, Charles S. 1890. On the regulation of the blood-supply of the brain. The Journal of Physiology 11(1–2). 85–158.17. DOI:  http://doi.org/10.1113/jphysiol.1890.sp000321

Salter, Katherine & Jutai, Jeffrey & Foley, Norine & Hellings, Chelsea & Teasell, Robert. 2006. Identification of aphasia post stroke: A review of screening assessment tools. Brain Injury 20(6). 559–568. DOI:  http://doi.org/10.1080/02699050600744087

Sams, Mikko & Kaukoranta, Elina & Hämäläinen, Matti & Näätänen, Risto. 1991. Cortical activity elicited by changes in auditory stimuli: Different sources for the magnetic N100m and mismatch responses. Psychophysiology 28(1). 21–29. DOI:  http://doi.org/10.1111/j.1469-8986.1991.tb03382.x

Samuel, Arthur G. 2020. Psycholinguists should resist the allure of linguistic units as perceptual units. Journal of Memory and Language 111. 104070. DOI:  http://doi.org/10.1016/j.jml.2019.104070

Sandrone, Stefano & Bacigaluppi, Marco & Galloni, Marco R. & Cappa, Stefano F. & Moro, Andrea & Catani, Marco & Filippi, Massimo & Monti, Martin M. & Perani, Daniela & Martino, Gianvito. 2014. Weighing brain activity with the balance: Angelo Mosso’s original manuscripts come to light. Brain 137(2). 621–633. DOI:  http://doi.org/10.1093/brain/awt091

Scharinger, Mathias & Domahs, Ulrike & Klein, Elise & Domahs, Frank. 2016. Mental representations of vowel features asymmetrically modulate activity in superior temporal sulcus. Brain and Language 163. 42–49. DOI:  http://doi.org/10.1016/j.bandl.2016.09.002

Scharinger, Mathias & Idsardi, William J. & Poe, Samantha. 2011. A comprehensive three-dimensional cortical map of vowel space. Journal of Cognitive Neuroscience 23(12). 3972–3982. DOI:  http://doi.org/10.1162/jocn_a_00056

Scharinger, Mathias & Monahan, Philip J. & Idsardi, William J. 2012. Asymmetries in the Processing of Vowel Height. Journal of Speech, Language, and Hearing Research 55(3). 903–918. DOI:  http://doi.org/10.1044/1092-4388(2011/11-0065)

Scharinger, Mathias & Monahan, Philip J. & Idsardi, William J. 2016. Linguistic category structure influences early auditory processing: Converging evidence from mismatch responses and cortical oscillations. NeuroImage 128. 293–301. DOI:  http://doi.org/10.1016/j.neuroimage.2016.01.003

Schluter, Kevin T. & Politzer-Ahles, Stephen & Al Kaabi, Meera & Almeida, Diogo. 2017. Laryngeal features are phonetically abstract: Mismatch negativity evidence from Arabic, English, and Russian. Frontiers in Psychology 8. 746. DOI:  http://doi.org/10.3389/fpsyg.2017.00746

Schluter, Kevin T. & Politzer-Ahles, Stephen & Almeida, Diogo. 2016. No place for /h/: An ERP investigation of English fricative place features. Language, Cognition and Neuroscience 31(6). 728–740. DOI:  http://doi.org/10.1080/23273798.2016.1151058

Scott, Sophie K. 2019. From speech and talkers to the social world: The neural processing of human spoken language. Science 366(6461). 58–62. DOI:  http://doi.org/10.1126/science.aax0288

Sergent, Justine & Zuck, Eric & Lévesque, Michel & MacDonald, Brennan. 1992. Positron emission tomography study of letter and object processing: Empirical findings and methodological considerations. Cerebral Cortex 2(1). 68–80. DOI:  http://doi.org/10.1093/cercor/2.1.68

Sharma, Anu & Dorman, Michael F. 1999. Cortical auditory evoked potential correlates of categorical perception of voice-onset time. The Journal of the Acoustical Society of America 106(2). 1078–1083. DOI:  http://doi.org/10.1121/1.428048

Sharma, Anu & Marsh, Catherine M. & Dorman, Michael F. 2000. Relationship between N1 evoked potential morphology and the perception of voicing. The Journal of the Acoustical Society of America 108(6). 3030–3035. DOI:  http://doi.org/10.1121/1.1320474

Smith, Caroline L. 1997. The devoicing of /z/ in American English: Effects of local and prosodic context. Journal of Phonetics 25(4). 471–500. DOI:  http://doi.org/10.1006/jpho.1997.0053

Sotero, Roberto C. & Trujillo-Barreto, Nelson J. 2007. Modelling the role of excitatory and inhibitory neuronal activity in the generation of the BOLD signal. NeuroImage 35(1). 149–165. DOI:  http://doi.org/10.1016/j.neuroimage.2006.10.027

Speck, Iva & Rottmayer, Valentin & Wiebe, Konstantin & Aschendorff, Antje & Thurow, Johannes & Frings, Lars & Meyer, Philipp T. & Wesarg, Thomas & Arndt, Susan. 2021. PET/CT background noise and its effect on speech recognition. Scientific Reports 11(1). 22065. DOI:  http://doi.org/10.1038/s41598-021-01686-5

Stevens, Kenneth N. 1998. Acoustic phonetics. The MIT Press. DOI:  http://doi.org/10.7551/mitpress/1072.001.0001

Sun, Yue & Poeppel, David. 2023. Syllables and their beginnings have a special role in the mental lexicon. Proceedings of the National Academy of Sciences 120(36). e2215710120. DOI:  http://doi.org/10.1073/pnas.2215710120

Talavage, Thomas M. & Gonzalez-Castillo, Javier & Scott, Sophie K. 2014. Auditory neuroimaging with fMRI and PET. Hearing Research 307. 4–15. DOI:  http://doi.org/10.1016/j.heares.2013.09.009

Vanhaudenhuyse, Audrey & Laureys, Steven & Perrin, Fabien. 2008. Cognitive event-related potentials in comatose and post-comatose states. Neurocritical Care 8(2). 262–270. DOI:  http://doi.org/10.1007/s12028-007-9016-0

Wernicke, Carl. 1874. Der aphasische Symptomencomplex: eine psychologische Studie auf anatomischer Basis. Cohn & Weigert.

Wijeakumar, Sobanawartiny & Huppert, Theodore J. & Magnotta, Vincent A. & Buss, Aaron T. & Spencer, John P. 2017. Validating an image-based fNIRS approach with fMRI and a working memory task. NeuroImage 147. 204–218. DOI:  http://doi.org/10.1016/j.neuroimage.2016.12.007

Wilcox, Teresa & Biondi, Marisa. 2015. fNIRS in the developmental sciences. WIREs Cognitive Science 6(3). 263–283. DOI:  http://doi.org/10.1002/wcs.1343

Winkler, István. 2007. Interpreting the mismatch negativity. Journal of Psychophysiology 21(3–4). 147–163. DOI:  http://doi.org/10.1027/0269-8803.21.34.147

Winkler, István & Lehtokoski, Anne & Alku, Paavo & Vainio, Martti & Czigler, István & Csépe, Valéria & Aaltonen, Olli & Raimo, Ilkka & Alho, Kimmo & Lang, Heikki & Iivonen, Antti & Näätänen, Risto. 1999. Pre-attentive detection of vowel contrasts utilizes both phonetic and auditory memory representations. Cognitive Brain Research 7(3). 357–369. DOI:  http://doi.org/10.1016/S0926-6410(98)00039-1

Wolff, Christian & Schröger, Erich. 2001. Human pre-attentive auditory change-detection with single, double, and triple deviations as revealed by mismatch negativity additivity. Neuroscience Letters 311(1). 37–40. DOI:  http://doi.org/10.1016/S0304-3940(01)02135-8

Yang, Zinong & Lewis, Laura D. 2021. Imaging the temporal dynamics of brain states with highly sampled fMRI. Current Opinion in Behavioral Sciences 40. 87–95. DOI:  http://doi.org/10.1016/j.cobeha.2021.02.005

Yi, Han Gyol & Leonard, Matthew K. & Chang, Edward F. 2019. The encoding of speech sounds in the superior temporal gyrus. Neuron 102(6). 1096–1110. DOI:  http://doi.org/10.1016/j.neuron.2019.04.023

Yu, Yan H. & Shafer, Valerie L. 2021. Neural representation of the English vowel feature [high]: Evidence from /ε/ vs. /ɪ/. Frontiers in Human Neuroscience 15. 629517. DOI:  http://doi.org/10.3389/fnhum.2021.629517

Zatorre, Robert J. & Evans, Alan C. & Meyer, Ernst & Gjedde, Albert. 1992. Lateralization of phonetic and pitch discrimination in speech processing. Science 256(5058). 846–849. DOI:  http://doi.org/10.1126/science.256.5058.846