1 Introduction

The internal structure of Verb Phrase (VP) has been a focal point of research throughout the history of generative grammar. The most widely accepted view, dating back to Larson (1988), posits that VP is decomposed into at least two layers: a higher vP (or VoiceP), which selects an external argument (e.g., Chomsky 1995; 2001; Kratzer 1996), and a lower lexical VP, which introduces an internal argument. More recently, a growing body of research has argued for a more complex, three-layered Voice-v-V structure (e.g., Pylkkänen 2002; 2008; Alexiadou et al. 2006; 2015; Ramchand 2008; Harley 2013; Legate 2014; the chapters of D’Alessandro et al. 2017 and references therein). While the specific formalizations vary among studies, this tripartite model has gained significant traction. For example, the structure of change-of-state causative verbs can be represented as follows:

    1. (1)

In (1), Voice introduces an external argument (typically an Agent) at its spec position, the little v encodes causation and functions as a verbalizer, and VRes encodes a result state and selects an internal argument (for expository reasons, I use V(P)Res for this V(P), distinguishing it from the traditional VP and emphasizing its result-state role). Building on this theoretical background, this paper argues for the tripartite VP structure on the basis of Japanese verbal phenomena. It focuses on two structural distinctions within the verbal domain: (i) the demarcation between VRes and v, and (ii) the separation of v and Voice. These distinctions enable us to account for several otherwise puzzling patterns in Japanese morphosyntax.

The discussion proceeds as follows. Section 2 examines Japanese VP anaphora and shows that transitivity mismatches support a distinction between vP and VPRes. Section 3 explores the asymmetric distribution of the light verbs su and nar, arguing for a split between v and Voice. Section 4 compares VP anaphora and VP ellipsis, providing further support for the proposed tripartite structure. Section 5 synthesizes the results into a typology of Japanese verbal domains. Section 6 concludes the paper.

2 Split between VRes and v: Japanese VP anaphora

This section explores the decomposition of VP into VRes (encoding result states) and vP (representing causal relations) through an analysis of Japanese VP anaphora (e.g., Nakau 1973; Hinds 1973; Tateishi 1994; Tanaka 2016). A representative example of Japanese VP anaphora is illustrated in (2):

    1. (2)
    1. Taro-ga
    2. Taro-nom
    1. [VP
    2.  
    1. ringo-o
    2. apple-acc
    1. tabe]-ta-node,
    2. eat-pst-because
    1. Ziro-mo
    2. Ziro-also
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1. ‘Because Taro ate an apple, Ziro did so, too.’

While traditional analyses since Nakau (1973) treated soo su as replacing the entire VP, Tanaka (2016) offers a compositional analysis where soo and su function as separate syntactic units.1

Following this compositional approach, I focus on VP anaphora with causative/inchoative verbs, demonstrating that Japanese permits transitivity mismatches in such contexts. I argue these mismatches follow if the anaphoric element soo copies the denotation of the antecedent VPRes, building on Bruening’s (2019) analysis of English do so.

2.1 Japanese VP anaphora with causative/inchoative verbs

Japanese has two pro-forms, soo su and soo nar, whose distribution correlates with antecedent transitivity:

    1. (3)
    1. a.
    1. Taro-ga
    2. Taro-nom
    1. doa-o
    2. door-acc
    1. sim-e-ta-ato,
    2. close-caus-pst-after
    1. Ziro-mo
    2. Ziro-also
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1. ‘After Taro closed the door, Ziro did so, too.’
    1.  
    1. b.
    1. Doa-ga
    2. door-nom
    1. sim-ar-ta-ato,
    2. close-inch-pst-after
    1. tonari-no
    2. next-gen
    1. doa-mo
    2. door-also
    1. soo
    2. so
    1. nar-ta.
    2. become-pst
    1. ‘After the door closed, the next door did so, too.’

In (3a), soo su refers to the transitive VP ‘close the door’ headed by a causative verb sim-e, while in (3b), soo nar refers to the intransitive VP with inchoative verb sim-ar. The reverse pattern is ungrammatical (under the interpretation where the pro-form matches the transitivity of its antecedent):2

    1. (4)
    1. a.
    1. *Taro-ga
    2.   Taro-nom
    1. doa-o
    2. door-acc
    1. sim-e-ta-ato,
    2. close-caus-pst-after
    1. Ziro-mo
    2. Ziro-also
    1. soo
    2. so
    1. nar-ta.
    2. become-pst
    1.   ‘After Taro closed the door, Ziro did so, too.’
    1.  
    1. b.
    1. *Doa-ga
    2.   door-nom
    1. sim-ar-ta-ato,
    2. close-inch-pst-after
    1. tonari-no
    2. next-gen
    1. doa-mo
    2. door-also
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1.   ‘After the door closed, the next door did so, too.’

This pattern admits two possible explanations. Under the first account, the pro-form must match the valency of its antecedent VP: if the antecedent is dyadic (or monadic), the target must also be dyadic (or monadic). This would explain why (4a) and (4b) are ungrammatical—they violate argument structure matching between antecedent and target. Under the second account, pro-form selection depends solely on the target clause’s own argument structure (soo su when an external argument is present, soo nar when absent), independently of the antecedent. On this view, (4a) and (4b) are ungrammatical because the wrong pro-form is selected for the target’s argument structure, not because of antecedent-target mismatch. These two accounts make different predictions. The first predicts that transitivity mismatches between antecedent and target should be impossible. The second predicts that such mismatches are possible in principle, provided the correct pro-form matches the target clause’s structure. The following subsection demonstrates that the latter prediction is borne out.

2.2 Transitivity mismatches

Before examining Japanese data, let us consider English VP anaphora with do so. Research has shown that English do so permits various identity mismatches, including transitivity mismatches between causative and inchoative constructions (example (a) from Bouton (1969: 246) and examples (b)–(c) from Houser (2010: 20)):

    1. (5)
    1. a.
    1. We planned to sink the boat, but we didn’t want it to do so while anyone was on it.
    1.  
    1. b.
    1. John told Steve to hang the horseshoe over the door, and it does so now.
    1.  
    1. c.
    1. Mary claimed that I closed the door, but it actually did so on its own.

In these examples, do so is interpreted as encoding inchoative events, despite the antecedent VPs containing causative verbs—a causative-to-inchoative mismatch.

Crucially, the reverse pattern—where the antecedent is inchoative and do so is interpreted as causative—is ungrammatical:

    1. (6)
    1. Houser (2010: 20)
    1.  
    1. a.
    1. *John wanted the horseshoe to hang over the door, so Steve did so.
    1.  
    1. b.
    1. *A bunch of books burned last night, and I heard that John did so.
    1.  
    1. c.
    1. *Mary claimed that the door closed on its own, but I actually did so.

These examples, with inchoative antecedents (hang, burn, close) and do so intended as causative, demonstrate a clear asymmetry: while do so can refer to an inchoative event when the antecedent is causative, it cannot refer to a causative event when the antecedent is inchoative.

Let us now turn to Japanese. Despite having morphologically distinct pro-forms, Japanese also allows causative-to-inchoative mismatches, similar to English:

    1. (7)
    1. a.
    1. Watasitati-wa
    2. we-top
    1. fune-o
    2. ship-acc
    1. sizum-e-yootosi-ta-ga,
    2. sink-caus-try-pst-but
    1. hito-ga
    2. person-nom
    1. iru-tokini
    2. be-when
    1. soo
    2. so
    1. nar-te
    2. become-ger
    1. hosiku-wa
    2. want-top
    1. nakat-ta.
    2. neg-pst
    1. ‘We tried to sink the boat, but we didn’t want it to do so while anyone was on it.’
    1.  
    1. b.
    1. John-ga
    2. John-nom
    1. Mary-ni
    2. Mary-dat
    1. taoru-o
    2. towel-acc
    1. doa-ni
    2. door-to
    1. kak-e-ru-yooni
    2. hang-caus-prs-imp
    1. iw-ta.
    2. say-pst
    1. Sosite,
    2. and
    1. ima
    2. now
    1. soo
    2. so
    1. nar-te
    2. become-ger
    1. i-ru.
    2. be-prs
    1. ‘John told Mary to hang the towel over the door. And it does so now.’
    1.  
    1. c.
    1. Mary-wa
    2. Mary-top
    1. watasi-ga
    2. I-nom
    1. doa-o
    2. door-acc
    1. sim-e-ta-to
    2. close-caus-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. katteni
    2. on.one’s.own
    1. soo
    2. so
    1. nar-ta.
    2. become-pst
    1. ‘Mary claimed that I closed the door, but it actually did so on its own.’

These examples parallel the English cases in (5). The antecedent clauses contain causative verbs (marked by -e morphology), while the target clauses use soo nar to encode inchoative events.3

Japanese VP anaphora also permits the reverse pattern that is ungrammatical in English: inchoative-to-causative mismatches.4 To illustrate this, consider:

    1. (8)
    1. a.
    1. Kinou
    2. yesterday
    1. fune-ga
    2. ship-nom
    1. sizum-Ø-ta-ga,
    2. sink-inch-pst-but
    1. zitu-wa
    2. actually-top
    1. watasitati-ga
    2. we-nom
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1. ‘Lit.: Yesterday, the boat sank, but actually we did so.’
    1.  
    1. b.
    1. Sakuban
    2. last.night
    1. tairyoo-no
    2. bunch-gen
    1. hon-ga
    2. book-nom
    1. moy-e-ta-ga,
    2. burn-inch-pst-but
    1. watasi-wa
    2. I-top
    1. John-ga
    2. John-nom
    1. soo
    2. so
    1. si-ta-to
    2. do-pst-comp
    1. kik-ta.
    2. hear-pst
    1. ‘Lit.: Last night, a bunch of books burned, but I heard that John did so.’
    1.  
    1. c.
    1. Mary-wa
    2. Mary-top
    1. katteni
    2. on.one’s.own
    1. doa-ga
    2. door-nom
    1. sim-ar-ta-to
    2. close-inch-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. watasi-ga
    2. I-nom
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1. ‘Lit.: Mary claimed that the door closed on its own, but actually I did so.’

Here, soo su encodes a causative event despite referring to inchoative verbs in the antecedent clauses (sizum-Ø ‘sink’, moy-e ‘burn’ and sim-ar ‘close’). This confirms that Japanese VP anaphora permits inchoative-to-causative mismatches, a pattern unavailable in English.5

In these transitivity-mismatch cases, the target pro-forms are construed as inchoative or causative variants of the antecedent verbs. Examples (7c) and (8c), for instance, are semantically equivalent to sentences using the corresponding inchoative or causative variants directly:

    1. (9)
    1. a.
    1. Mary-wa
    2. Mary-top
    1. watasi-ga
    2. I-nom
    1. doa-o
    2. door-acc
    1. sim-e-ta-to
    2. close-caus-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. katteni
    2. on.one’s.own
    1. sim-ar-ta.
    2. close-inch-pst
    1. ‘Mary claimed that I closed the door, but it actually did so on its own.’
    1.  
    1. b.
    1. Mary-wa
    2. Mary-top
    1. katteni
    2. on.one’s.own
    1. doa-ga
    2. door-nom
    1. sim-ar-ta-to
    2. close-inch-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. watasi-ga
    2. I-nom
    1. sim-e-ta.
    2. close-caus-pst
    1. ‘Mary claimed that the door closed on its own, but actually I did so.’

This suggests that despite transitivity discrepancies, the pro-form maintains a semantic link to its antecedent. A plausible analysis, consistent with Tanaka (2016), is that (a) soo itself, rather than the complex form soo su or soo nar, is anaphoric to the preceding event and refers back to a result state encoded by the antecedent verb, and (b) the verbal elements su or nar are merely morphological realizations determined by the transitivity of the target clause.

The following subsection formalizes this analysis by adapting and extending Bruening’s (2019) framework for English do so anaphora, showing how soo targets the result state encoded in VRes and how the verbal elements su/nar realize vcaus/inch in the structure.

2.3 Deriving VP anaphora in transitivity mismatches

2.3.1 Bruening’s (2019) analysis of English do so

Bruening (2019) proposes an analysis of English do so anaphora that accounts for its compatibility with A-movement contexts including unaccusatives and passives. His key insight is that do so consists of the intransitive verb do and an adjunct so (cf. Bouton 1969; Ross 1972; Houser 2010), denoting λe.f(e), where function f is replaced by finding a linguistic antecedent (Hankamer & Sag 1976).

For transitive VP antecedents as in (10), Bruening’s analysis works as follows:

    1. (10)
    1. You need to decorate the eggs, but before you do so, … (Bruening 2019: 31(94))
    1.  
    1. a.
    1.  
    1. (i)
    1. ⟦VP⟧= λe.decorate(e,the eggs)
    1.  
    1. (ii)
    1. ⟦Voicetr⟧= λf⟨v,t⟩λxλe.f(e) ∧ Init(e,x)
    1.  
    1. (iii)
    1. ⟦VoiceP⟧= λe.decorate(e,the eggs) ∧ Init(e,you)
    1.  
    1. b.
    1.  
    1. (i)
    1. ⟦VPd⟧= λe.f(e)
    1.  
    1. (ii)
    1. Copy: ⟦VPd⟧ = λe.decorate(e,the eggs)
    1.  
    1. (iii)
    1. ⟦VoiceP⟧= λe.decorate(e,the eggs) ∧ Init(e,you)

In the antecedent clause (10a), transitive Voice (Voicetr) introduces an external argument with an Initiator role (agents/causers) at its specifier position. In the target clause (10b), VPd (do so) initially denotes λe.f(e), then function f is replaced by the semantic value of the antecedent VP, yielding λe.decorate(e, the eggs). The remaining semantic composition proceeds as in the antecedent structure.

Another crucial component of Bruening’s analysis is the A-Movement Rule, which handles passives and unaccusatives:

    1. (11)
    1. A-Movement Rule (Bruening 2019: 28)
    2. If α is a branching node whose daughters are β and a head H such that H is an A-movement head with index 1, then ⟦α⟧g = λx∈De.⟦H⟧(⟦β⟧g(1→x))

Bruening identifies unaccusative Voice and passive Voice as A-movement heads under this rule. This rule serves as a mechanism for lambda abstraction in A-movement contexts, allowing elements that have undergone A-movement to be properly interpreted in their derived positions.

To illustrate this rule’s application, consider an unaccusative antecedent:

    1. (12)
    1. The towels dried, but before they did so, … (Bruening 2019: 32(95))
    1.  
    1. a.
    1. ⟦VP⟧= λe.dry(e,t1)
    1.  
    1. b.
    1. ⟦Voiceun⟧= λ fλe.f(e)
    1.  
    1. c.
    1. ⟦VoiceP(1)⟧= λxλe.dry(e,x) (by the A-Movement Rule)
    1.  
    1. d.
    1. ⟦VoiceP(2)⟧= λe.dry(e,the towels)

Unaccusative Voice (Voiceun) differs from transitive Voice in that it introduces no Initiator (12b). As an A-movement head, it triggers the A-Movement Rule, lambda-abstracting over the VP-internal trace (12c). Syntactically, Voiceun induces A-movement of the underlying object NP to the specifier position of VoiceP, where semantically it saturates the lambda-abstracted variable (12d).

Based on this process, the corresponding do so sentence derives as follows:

    1. (13)
    1. The towels dried, but before they did so, … (Bruening 2019: 32(95))
    1.  
    1. a.
    1. ⟦VPd⟧= λe.f(e)
    1.  
    1. b.
    1. Copy: ⟦VPd⟧ = λe.dry(e,t2)
    1.  
    1. c.
    1. ⟦VoiceP(1)⟧= λxλe.dry(e,x) (by the A-Movement Rule)
    1.  
    1. d.
    1. ⟦VoiceP(2)⟧= λe.dry(e,they)

The copied denotation here includes a trace. Voiceun triggers the A-Movement Rule, creating lambda abstraction over the trace. VoiceP(1) becomes a predicate of individuals applying to the NP in Spec,VoiceP (13d). This is how a base-generated subject of do so is interpreted as a logical object when the antecedent supplies a trace-containing function.

In summary, Bruening’s analysis provides a unified account for do so anaphora in both transitive and unaccusative constructions through three components: (i) semantic copying, (ii) the A-Movement Rule, and (iii) distinct Voice heads.

2.3.2 Japanese VP anaphora and transitivity mismatches

Building upon Bruening’s (2019) framework, we now analyze Japanese VP anaphora with focus on transitivity mismatches. As established in section 2.2, Japanese exhibits both causative-to-inchoative and inchoative-to-causative mismatches, suggesting that soo anaphorically refers to a result state encoded in the antecedent, while su/nar are morphological realizations determined by the transitivity of the target clause.

To account for these observations, I propose decomposing the traditional VP into VPRes (encoding a result state) and vP (encoding causal relation). The little v head bears either an inchoative or causative semantic value:

    1. (14)
    1. a.
    1. ⟦vinch⟧= λf⟨s,t⟩λe∃s.caus(e,s) ∧ f(s)
    1.  
    1. b.
    1. ⟦vcaus⟧= λfs,tλe∃s.activity(e) ∧ caus(e,s) ∧ f(s)

This analysis posits a causative relation between an event variable e and a state variable s in both lexical causatives and inchoatives, consistent with Kratzer (2005) and Schäfer (2025). The key distinction is that causative v includes an activity predicate, which is absent in inchoative v, capturing the semantic difference between these event types.

Under these proposals, the causative and inchoative variants with the root sim- ‘close’ are analyzed as follows:

    1. (15)
    1. Causative6
    1.  
    1. a.
    1. ⟦VPRes⟧= λs.close(s,the door)
    1.  
    1. b.
    1. ⟦vcaus⟧= λ f⟨s,t⟩ λe∃s.activity(e) ∧ caus(e,s) ∧ f(s)
    1.  
    1. c.
    1. ⟦vPcaus⟧= λe∃s.activity(e) ∧ caus(e,s) ∧ close(s,the door)
    1.  
    1. d.
    1. ⟦Voicetr⟧= λ f⟨v,t⟩λxλe.f(e) ∧ Init(e,x)
    1.  
    1. e.
    1. ⟦VoiceP⟧= λe∃s.activity(e) ∧ caus(e,s) ∧ close(s,the door) ∧ Init(e,Taro)
    1. (16)
    1. Inchoative
    1.  
    1. a.
    1. ⟦VPRes⟧= λs.close(s,t1)
    1.  
    1. b.
    1. ⟦vinch⟧= λ f⟨s,t⟩ λe∃s.caus(e,s) ∧ f(s)
    1.  
    1. c.
    1. ⟦vPinch⟧= λe∃s.caus(e,s) ∧ close(s,t1)
    1.  
    1. d.
    1. ⟦Voiceun⟧= λ fλe.f(e)
    1.  
    1. e.
    1. ⟦VoiceP(1)⟧= λxλe∃s.caus(e,s) ∧ close(s,x) (by the A-Movement Rule)
    1.  
    1. f.
    1. ⟦VoiceP(2)⟧= λe∃s.caus(e,s) ∧ close(s,the door)

These structures reveal key distinctions between the variants: transitivity-dependent morphology realizes the v head (Akimoto 2018), with vcaus realized as -e and vinch as -ar. Semantically, vcaus introduces both a causing event and an activity predicate, while vinch introduces only a causing event. The structures also differ in their Voice heads: causatives employ Voicetr, which introduces an Initiator, while inchoatives use Voiceun, which does not. Following Bruening (2019), the inchoative structure contains a trace within VPRes, associated with the internal argument that undergoes A-movement to Spec,VoiceP. Voiceun, as an A-movement head, triggers the A-Movement Rule, lambda-abstracting over this trace and thereby explaining how the internal argument becomes the subject of the inchoative while maintaining its thematic relation to the result state.7

Building on this structural analysis, I now examine how soo interacts with these VP structures to account for transitivity mismatches. I propose that the anaphoric element soo has the following denotation, parallel to English do so in Bruening (2019):

    1. (17)
    1. ⟦VPsoo⟧= λe.f(e)

Function f is replaced by the semantic content of the antecedent VPRes. This allows soo to preserve that result state while the surrounding verbal structure (v and Voice) determines the transitivity of the clause. Consider first a causative-to-inchoative mismatch case in (7c), reproduced in (18).

    1. (18)
    1. Mary-wa
    2. Mary-top
    1. watasi-ga
    2. I-nom
    1. doa-o
    2. door-acc
    1. sim-e-ta-to
    2. close-caus-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. katteni
    2. on.one’s.own
    1. soo
    2. so
    1. nar-ta.
    2. become-pst
    1. ‘Mary claimed that I closed the door, but it actually did so on its own.’
    1.  
    1. a.
    1.  
    1. b.
    1. ⟦VPsoo⟧ = λe.f(e)
    1.  
    1. c.
    1. ⟦VPRes⟧ = λs.close(s,the door)
    1.  
    1. d.
    1. Copy: ⟦VPsoo⟧ = λs.close(s,the door)
    1.  
    1. e.
    1. ⟦vinch⟧ = λ f⟨s,t⟩ λe∃s.caus(e,s) ∧ f(s)
    1.  
    1. f.
    1. ⟦vPinch⟧ = λe∃s.caus(e,s) ∧ close(s,the door)
    1.  
    1. g.
    1. ⟦Voiceun⟧ = λ fλe.f(e)
    1.  
    1. h.
    1. ⟦VoiceP⟧ = λe∃s.caus(e,s) ∧ close(s,the door)

The derivation proceeds as follows: VPsoo initially has a vacuous denotation (18b). Its semantic content is derived by replacing function f with the denotation of the antecedent VPRes (18c), resulting in (18d). Crucially, this operation copies only the result state, excluding the causative semantics introduced by vcaus in the antecedent. This mechanism preserves the core meaning (‘the closing of the door’) while allowing an event-structure mismatch: the causing event in the target is instead introduced by vinch (18f), which lacks the activity predicate and external argument of the causative variant. Voiceun applies without introducing an external argument. Since the copied VPRes from the causative antecedent contains no trace, the A-Movement Rule does not apply here, despite the presence of Voiceun. The inchoative v is realized morphologically as nar, contrasting with the regular inchoative morphology -ar in verbs like sim-ar.

A consequence of the causative-to-inchoative analysis in (18) is that the target clause does not admit an overt subject. Since the copied VPRes supplies ‘the door’ as the logical subject of the target clause, this argument is already fixed, and an independently realized overt subject is not available, as (19) shows.

    1. (19)
    1. Mary-wa
    2. Mary-top
    1. watasi-ga
    2. I-nom
    1. kono
    2. this
    1. doa-o
    2. door-acc
    1. sim-e-ta-to
    2. close-caus-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1.  
    1. ‘Mary claimed that I closed this door, but …’
    1.  
    1. a.
    1. *?zissai-wa
    2.    actually-top
    1. sore-ga
    2. it-nom
    1. katteni
    2. on.one’s.own
    1. soo
    2. so
    1. nar-ta.
    2. become-pst
    1.   ‘it actually did so on its own.’
    1.  
    1. b.
    1. *?zissai-wa
    2.    actually-top
    1. ano doa-ga
    2. that door-nom
    1. katteni
    2. on.one’s.own
    1. soo
    2. so
    1. nar-ta.
    2. become-pst
    1.   ‘that door actually did so on its own.’

In (19a), the overt pronoun sore ‘it’ appears as the subject, and in (19b), the full NP ano doa ‘that door’ does; both are unacceptable. With ‘this door’ already fixed as the logical subject, there is no room for an additional overt subject, whether it re-expresses that argument (sore) or introduces a distinct one (ano doa).8

However, an apparent exception exists. Consider the following example:

    1. (20)
    1. a.
    1. Mary-ga
    2. Mary-nom
    1. ano doa-wa
    2. that door-top
    1. zibun-ga
    2. self-nom
    1. sim-e-ta-to
    2. close-caus-pst-comp
    1. iw-ta.
    2. say-pst
    1. ‘Mary claimed that she closed that door.’
    1.  
    1. b.
    1. zissai
    2. actually
    1. kono doa-wa
    2. this door-top
    1. soo
    2. so
    1. nar-te-i-nakat-ta.
    2. become-ger-be-neg-pst
    1. ‘Lit.: Actually this door didn’t do so.’

Here, an overt subject (kono doa-wa ‘this door-top’) appears grammatically in the target clause (20b). The key difference is that in (20a), the antecedent object appears with the topic marker -wa, functioning as a contrastive topic rather than a simple accusative object. Topic elements undergo movement from their base position to a higher position (Saito 1985), and the VPRes in the antecedent contains a trace of the topic-marked phrase: [VPRes t1 close], where t1 is the trace of the topicalized object. When this VPRes is copied onto soo, the unaccusative Voice head of the target clause, being an A-movement head, triggers the A-Movement Rule, thereby licensing an independently realized topic-marked subject. The structural representation and semantic derivation of (20b) proceed as follows:

    1. (21)
    1. a.
    1.  
    1. b.
    1. ⟦VPRes⟧= λs.close(s,t1)
    1.  
    1. c.
    1. ⟦VPsoo⟧= λe.f(e)
    1.  
    1. d.
    1. Copy: ⟦VPsoo⟧ = λs.close(s,t1)
    1.  
    1. e.
    1. ⟦vPinch⟧= λe∃s.caus(e,s) ∧ close(s,t1)
    1.  
    1. f.
    1. ⟦Voiceun⟧= λ fλe.f(e)
    1.  
    1. g.
    1. ⟦VoiceP(1)⟧= λxλe∃s.caus(e,s) ∧ close(s,x) (by the A-Movement Rule)
    1.  
    1. h.
    1. ⟦VoiceP(2)⟧= λe∃s.caus(e,s) ∧ close(s,this door)

This accounts for the apparent exception: a trace rather than a full NP in the copied VPRes licenses an independently realized subject in the target clause.9

Next, I analyze an inchoative-to-causative mismatch case, reproduced from (8c) and illustrated with its structures and semantic compositions in (22):

    1. (22)
    1. Mary-wa
    2. Mary-top
    1. katteni
    2. on.one’s.own
    1. doa-ga
    2. door-nom
    1. sim-ar-ta-to
    2. close-inch-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. watasi-ga
    2. I-nom
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1. ‘Mary claimed that the door closed on its own, but actually I did so.’
    1.  
    1. a.
    1.  
    1. b.
    1. ⟦VPRes⟧ = λs.close(s,t1)
    1.  
    1. c.
    1. Copy: ⟦VPsoo⟧ = λe.f(e) → λs.close(s,t1)
    1.  
    1. d.
    1. ⟦vcaus⟧ = λ f⟨s,t⟩ λe∃s.activity(e) ∧ caus(e,s) ∧ f(s)
    1.  
    1. e.
    1. ⟦vPcaus⟧ = λe∃s.activity(e) ∧ caus(e,s) ∧ close(s,t1)
    1.  
    1. f.
    1. ⟦Voicetr⟧ = λ f⟨v,t⟩λxλe.f(e) ∧ Init(e,x)
    1.  
    1. g.
    1. ⟦VoiceP⟧ = λe∃s.activity(e) ∧ caus(e,s) ∧ close(s,t1) ∧ Init(e,I)

The antecedent shows a typical unaccusative structure where the internal argument has undergone A-movement to Spec,VoiceP. VPRes in the antecedent contains a trace t1 (22b), which is copied onto VPsoo (22c). The target then projects transitive layers, vcaus and Voicetr, introducing causative semantics (22e) and (22g). In this context, causative v is realized as su rather than the typical causative morpheme -e found in regular causative verbs, paralleling how nar appears in place of -ar in the causative-to-inchoative case.

This configuration differs from the causative–inchoative mismatch in an important respect. Since Voicetr is not an A-movement head,10 the A-Movement Rule is not available here. Voicetr instead introduces an external argument with an Initiator role in its specifier. The question, then, is how the trace t1 in the copied VPRes is interpreted.11 Since there is no local binder for this trace within the target VoiceP, the interpretation would be problematic if the trace t1 were treated as an ordinary A-trace/anaphor.

As Bruening (2019: 35) shows, VP-internal traces have two possible interpretations: (i) as variables abstracted over by the A-Movement Rule, or (ii) as copies under the copy theory of movement (Chomsky 1993), interpreted as identical to the moved NP.12

In the present inchoative-to-causative case, the copy interpretation obtains. The trace t1 in the copied VPRes is interpreted as identical to ‘the door’ from the antecedent (22a). Accordingly, the target VoiceP receives the denotation in (23).

    1. (23)
    1. ⟦VoiceP⟧= λe∃s.activity(e) ∧ caus(e,s) ∧ close(s,the door) ∧ Init(e,I).

This interpretation specifies that the causing event denoted by soo su has the same theme argument as the inchoative event in the antecedent—namely, ‘the door.’13

2.4 Cross-linguistic variation in transitivity mismatches: The case of English

English, unlike Japanese, prohibits transitivity mismatches in the inchoative-causative order. Bruening (2019: 39–40) proposes that causative variants incorporate an additional functional projection between VP and VoiceP—CAUSP (equivalent to vcaus in our analysis)—while inchoative variants lack this projection. Bruening argues that CAUS cannot merge above do so because English CAUS requires transitive VPs with direct objects. Since do so is inherently intransitive, it fails to satisfy these requirements, blocking inchoative-causative mismatches.

This analysis, however, raises a question about causative-causative configurations:

    1. (24)
    1. Mary claimed that I closed the door, but actually she did so.

In this example, both clauses exhibit causative structures. If the target do so structure lacks the CAUSP layer, how does it secure a causative interpretation?

To resolve this apparent contradiction, I propose that when the antecedent includes CAUSP, do so can target either CAUSP or VP as its antecedent. With causative antecedents, do so has two options: (i) target VP, yielding a causative-inchoative mismatch, or (ii) target CAUSP, yielding the matched causative-causative configuration in (24).

Conversely, with inchoative antecedents, do so can only target VP since the inchoative structure lacks CAUSP. This permits only matched inchoative-inchoative orders:

    1. (25)
    1. Mary claimed that the door closed on its own, and it actually did so.

This account not only elucidates how the target do so can encode causative semantics in causative-causative configurations, but also explains the systematic prohibition against the transitivity mismatch in the inchoative-causative order in English.

As such, the cross-linguistic variation between Japanese and English stems from two structural differences: (i) In Japanese, unlike English, a causation-related head is projected between VP and VoiceP in both causative and inchoative variants—vcaus for causatives and vinch for inchoatives—which are morphologically realized as su and nar, respectively; (ii) In Japanese, the element soo independently targets the antecedent VP.

Whether these properties are typologically correlated remains an open question. Future research should investigate whether VP anaphora in other languages permits similar mismatches and whether such languages share the structural properties identified in Japanese.

2.5 Summary of 2

This section has shown that Japanese VP anaphora permits bidirectional transitivity mismatches—both causative-to-inchoative and inchoative-to-causative—the latter prohibited in English. These follow from decomposing the traditional VP into VPRes and vP: the anaphor soo targets VPRes, preserving core content while accommodating different argument structures, with the v-head realized as su (vcaus) or nar (vinch) atop VPRes. The cross-linguistic difference follows from Japanese projecting causation-related heads in both variants and from the independent targeting capacity of soo, which together yield the observed variation.

3 Split between v and Voice: Light verb asymmetries

Having established the decomposition of the traditional VP into VPRes and vP in the preceding section, I now turn to the internal structure of the conventional vP. While the little v head has traditionally been assumed to serve multiple functions—introducing external arguments, licensing accusative case, and encoding causative events (Chomsky 1995)—I argue that these functions should be distributed across two distinct projections: Voice, which introduces external arguments, and v, which encodes causation (and licenses accusative case).

This theoretical division, previously advocated by Pylkkänen (2002; 2008), Harley (2013; 2017), and Legate (2014), receives further empirical support from the distribution of light verbs in Japanese. The asymmetric behavior, as shown below, requires distinguishing Voice and v, providing further evidence for the tripartite VP structure.

3.1 On the distribution of su and nar

3.1.1 Non-finite stative predicates vs. verbal nouns

Evidence for decomposing the traditional vP into VoiceP and vP emerges from the distribution of su ‘do’ and nar ‘become’ in Japanese. With non-finite stative predicates, these verbs show systematic alternation based on transitivity:

    1. (26)
    1. a.
    1. John-ga
    2. John-nom
    1. kabe-o
    2. wall-acc
    1. akaku
    2. red
    1. si-ta.
    2. do-pst
    1. ‘John made the wall red.’
    1.  
    1. b.
    1. Kabe-ga
    2. wall-nom
    1. akaku
    2. red
    1. nar-ta.
    2. become-pst
    1. ‘The wall became red.’
    1. (27)
    1. a.
    1. John-ga
    2. John-nom
    1. heya-o
    2. room-acc
    1. kirei-ni
    2. clean-prt
    1. si-ta.
    2. do-pst
    1. ‘John made the room clean.’
    1.  
    1. b.
    1. Heya-ga
    2. room-nom
    1. kirei-ni
    2. clean-prt
    1. nar-ta.
    2. become-pst
    1. ‘The room became clean.’

Example (26) illustrates the alternation with a canonical adjective (akaku ‘red’), while (27) shows the same alternation with a nominal adjective (kirei-ni ‘clean’). In this alternation, transitive variants (a) involve su appearing with two arguments—an external argument (nominative) and an internal argument (accusative). Intransitive variants (b) employ nar, where the internal argument becomes the subject with nominative case.

This correlation indicates that the su/nar alternation is linked to argument structure, particularly the presence/absence of an external argument. The stative predicate remains constant across variants, suggesting that the alternation is governed by syntactic configuration rather than the semantic content of the result state.

Based on such data, Sakai et al. (2004) identify su and nar in stative predicate constructions as light verbs rather than lexical verbs, demonstrating that the stative predicate—not su/nar—plays the central role in theta-role assignment to the theme argument.14 This becomes evident when we compare stative predicate constructions with su/nar to those formed with a lexical verb instead:

    1. (28)
    1. a.
    1. John-ga
    2. John-nom
    1. kabe-o
    2. wall-acc
    1. (akaku)
    2. red
    1. som-e-ta.
    2. dye-caus-pst
    1. ‘John dyed the wall (red).’
    1.  
    1. b.
    1. Kabe-ga
    2. wall-nom
    1. (akaku)
    2. red
    1. som-ar-ta.
    2. dye-inch-pst
    1. ‘The wall got dyed (red).’
    1. (29)
    1. a.
    1. John-ga
    2. John-nom
    1. heya-o
    2. room-acc
    1. (kirei-ni)
    2. clean-prt
    1. katazuk-e-ta.
    2. tidy.up-caus-pst
    1. ‘John tidied up the room (neatly).’
    1.  
    1. b.
    1. Heya-ga
    2. room-nom
    1. (kirei-ni)
    2. clean-prt
    1. katazuk-Ø-ta.
    2. tidy.up-inch-pst
    1. ‘The room got tidied up (neatly).’

The crucial distinction lies in the optionality of the predicate: with lexical verbs like som-e/som-ar ‘dye (vt/vi)’, the stative predicate can be omitted, so the internal argument is theta-marked by the verb. In contrast, with su/nar, it is obligatory (30)–(31), showing that here the theme is theta-marked by the predicate, not by su or nar, confirming their light-verb status.

    1. (30)
    1. a.
    1. John-ga
    2. John-nom
    1. kabe-o
    2. wall-acc
    1. *(akaku)
    2. red
    1. si-ta.
    2. do-pst
    1. ‘John made the wall red.’
    1.  
    1. b.
    1. Kabe-ga
    2. wall-nom
    1. *(akaku)
    2. red
    1. nar-ta.
    2. become-pst
    1. ‘The wall became red.’
    1. (31)
    1. a.
    1. John-ga
    2. John-nom
    1. heya-o
    2. room-acc
    1. *(kirei-ni)
    2. clean-prt
    1. si-ta.
    2. do-pst
    1. ‘John made the room clean.’
    1.  
    1. b.
    1. Heya-ga
    2. room-nom
    1. *(kirei-ni)
    2. clean-prt
    1. nar-ta.
    2. become-pst
    1. ‘The room became clean.’

Based on these observations, Sakai et al. (2004) propose the following syntactic structure for stative predicate constructions, where su and nar instantiate Chomsky’s (2001) transitive v* and intransitive v, respectively:15

    1. (32)
    1. a.
    1.  
    1. b.

In this binary approach, su realizes transitive v*, which introduces an external argument and assigns accusative case, while nar realizes intransitive v, which does neither.

Sakai et al.’s (2004) analysis parallels analyses of Japanese morphological transitivity alternation proposed by Nishiyama (1998) and Hasegawa (2001), as seen in pairs like:

    1. (33)
    1. a.
    1. John-ga
    2. John-nom
    1. kabe-o
    2. wall-acc
    1. kowa-si-ta.
    2. break-caus-pst
    1. ‘John broke the wall.’
    1.  
    1. b.
    1. Kabe-ga
    2. wall-nom
    1. kowa-re-ta.
    2. break-inch-pst
    1. ‘The wall broke.’

Here, the causative variant uses the morpheme -s while the inchoative uses -re, both attached to the root kowa- ‘break’. Nishiyama (1998) and Hasegawa (2001) propose the following structures for these alternations:

    1. (34)
    1. a.
    1.  
    1. b.

This parallel treatment suggests that just as the causative morpheme -s and inchoative morpheme -re realize v* and v respectively in morphological alternations, su and nar can be analyzed as their light verb counterparts in periphrastic constructions, providing a unified account of both alternation types in Japanese.

While the binary distinction between v* and v successfully accounts for light verb distribution in stative predicate constructions, it proves inadequate for capturing their behavior with verbal nouns (VNs). Since Grimshaw & Mester (1988), su has been widely recognized as a light verb when combined with VNs such as Sino-Japanese verbs (e.g., Miyagawa 1989; Saito & Hoshi 2000; Fukui & Sakai 2003):

    1. (35)
    1. a.
    1. Taro-ga
    2. Taro-nom
    1. Amerika-ni
    2. America-to
    1. ryokoo-o
    2. travel-acc
    1. si-ta.
    2. do-pst
    1.  
    1. b.
    1. Taro-ga
    2. Taro-nom
    1. Amerika-ni
    2. America-to
    1. ryokoo-si-ta.
    2. travel-do-pst
    1. ‘Taro took a trip to the United States.’

The present paper focuses on the non-accusative VN-su construction of the type exemplified in (35b), as opposed to the accusative VN-o suru pattern exemplified in (35a), which is set aside here. What matters for this construction is that some VNs can be used both transitively and intransitively. As noted by Kageyama (1996: 202), this class includes items such as kakudai ‘expansion’, syukusyoo ‘reduction’, and zitugen ‘realization’. Even in these cases, however, the light verb remains su in both the transitive and intransitive uses; it does not alternate with nar.

    1. (36)
    1. a.
    1. Seifu-ga
    2. government-nom
    1. yusyutu-ryoo-o
    2. export-volume-acc
    1. {kakudai/syukusyoo}-si-ta.
    2. expansion/reduction-do-pst
    1. ‘The government expanded/reduced the export volume.’
    1.  
    1. b.
    1. Yusyutu-ryoo-ga
    2. export-volume-nom
    1. {kakudai/syukusyoo}-{si/*nar}-ta.
    2. expansion/reduction-do/become-pst
    1. ‘The export volume expanded/decreased.’

This pattern reveals a fundamental asymmetry in the distribution of light verbs: while su participates in transitivity alternation with nar when combined with stative predicates, it fails to do so when combined with bare VNs, regardless of whether the VN itself allows both transitive and intransitive uses. This asymmetry is summarized in Table 1.

Table 1: Distribution of light verbs su and nar across predicate types.

Complement of the light verb transitive intransitive
Adjectival Stative Predicates su nar
Verbal Nouns su su

Notably, a small number of VNs also function as -ni-marked stative predicates, though speaker judgments vary. One such item is tyuusi ‘cancellation’, illustrated in (37)–(38). Since the intransitive bare-VN use in (37b) and the transitive -ni-marked use in (38a) appear to be the forms most subject to this variation, naturally attested web examples are provided for them.

    1. (37)
    1. a.
    1. Monkasyoo-ga
    2. Ministry.of.Education-nom
    1. ibento-o
    2. event-acc
    1. tyuusi-si-ta
    2. cancellation-do-pst
    1. baai
    2. case
    1.  
    1. ‘In the case where the Ministry of Education canceled the event …’
    1.  
    1. b.
    1. Ibento-ga
    2. event-nom
    1. tyuusi-si-ta
    2. cancellation-do-pst
    1. baai
    2. case
    1.  
    1. ‘In the case where the event was canceled …’
    2. (https://www.tokyo-hokensyakyougikai.jp/contents/report/2021/pdf/02/04_activity_group.pdf)
    1. (38)
    1. a.
    1. Monkasyoo-ga
    2. Ministry.of.Education-nom
    1. ibento-o
    2. event-acc
    1. tyuusi-ni
    2. cancellation-prt
    1. si-ta
    2. do-pst
    1. koto
    2. fact
    1.  
    1. ‘The fact that the Ministry of Education canceled the event …’
    2. (https://www.bengo4.com/c_18/n_16116/)
    1.  
    1. b.
    1. Ibento-ga
    2. event-nom
    1. tyuusi-ni
    2. cancellation-prt
    1. nar-ta
    2. become-pst
    1. koto
    2. fact
    1.  
    1. ‘The fact that the event was canceled …’

Thus, tyuusi occurs with su in its bare VN use, but shows the expected su/nar alternation when used as a -ni-marked stative predicate.

Returning to the main asymmetry summarized in Table 1, it challenges Sakai et al.’s binary analysis, which treats su and nar as realizations of v* and v, respectively. Under their analysis, su in VN contexts should alternate with nar, contrary to observations. While Sakai et al. (2004: 372) suggest that “[Noun+suru] as a whole has similar functions to [V+v*] complex,” this fails to explain why [Noun+naru] cannot function similarly to [V+v] in unaccusative variants.

3.1.2 On the locus of transitivity

The key to resolving this puzzle lies in recognizing that in VN-based alternations, transitivity is an inherent property of the VNs themselves, rather than a function of the light verb (e.g., Miyagawa 1989; Kageyama 1991). While this distinction is not immediately apparent with VNs like kakudai/syukusyoo ‘expansion/reduction’ in (36), which can appear in both contexts without morphological change, most VNs show clear transitivity-dependent distribution:

    1. (39)
    1. a.
    1. Sosiki-ga
    2. gang-nom
    1. kuruma-o
    2. car-acc
    1. {bakuha/*bakuhatu}-si-ta.
    2. explosion-do-pst
    1. ‘A gang exploded the car.’
    1.  
    1. b.
    1. Kuruma-ga
    2. car-nom
    1. {bakuhatu/*bakuha}-si-ta.
    2. explosion-do-pst
    1. ‘The car exploded.’

Here, bakuha and bakuhatu, while sharing the core meaning ‘explosion’, differ in their inherent transitivity: bakuha is inherently transitive, while bakuhatu is inherently unaccusative.

This lexically encoded transitivity becomes even more evident in deverbal N–V compounds, where a noun combines with the infinitive form (renyookei) of a verb. Some of these compounds take the light verb su, paralleling Sino-Japanese verbs, and exhibit morphologically transparent transitivity alternations.16,17

    1. (40)
    1. a.
    1. Seifu-ga
    2. government-nom
    1. tabako-o
    2. cigarette-acc
    1. ne-ag-e-si-ta.
    2. price-rise-caus-do-pst
    1. ‘The government raised the price of cigarettes.’
    1.  
    1. b.
    1. Tabako-ga
    2. cigarette-nom
    1. ne-ag-ari-si-ta.
    2. price-rise-inch-do-pst
    1. ‘The price of cigarettes increased.’

The compound consisting of the noun ne ‘price’ and the verb root ag- ‘rise’ shows systematic morphological alternation: causative morpheme -e in transitive contexts and inchoative morpheme -ar in unaccusative contexts.18 This pattern precisely mirrors the simplex verb pair ag-e ‘raise’ and ag-ar ‘rise’:

    1. (41)
    1. a.
    1. Seifu-ga
    2. government-nom
    1. tabako-no
    2. cigarette-gen
    1. nedan-o
    2. price-acc
    1. ag-e-ta.
    2. rise-caus-pst
    1. ‘The government raised the price of cigarettes.’
    1.  
    1. b.
    1. Tabako-no
    2. cigarette-gen
    1. nedan-ga
    2. price-nom
    1. ag-ar-ta.
    2. rise-inch-pst
    1. ‘The price of cigarettes increased.’

These parallel morphological alternation patterns indicate that transitivity in VN-su predicates is encoded within the VNs themselves, not in the light verb su.

For stative predicates, in contrast, transitivity must be encoded in the light verb itself, as evidenced by the morphological su/nar alternation. This necessity stems from adjectival predicates lacking internal mechanisms for transitivity alternations:

    1. (42)
    1. a.
    1.   Heya-ga
    2.   room-nom
    1. {kirei-da/atataka-i}.
    2. {clean-cop/warm-prs}
    1.   ‘The room is {clean/warm}.’
    1.  
    1. b.
    1. *John-ga
    2.   John-nom
    1. heya-o
    2. room-acc
    1. {kirei-da/atataka-i}.
    2. {clean-cop/warm-prs}
    1.   ‘Int.: John made the room {clean/warm}.’

Both nominal adjective kirei-da and canonical adjective atataka-i are inherently one-place predicates that can only take a subject, as in (42a). Their inability to license an external argument and accusative object is shown by the ungrammaticality of (42b).19

To summarize, Japanese light verbs exhibit two distinct patterns depending on the predicate type they combine with:

    1. (43)
    1. a.
    1. When the predicate encodes transitivity inherently (e.g., VNs), the light verb surfaces uniformly as su, regardless of the transitivity value.
    1.  
    1. b.
    1. When the predicate lacks transitivity specification (e.g., adjectival stative predicates), the light verb encodes transitivity through morphological su/nar alternation.

This generalization reveals a fundamental architectural property of the Japanese verbal domain: the morphological realization of light verbs systematically reflects the locus of transitivity specification in the predicate-light verb complex.

3.2 Proposal

The preceding sections have revealed two distinct patterns in the distribution of light verbs su and nar: with stative predicates, they alternate based on transitivity, while with VNs, only su appears regardless of transitivity. This asymmetry correlates directly with whether transitivity is encoded in the predicate itself—VNs inherently specify transitivity, whereas stative predicates do not.

I propose that this distribution reflects a fundamental architectural property of the verbal domain. Specifically, the traditional vP should be decomposed into two distinct functional layers above VPRes: VoiceP and vP. Under this analysis, transitivity emerges from the interaction between Voice, which licenses external arguments, and v, which specifies event type (causative, inchoative, etc.). Japanese light verbs fall into two categories: v-type light verbs that alternate between su and nar based on the specification of v (vcaus vs. vinch), and Voice-type light verbs (consistently su) that show no such alternation.

3.2.1 Tripartite VP analysis

To illustrate this proposal, let us examine the structures of stative predicate constructions like those in (27a) and (27b):

    1. (44)
    1. a.
    1.  
    1. b.

In the transitive variant (44a), vcaus combines with AP to encode a causative event, while transitive Voice (Voicetr) introduces the external argument. In the intransitive variant (44b), vinch takes AP as its complement to denote an inchoative event, with unaccusative Voice (Voiceun) projecting no specifier position. Morphologically, these structures are realized through head movement: v merges with Voice, with vcaus surfacing as su and vinch as nar, while Voice is consistently null (Ø). This analysis distributes the determination of argument structure across two functional heads: v specifies event type (causative/inchoative), while Voice regulates external argument presence, providing a more nuanced account than Sakai et al.’s (2004) binary v*/v distinction.

Turning to the VN-su construction exemplified in (36), I propose the following structures:

    1. (45)
    1. a.
    1.  
    1. b.

In both variants, the VN root projects its own phrasal structure (VNP), selects its internal argument, and undergoes head movement to v.20 In the transitive variant, Voicetr introduces an external argument above vPcaus; in the intransitive variant, Voiceun merges with vPinch without introducing an external argument. Crucially, the morphological realization pattern here is inverted compared to stative predicate constructions: with VNs, v is consistently null (Ø) while Voice invariably surfaces as su, regardless of transitivity. This contrasts with stative predicate constructions where v is overtly realized as su/nar while Voice is null.

The asymmetric distribution of the light verbs su and nar then follows from the distinct structural positions these light verbs occupy within the verbal domain. This distribution is summarized in Table 2.

Table 2: Morphological realization of vcaus/inch and Voice across predicate types.

Complement of the light verb vcaus/inch Voice
Adjectival Stative Predicates su / nar Ø
Verbal Nouns Ø su

This raises three interrelated questions: (i) Why does Voice realize as su with VNs but remain null with stative predicates? (ii) Why does v alternate between su/nar with stative predicates but surface as Ø with VNs? (iii) What mechanisms regulate these patterns?

To account for these patterns, I adopt the late insertion model of Distributed Morphology (Halle & Marantz 1993). In the following sections, I examine the spell-out conditions governing the morphological realization of these functional heads, demonstrating how this framework captures both invariant and alternating patterns in the data.

3.2.2 Voice morphology

To account for the morphological realization of the Voice head, I propose the following Vocabulary Insertion (VI) rules:

    1. (46)
    1. Vocabulary Insertion rules for Voice
    1.  
    1. a.
    1. Voice ↔ Ø / [v ___]Voice
    1.  
    1. b.
    1. Voice ↔ su

These rules specify that Voice is realized as null (Ø) when it enters into a local relationship with v by forming a complex head, and as su elsewhere. This context-sensitive distribution follows standard locality conditions and the Elsewhere Principle (Kiparsky 1973), with the more specific rule in (46a) taking precedence over the general case in (46b). Rule (46a) accounts for Voice being realized as Ø in stative predicate constructions (44), where Voice forms a local configuration with v via head movement.

However, we must explain why Voice consistently surfaces as su in VN-su constructions (45). The VI rules in (46) suggest that in these constructions, Voice and v do not form a complex head. Following Hayashi (2015), I propose that this pattern results from a structural constraint on VNs: they are fixed at the vP level and cannot undergo raising to Voice. Consequently, Voice fails to enter into a local relationship with v, triggering the elsewhere realization as su.

This analysis suggests that the VN and the light verb su are structurally independent entities (Kageyama 1993; Miyamoto & Kishimoto 2016).21 The structural independence of VNs and the light verb su finds empirical support in ellipsis phenomena.22 Hayashi (2015) observes a significant contrast between Native-Japanese verbs (NJVs) and Sino-Japanese verbs (SJVs). Consider first:

    1. (47)
    1. a.
    1.   Taro-wa
    2.   Taro-top
    1. Nihon-e
    2. Japan-to
    1. ki-ta
    2. come-pst
    1. kedo,
    2. but
    1.   ‘Taro came to Japan, but …’
    1.  
    1. b.
    1.   Ziro-wa
    2.   Ziro-top
    1. [Compl
    2.  
    1. e
    2.  
    1. ]
    2.  
    1. ko-nakat-ta.
    2. come-neg-pst
    1.   ‘Ziro didn’t come to Japan.’
    1.  
    1. c.
    1. *Ziro-wa
    2.   Ziro-top
    1. [Compl
    2.  
    1. e
    2.  
    1. ]
    2.  
    1. [NJV
    2.  
    1. e
    2.  
    1. ]
    2.  
    1. nakat-ta.
    2. neg-pst
    1.   ‘Ziro didn’t come to Japan.’

In both target clauses, the complement of the verb—the goal PP Nihon-e ‘to Japan’—is elided, indicated by [Compl e ]. The two clauses differ minimally in whether the NJV is present in the surface string: while (47b), with the pronounced NJV ko, is well-formed, (47c), with an elided NJV, is ungrammatical. Contrast this with SJVs:

    1. (48)
    1. a.
    1. Taro-wa
    2. Taro-top
    1. Nihon-e
    2. Japan-to
    1. kikoku-si-ta
    2. returning-do-pst
    1. kedo,
    2. but
    1. ‘Taro went back to Japan, but …’
    1.  
    1. b.
    1. Ziro-wa
    2. Ziro-top
    1. [Compl
    2.  
    1. e
    2.  
    1. ]
    2.  
    1. kikoku-si-nakat-ta.
    2. returning-do-neg-pst
    1. ‘Ziro didn’t go back to Japan.’
    1.  
    1. c.
    1. Ziro-wa
    2. Ziro-top
    1. [Compl
    2.  
    1. e
    2.  
    1. ]
    2.  
    1. [SJV
    2.  
    1. e
    2.  
    1. ]
    2.  
    1. si-nakat-ta.
    2. do-neg-pst
    1. ‘Ziro didn’t go back to Japan.’

Here, both variants—with the pronounced VN kikoku in (48b) and with the elided VN in (48c)—are grammatical, with the light verb su stranded in the latter case.

This asymmetry in anaphoric properties led Hayashi (2015) to posit differential patterns of predicate raising: NJVs undergo head movement to a higher functional position, while SJVs remain in their base position. The current tripartite VP analysis refines this insight by providing a more articulated structural account: NJVs raise successively through v to Voice, while SJVs (and other VNs) raise only as far as v, not to Voice.23

    1. (49)
    1. a.
    1.  
    1. b.

The proposed analysis explains the grammatical contrast between (47c) and (48c) under the assumption that VP ellipsis targets vP in the tripartite structure (e.g., Otani & Whitman 1991; Funakoshi 2016).24 With SJVs, the VN remains within the ellipsis domain (vP) alongside its complement (49b). Consequently, vP-level ellipsis elides the VN while stranding Voice (realized as su), correctly predicting the grammaticality of (48c). With NJVs, by contrast, V raises to Voice via v, positioning it outside the domain of vP ellipsis (49a). It thus remains phonologically realized even when vP is elided—hence the ungrammaticality of (47c).25

In sum, Voice in Japanese exhibits sensitivity to structural locality with v: when in a local configuration within a complex head, it is realized as null (Ø); when structurally isolated due to the absence of V-v to Voice movement, it surfaces as the elsewhere form su.

Before examining v’s morphological realization, two comments are necessary regarding (i) NJVs and the light verb su, and (ii) v-to-Voice movement constraints. First, NJVs systematically exclude su in canonical contexts, yet require it in object honorification:

    1. (50)
    1. a.
    1. John-ga
    2. John-nom
    1. sensei-o
    2. teacher-acc
    1. tasuke-(*si)-ta.
    2. help-do-pst
    1. ‘John helped his teacher.’
    1.  
    1. b.
    1. John-ga
    2. John-nom
    1. sensei-o
    2. teacher-acc
    1. o-tasuke-*(si)-ta.
    2. hon-help-do-pst
    1. ‘John helped his teacher.’

The pattern in (50a) follows from our analysis: NJV roots undergo successive head movement through v to Voice, establishing a local relationship that triggers Voice’s null realization via rule (46a). The object honorification exception in (50b) can be explained following Ikawa (2022): the honorific head hon projects between vP and VoiceP, blocking V-v to Voice movement:

    1. (51)

In (51), the honorific head hon intervenes between vP and VoiceP, blocking head movement of the V-v complex (containing the NJV tasuke ‘help’) to Voice. Voice is therefore not in a local relation with V-v and surfaces as su via rule (46b). More generally, when an intervening element disrupts v-Voice locality, Voice is realized as su.

Regarding the theoretical underpinnings of head movement constraints, this analysis remains agnostic about the fundamental question of why certain predicate types such as SJVs resist head movement in the grammar. A comprehensive account would require cross-linguistic comparison and investigation of a broader range of predicates—endeavors beyond the scope of the present study. Nevertheless, Hayashi’s (2015) feature-based account offers a promising theoretical direction worthy of consideration.

Hayashi proposes that verb raising is triggered by an uninterpretable feature [uNative] on the functional little v head (equivalent to Voice in our tripartite structure). Consequently, verbs bearing an interpretable feature [iNative]—specifically, verbs of Native Japanese origin—undergo obligatory head raising, while other predicates such as SJVs cannot participate in this movement operation due to their lack of the requisite feature specification. Within this framework, the light verb su is inserted precisely to value the [uNative] feature on the higher functional head.

This account encounters challenges with deverbal compounds like ne-age(-su) ‘price-rise-do’, whose head is native but still requires su. One potential refinement would posit that non-head elements in compounds function as interveners for head movement, accounting for both deverbal compounds and object honorification constructions.26 However, this refinement raises questions about what constitutes an intervener, as V-V compounds (e.g., tataki-kowas ‘hit-break’) pattern with ordinary NJVs despite their complex structure. This suggests that structural complexity alone cannot explain the observed patterns.

In sum, the realization of su correlates with the absence of v-to-Voice movement, though a full theoretical interpretation—ideally integrating feature-based and structural-intervention approaches—remains for future work.

3.2.3 v morphology

In our tripartite analyses (44) and (45), vcaus/inch is realized as su/nar in stative predicate constructions but as Ø in VN-su constructions. However, this simple pattern fails to account for two critical phenomena: first, transitivity-alternating NJVs like kowa-s ‘break-caus’ and kowa-re ‘break-inch’, where causative/inchoative morphology surfaces in verb root-sensitive patterns; second, deverbal compounds like ne-ag-e-su ‘price-rise-caus-do’ and ne-ag-ari-su ‘price-rise-inch-do’ that exhibit transitivity alternation while still requiring su.

To account for these patterns, I propose the following VI rules:

    1. (52)
    1. Vocabulary Insertion rules for vcaus
    1.  
    1. a.
    1. vcaus ↔ -S / [ V[+Native] ]v {e.g., kowa- ‘break’, ag- ‘rise’, …}
    1.  
    1. b.
    1. vcaus ↔ Ø / [ V[−Native] ]v
    1.  
    1. c.
    1. vcaus ↔ su
    1. (53)
    1. Vocabulary Insertion rules for vinch
    1.  
    1. a.
    1. vinch ↔ -R / [ V[+Native] ]v {e.g., kowa- ‘break’, ag- ‘rise’, …}
    1.  
    1. b.
    1. vinch ↔ Ø / [ V[−Native] ]v
    1.  
    1. c.
    1. vinch ↔ nar

These rules instantiate a principled hierarchy of morphological realization. When v combines with a transitivity-alternating, native verb root ([+Native]), it surfaces as the abstract morphemes -S/-R (52a)/(53a) (e.g., Oseki 2021), which undergo root-conditioned allomorphy (e.g., -S/-R becomes -s/-re with kowa ‘break’ but -e/-ar with ag ‘rise’). This allomorphic variation is mediated through morpho-phonological operations (see e.g., Inoue 1976; Akimoto 2018; cf. Miyagawa 1994; 1998; Harley 2008).27,28 When v combines with a non-native verb root ([−Native]), it is realized as Ø (52b)/(53b). In environments where v does not establish a local relationship with any verbal root, the elsewhere forms su/nar emerge (52c)/(53c).

The proposed rules apply across four construction types: NJVs, stative predicate constructions (SPCs), SJVs, and deverbal compounds (DCs). Let us first consider the morphosyntactic structures of constructions with NJVs and SPCs, as illustrated in (54a) and (54b), respectively.

    1. (54)
    1. a.
    1. NJVs
    1.  
    1. b.
    1. SPCs

In NJV constructions (54a), a complex head comprising V-v-Voice is formed through successive head movement. Crucially, v establishes a local configuration with a [+Native] verb root (e.g., kowa- ‘break’), triggering the insertion of causative/inchoative morphemes (-S/-R) via rules (52a)/(53a). These abstract morphemes then undergo root-conditioned allomorphy, surfacing as -s and -re. SPCs (54b) present a different configuration: v forms a local relationship with Voice (realized as Ø) but not with any verbal root, as its complement is an adjectival phrase that cannot undergo head movement to v. Consequently, the elsewhere forms su/nar emerge via rules (52c)/(53c).29

The structural configurations of SJVs and DCs exhibit distinct yet systematically related patterns of morphological realization:

    1. (55)
    1. a.
    1. SJVs
    1.  
    1. b.
    1. DCs

In SJVs (55a), the verbal root bears the [−Native] feature (e.g., kakudai ‘expansion’), reflecting their Sino-Japanese origin. This triggers the null realization of vcaus/inch via rules (52b)/(53b), resulting in no overt morphological exponent at the v node. DCs (55b), in contrast, involve a more complex architecture: a nominal element (N) merges with a verbal root (V) to form a compound base, which raises to v. Here, v establishes a local relation with a V node bearing the [+Native] feature, licensing the insertion of causative/inchoative morphemes (-e/-ar) via rules (52a)/(53a), paralleling the morphological pattern of NJVs. Despite their structural differences, both SJVs and DCs share a crucial property: their V-v complexes cannot undergo further head movement to Voice—likely due to locality constraints or intervening material. As a result, Voice remains in situ, necessitating its realization as the elsewhere morpheme su via rule (46b).

The pattern observed here—where the presence or absence of head movement determines whether a predicate surfaces as a simplex form or requires a light verb—aligns with cross-linguistic patterns of complex predicate formation. Hale & Keyser (1993; 2002) argue that many unergatives derive from an underlying N+V configuration, with languages differing in whether the nominal element incorporates into the verb. English displays incorporation, yielding morphologically simplex verbs (e.g., cry from N-to-V incorporation into do), whereas Persian exhibits the opposite pattern: the nominal remains in situ and an overt light verb is required, as in gerye kardan ‘cry do’ (Folli et al. 2005).

Japanese fits naturally within this typology. NJVs undergo full head movement to Voice (54a), producing morphologically integrated forms without an overt light verb, while VNs such as SJVs move only to v (55a), necessitating the overt realization of Voice as su. Thus, the degree of head movement directly determines morphological realization, paralleling the incorporation vs. light verb contrast observed cross-linguistically.30

3.2.4 On mismatches between v and Voice

The preceding analysis has focused on configurations where v and Voice align in transitivity: vcaus combines with Voicetr, and vinch with Voiceun. However, the tripartite architecture logically permits mismatches: vcaus could combine with Voiceun, and vinch with Voicetr. This subsection demonstrates that at least the former type of mismatch is well-motivated in the literature, providing further evidence for treating v and Voice as independent functional projections.31

The first type of vcaus + Voiceun mismatch involves causative constructions where the external argument is interpreted as an experiencer or affectee rather than as an agent or causer. Consider (56):

    1. (56)
    1. a.
    1. Taro-wa
    2. Taro-top
    1. (kazi-de)
    2. fire-by
    1. ie-o
    2. house-acc
    1. yak-Ø-ta.
    2. burn-caus-pst
    1. ‘Taro got his house burned down (in a fire).’
    1.  
    1. b.
    1. Taro-wa
    2. Taro-top
    1. (kazi-de)
    2. fire-by
    1. ie-ga
    2. house-nom
    1. yak-e-ta.
    2. burn-inch-pst
    1. ‘As for Taro, his house burned down (in a fire).’

In (56a), the subject Taro is not construed as an agent or causer but rather as an experiencer or affectee—someone affected by the burning event. This non-agentive interpretation is similar to that of the standard inchoative in (56b). Asami (2024) proposes that experiencer subjects in causatives like (56a) are licensed by the functional head Affect (Bosse et al. 2012), which projects between vcaus and Voiceun in the present framework:32,33

    1. (57)

Under this analysis, vcaus introduces an event that brings about the eventuality denoted by its complement (i.e., the event of Taro’s house burning down), but Voiceun does not introduce an Initiator. Instead, Affect licenses the external argument as an experiencer/affectee. The structure in (57) can be semantically paraphrased as: there is a set of events that causes Taro’s house to burn, and this situation affects Taro.34 This configuration exemplifies vcaus + Voiceun, demonstrating that v and Voice need not align in transitivity.

It should be noted that in (57), the structure for (56a), Affect projects between vcaus and Voiceun, yet it does not block V-v from reaching Voice. The question, then, is why Affect, intervening between v and Voice, does not block this movement, given that hon—in the same structural position in object honorification (51)—does, forcing the realization of su on Voice.35 I suggest two possibilities. One is phonological: head movement has often been argued to be sensitive to overt intervening heads (see e.g., Kandybowicz 2015 on Asante Twi). hon is an overt head, being realized as the prefix o-, whereas Affect lacks any phonological exponent.36 The other is categorial. Affect is responsible for introducing an argument (i.e., experiencer) just like Voice. hon, by contrast, introduces no argument and contributes nothing to event structure. If head movement depends on a local relation among verbal heads, hon may disrupt that relation whereas Affect may not. Notably, -tari, which will be argued in section 5 to be an intervener, patterns with hon in being both overt and non-argument-introducing. On either view, Affect is correctly predicted not to block this movement, whereas hon and -tari, being both overt and non-argument-introducing, are predicted to do so; which property is operative I leave for future research.

The second type of vcaus + Voiceun mismatch is what Schäfer (2025) terms transitive anticausatives, which are observed cross-linguistically:

    1. (58)
    1. a.
    1. The clouds altered their shape.
    1.  
    1. b.
    1. The shape of the clouds altered.

(58a) takes a transitive structure but is semantically equivalent to the unaccusative (58b). The nominative subject in (58a) does not denote an agent or causer; rather, both sentences describe a spontaneous change event. Japanese exhibits a similar pattern with causative-inchoative alternating verbs:

    1. (59)
    1. a.
    1. (Kaze-de)
    2. wind-by
    1. kumo-ga
    2. cloud-nom
    1. katati-o
    2. shape-acc
    1. kaw-e-ta.
    2. change-caus-pst
    1. ‘The clouds changed their shape (due to wind).’
    1.  
    1. b.
    1. (Kaze-de)
    2. wind-by
    1. kumo-no
    2. cloud-gen
    1. katati-ga
    2. shape-nom
    1. kaw-ar-ta.
    2. change-inch-pst
    1. ‘The shape of the clouds changed (due to wind).’

In (59a), the verb takes the causative form kaw-e, yet the nominative subject kumo ‘clouds’ is not a causer. This is evidenced by the fact that the actual causer can be expressed as an adjunct PP kaze-de ‘by wind’ in both examples—a pattern unexpected if the subject were introduced by Voicetr as a causer. In line with Schäfer (2025), I analyze transitive anticausatives as involving vcaus combined with semantically contentless Voiceun:

    1. (60)

In (60), vcaus introduces the change-of-state event affecting the shape of the clouds, while the subject kumo ‘clouds’ is not assigned an Initiator role by Voice. Rather it originates in the possessor position of the theme katati ‘shape’ and raises to Spec,VoiceP via possessor raising.

These two types of vcaus + Voiceun mismatches demonstrate that v and Voice can vary independently. The converse pattern—vinch + Voicetr—is also logically possible within the tripartite architecture but lacks clear empirical instantiation in Japanese.37

Nevertheless, the well-established existence of vcaus + Voiceun patterns already demonstrates that v and Voice can vary independently in transitivity. This flexibility—whereby event structure (causative vs. inchoative) and Initiator licensing (present vs. absent) are controlled by distinct functional heads—would be unexpected under bipartite approaches that conflate v and Voice into a single head. The tripartite architecture naturally accommodates such mismatches, providing a more empirically adequate framework.

3.3 Summary of 3

This section has established that the asymmetric distribution of su ‘do’ and nar ‘become’—resistant to explanation under the traditional bipartite VP structure (e.g., Sakai et al. 2004)—follows from bifurcating the conventional vP into vP and VoiceP. Japanese light verbs then fall into two types according to their syntactic locus. The v-type, prototypically in stative predicate constructions, alternates between su and nar as governed by the type of v (vcaus or vinch); the Voice-type, associated with VNs such as Sino-Japanese verbs and deverbal compounds, invariably surfaces as su.

This distinction parallels the one drawn by Si (2021) between heavy light verbs, which contribute event-structural semantics (our v-type), and light light verbs, which serve as largely functional elements (our Voice-type). The tripartite structure thus instantiates Si’s Split light verb hypothesis, on which Si (2021: 219) treats “v” not as “ONE” head but as a “rich structural” zone, lending empirical support to cartographic approaches: the seemingly arbitrary distributions of su and nar emerge from systematic interactions of structure, head movement, and featural specification. This bifurcation may reflect a universal property of verbal domains, inviting comparison with other rich light-verb systems such as Persian, Hindi-Urdu, and Korean.

4 Another asymmetry of light verbs: VP anaphora vs. VP ellipsis

The tripartite VP analysis extends to another domain exhibiting asymmetric light verb distribution: the contrast between VP anaphora (2) and VP ellipsis (3.2.2). Their differential behavior with SJVs is particularly revealing, especially in how they interact with causative versus inchoative variants.

Consider first causative SJVs, as exemplified below:

    1. (61)
    1. a.
    1. A-sya-wa
    2. A-company-top
    1. [VP
    2.  
    1. biru-o
    2. building-acc
    1. bakuha]-si-ta-kedo,
    2. explosion-do-pst-but
    1.  
    1. ‘Company A blew up the building, but …’
    1.  
    1. b.
    1. B-sya-wa
    2. B-company-top
    1. [VP
    2.  
    1. soo]
    2. so
    1. si-nakat-ta.
    2. do-neg-pst
    1. ‘Company B did not do so.’
    1.  
    1. c.
    1. B-sya-wa
    2. B-company-top
    1. [VP Δ ]
    2.  
    1. si-nakat-ta.
    2. do-neg-pst
    1. ‘Company B didn’t blow up the building.’

In the antecedent clause (61a), the VP contains the object biru ‘building’ and the SJV bakuha ‘explosion’. This VP serves as antecedent for both the anaphoric element soo in (61b) and the ellipsis site in (61c). Crucially, both VP anaphora and VP ellipsis consistently employ the light verb su in this causative context.

A different pattern emerges when the antecedent clause contains an inchoative SJV. Consider the following paradigm with the SJV zyoohatu ‘evaporation’:

    1. (62)
    1. a.
    1. Ano-ekitaii-wa
    2. that-liquid-top
    1. 100-do-de
    2. 100-degree-at
    1. [VP
    2.  
    1. ti
    2.  
    1. zyoohatu]-si-ta-kedo,
    2. evaporation-do-pst-but
    1.  
    1. ‘That liquid evaporated at 100 degrees, but … .’
    1.  
    1. b.
    1. kono-ekitai-wa
    2. this-liquid-top
    1. 100-do-de-mo
    2. 100-degree-at-even
    1. [VP
    2.  
    1. soo
    2. so
    1. ]
    2.  
    1. {nara/*si}-nakat-ta.
    2. become/do-neg-pst
    1. ‘This liquid did not do so, even at 100 degrees.’
    1.  
    1. c.
    1. kono-ekitai-wa
    2. this-liquid-top
    1. 100-do-de-mo
    2. 100-degree-at-even
    1. [VP
    2.  
    1. Δ
    2.  
    1. ]
    2.  
    1. {si/*nara}-nakat-ta.
    2. do/become-neg-pst
    1. ‘This liquid did not evaporate, even at 100 degrees.’

Here a clear asymmetry emerges. In VP anaphora, the inchoative antecedent requires nar and excludes su (62b), whereas VP ellipsis shows the opposite pattern, requiring su and excluding nar (62c). This sharply contrasts with the causative SJV paradigm in (61), where both constructions employ su.38

This systematic asymmetry is summarized in Table 3. While VP ellipsis invariably employs the light verb su irrespective of the transitivity specification of the antecedent SJV, VP anaphora exhibits a principled alternation between su and nar, reflecting the transitivity properties of its antecedent.

Table 3: Distribution of light verbs in VP anaphora versus VP ellipsis constructions.

causative SJV inchoative SJV
VP anaphora su nar
VP ellipsis su su

This asymmetric distribution emerges naturally from the proposed tripartite VP analysis without necessitating additional stipulations. First, the antecedent SJV has the following structure in our tripartite VP framework, where v is realized as Ø and Voice as su:

    1. (63)
    1. Antecedent SJV

As proposed in section 2, VP anaphora specifically targets VPRes, constructing an independent VP structure, as illustrated in (64):

    1. (64)
    1. VP anaphora

In VP anaphora constructions, VPsoo initially copies the semantic content of the antecedent VPRes. As soo functions as an adverbial element (Tanaka 2016), it cannot undergo head raising to v. Consequently, vcaus/inch lacks a local relationship with any verbal root and is realized as su or nar, explaining the transitivity-sensitive alternation in (61b) and (62b). This analysis correctly predicts that transitivity mismatches also arise with SJVs:

    1. (65)
    1. a.
    1. John-ga
    2. John-nom
    1. gamen-o
    2. screen-acc
    1. kakudai-si-yootosi-ta-ga,
    2. expansion-do-try-pst-but
    1. soo
    2. so
    1. nara-nakat-ta.
    2. become-neg-pst
    1. ‘John tried to expand the screen, but it didn’t do so.’
    1.  
    1. b.
    1. Mary-wa
    2. Mary-top
    1. katteni
    2. on.one’s.own
    1. gamen-ga
    2. screen-nom
    1. kakudai-si-ta-to
    2. expansion-do-pst-comp
    1. iw-ta-ga,
    2. say-pst-but
    1. zissai-wa
    2. actually-top
    1. watasi-ga
    2. I-nom
    1. soo
    2. so
    1. si-ta.
    2. do-pst
    1. ‘Lit.: Mary claimed that the screen expanded on its own, but actually I did so.’

Regarding VP ellipsis phenomena, we maintain that it targets the vP domain within the tripartite VP structure. This analysis yields two theoretically viable configurations: either through LF-copying (66a) (Fiengo & May 1994; Oku 1998; Sato 2015) or PF-deletion (66b) (Merchant 2001). Though we do not adjudicate between these frameworks, both capture the invariant realization of Voice as su in (61c) and (62c).

    1. (66)
    1. VP ellipsis
    1.  
    1. a.
    1. LF-Copy
    1.  
    1. b.
    1. PF-deletion

Under the LF-copying approach (66a), the vP node contains an empty category [e] whose internal structure is reconstructed through antecedent copying at LF. Voice remains structurally isolated, lacking a local relationship with any v head, and consequently surfaces as su per our VI rules. Under the PF-deletion approach (66b), the same vP structure is derived in overt syntax and subsequently deleted at PF, with Voice realized as su parallel to the antecedent structure in (63).

The contrastive behavior of VP anaphora and VP ellipsis thus receives a principled explanation in terms of the domains they target: VP anaphora targets VPRes, so that v is realized as su or nar according to the transitivity of the target clause, whereas VP ellipsis targets vP, stranding Voice, which is invariably realized as su. The two operations therefore target distinct domains, providing further support for distinguishing VPRes, vP, and VoiceP.

5 Verbal domains in Japanese

This section synthesizes the principal findings of this investigation on the verbal domain in Japanese. Throughout this study, I have defended a tripartite VP structure—comprising VPRes, vPcaus/inch, and VoiceP—based on evidence from diverse constructions: VP anaphora (VPA), VP ellipsis (VPE), stative predicate constructions (SPCs), NJVs, SJVs, deverbal compounds (DCs), and object honorification (OH).

The proposed architecture yields three configurational possibilities for verbal domain formation:

    1. (67)
    1. Three types of verbal domains

Configuration (67a) illustrates complete head movement, where VRes raises through v to Voice, forming a maximally integrated morphosyntactic complex. Configuration (67b) represents constructions where the VPRes layer is absent, resulting in a v-Voice complex. Configuration (67c) depicts partial head movement, where VRes raises to v but not to Voice, yielding two distinct morphosyntactic domains within the verbal architecture.

Within this framework, vcaus/inch has three realization patterns: (i) -S/-R (causative/ inchoative morphology), (ii) su/nar (transitivity-alternating light verbs), and (iii) -Ø (null morpheme). Voice has two patterns: (i) -Ø (null) or (ii) su (invariant light verb). These are encoded through the following VI rules:

    1. (68)
    1. VI rules for Voice (=(46))
    1.  
    1. a.
    1. Voice ↔ Ø / [ v___ ]Voice
    1.  
    1. b.
    1. Voice ↔ su
    1. (69)
    1. VI rules for vcaus (=(52))
    1.  
    1. a.
    1. vcaus ↔ -S / [ V[+Native]___ ]v {e.g., kowa- ‘break’, ag- ‘rise’, …}
    1.  
    1. b.
    1. vcaus ↔ Ø / [ V[−Native]___ ]v
    1.  
    1. c.
    1. vcaus ↔ su
    1. (70)
    1. VI rules for vinch (=(53))
    1.  
    1. a.
    1. vinch ↔ -R / [ V[+Native]___ ]v {e.g., kowa- ‘break’, ag- ‘rise’, …}
    1.  
    1. b.
    1. vinch ↔ Ø / [ V[−Native]___ ]v
    1.  
    1. c.
    1. vinch ↔ nar

The interaction between these morphological realization patterns and the configurational possibilities of the verbal domain yields four attested morphosyntactic patterns in Japanese, systematically summarized in Table 4.

Table 4: Verbal domain configurations and morphological realizations in Japanese.

Verbal Domain(s) Vocabulary Predicate Types
vcaus/inch Voice
a. V[+Native]-v-Voice -S/-R Ø NJVs
b. v-Voice su/nar Ø SPC, VPA
c. V[−Native]-v Voice Ø su SJVs (incl. VPE)
d. V[+Native]-v Voice -S/-R su DCs, OH

Pattern (a) involves an NJV root undergoing complete head movement. vcaus/inch surfaces as -S/-R while Voice remains null. Pattern (b) lacks a verbal root, with vcaus/inch realized as su/nar and Voice as null. Pattern (c) features an SJV root with partial head movement to v but not Voice, resulting in null v and su for Voice. Pattern (d) shows an NJV root with partial head movement, with v realized as -S/-R and Voice as su (DCs, OH).39

Beyond these attested patterns, our framework predicts two additional logical possibilities. Pattern (e) would have both v and Voice realized as null (Ø), requiring SJVs to undergo complete head movement to Voice, forming a V-v-Voice complex. This is systematically excluded in Japanese due to constraints on head movement with non-Native verbal roots. Pattern (f) would involve both v and Voice as overt light verbs (su/nar for v and su for Voice). This is precluded in standard environments because v-Voice locality triggers rule (68a), resulting in null Voice.

However, our framework predicts the emergence of pattern (f) in contexts where v-to-Voice movement is blocked by an intervening element as illustrated in (71), similar to what we observed in OH cases:

    1. (71)

This structural configuration might occur in the -tari construction, as exemplified in (72):

    1. (72)
    1. Kyoo-wa
    2. today-top
    1. [ronbun-o
    2. paper-acc
    1. kak-tari],
    2. write-tari
    1. [heya-o
    2. room-acc
    1. kirei-ni
    2. clean-prt
    1. si-tari]
    2. do-tari
    1. si-ta.
    2. do-pst
    1. ‘Today I wrote a paper and made the room clean.’

In this construction, two verb phrases are conjoined through the particle -tari, which attaches to the verbal stems—specifically, the NJV kak ‘write’ in the first conjunct and the light verb si of the stative predicate construction kirei-ni su ‘make clean’ in the second. In the -tari construction, the entire coordinated structure must be followed by the light verb su, resulting in the manifestation of two light verbs in sequence in (72).

One might analyze the second si in (72) as a tense-support dummy verb rather than a light verb. However, a suppletion diagnostic demonstrates that it is indeed a light verb. The potential form of su suppletively becomes deki(ru) ‘can,’ but crucially, this suppletion applies to light verb su but not to dummy verb su (Kageyama 1993). Consider a VN-su construction with the focus particle -sae ‘even’:

    1. (73)
    1. a.
    1.   benkyoo-si-sae
    2.   study-do-even
    1. si-ta.
    2. do-pst
    1. ‘(Someone) even studied.’
    1.  
    1. b.
    1.   benkyoo-deki-sae
    2.   study-can-even
    1. si-ta.
    2. do-pst
    1.   ‘(Someone) could even study.’
    1.  
    1. c.
    1. *benkyoo-si-sae
    2.   study-do-even
    1. deki-ta.
    2. can-pst

In (73a), the first si is a light verb and the second si is arguably a dummy verb. In the potential form (73b), only the light verb si can undergo suppletion to deki, while the dummy verb si cannot, as shown by the ungrammaticality of (73c). This diagnostic extends to (72): the second si also allows suppletion:

    1. (74)
    1. Kyoo-wa
    2. today-top
    1. [ronbun-o
    2. paper-acc
    1. kak-tari],
    2. write-tari
    1. [heya-o
    2. room-acc
    1. kirei-ni
    2. clean-prt
    1. si-tari]
    2. do-tari
    1. deki-ta.
    2. can-pst
    1. ‘Today I could write a paper and make the room clean.’

This confirms that the second si in (72) is a light verb, not a dummy verb.

Given this, one possible analysis would be to propose that -tari subcategorizes for vP as its complement (see also Smith & Kobayashi 2017) and possibly functions as a syntactic intervener that blocks the head movement of vcaus/inch to Voice, as represented in (75):

    1. (75)

If correct, -tari prevents v-Voice locality, potentially explaining why the phonological form su appears twice—once as the realization of vcaus and once as the realization of Voice—in these constructions.40 This in turn provides further morphological evidence for the separation of Voice and v, since the two heads are independently realized despite their phonological identity.

In sum, our framework accounts not only for patterns (a)–(d) but also for the absence of pattern (e) and restricted distribution of pattern (f), providing strong evidence for its validity. The morphosyntactic patterns in Japanese emerge systematically from the interaction of hierarchical structure, head movement, and VI rules.

A final consequence of the present proposal concerns the relationship between passive and active Voice in Japanese. Consider the passivization of VN-su constructions:

    1. (76)
    1. a.
    1. kaigi-ga
    2. meeting-nom
    1. enki-s-are-ta.
    2. postponement-do-pass-pst
    1. ‘The meeting was postponed.’
    1.  
    1. b.
    1. sono
    2. that
    1. teian-ga
    2. proposal-nom
    1. saiyo-s-are-ta.
    2. adoption-do-pass-pst
    1. ‘That proposal was adopted.’

These examples reveal a striking pattern: the light verb su, which realizes Voice, systematically co-occurs with the passive morpheme -rare. This co-occurrence is theoretically significant. If passivization involved the replacement of active Voice by passive Voice, we would expect su to be absent from passive constructions. Its obligatory presence instead indicates that passive Voice does not replace active Voice but rather projects as an independent functional layer above it, which converges with recent work treating passive Voice as a projection above active Voice (e.g., Asami 2025).41

6 Conclusion

This study has argued that Japanese verbal syntax requires a tripartite VP structure—VPRes, vPcaus/inch, and VoiceP—aligning with Pylkkänen (2002), Harley (2013; 2017), Legate (2014), and others. VP anaphora (2) showed that soo targets result states in VPRes, explaining transitivity mismatches that traditional approaches cannot. Light verb distributions (3) motivated separating v and Voice: v-type light verbs (su/nar) lexicalize v and alternate with transitivity, whereas Voice-type light verbs (invariant su) lexicalize Voice and do not—a distinction that resolves asymmetries problematic for bipartite approaches. The VP anaphora/VP ellipsis contrast (4) further showed that the two operations target different domains (VPRes vs. vP), and the typology in section 5 derived four attested patterns from the interaction of structure, head movement, and featural specification. The framework thus shows that the surface complexity of Japanese verbal morphology reflects a principled system, contributing to the cross-linguistic understanding of verbal architecture.

Abbreviations

acc = accusative, caus = causative, comp = complementizer, cop = copula, dat = dative, gen = genitive, ger = gerund, hon = honorific, imp = imperative, inch = inchoative, neg = negative, nom = nominative, pass = passive, prs = present, prt = particle, pst = past, tari = -tari (coordinating/listing particle), top = topic

Data availability

No dataset or supplementary files are associated with this paper; all data supporting the analysis are contained within the article.

Ethics and consent

Not applicable. This study did not involve human-subjects research requiring ethical approval.

Funding information

This work was supported by the Japan Society for the Promotion of Science (JSPS) KAKENHI Grant Number JP22K13104.

Acknowledgements

I would like to thank the three anonymous reviewers and the handling editor of Glossa for their constructive comments and helpful suggestions, which have significantly improved the paper. Earlier versions of this work were presented at the Morphology & Lexicon Forum 2023 (Shiga University) and at the 3rd meeting of the Setagaya Linguistics Circle (Tsuda University, 2023). I am grateful to the audiences at both venues for their valuable comments, and I would especially like to thank Ryoichiro Kobayashi for reading and commenting on an earlier draft. All remaining errors are my own.

Competing interests

The author has no competing interests to declare.

Notes

  1. The traditional view that the pro-form targets a VP, not just the verb, goes back to Nakau (1973), who noted that an identical object generally cannot co-occur with soo su. It has been observed, however, that an overt object may occur with the pro-form when interpreted contrastively. I follow Tateishi (1994) in analyzing such cases as involving scrambling of the contrastive object (focus-driven A-like movement; see fn. 8), leaving a variable inside VP. [^]
  2. As a reviewer notes, (4b) can be grammatical if an implicit agent (pro) is inferred in the target clause, yielding a ‘someone closed the next door’ reading. The asterisk reflects ungrammaticality under the intended inchoative reading. [^]
  3. The morpheme -e also appears in inchoative verbs (e.g., yak-e-ru ‘burn (vi)’). It is also well known that -e originates from historical conjugation class changes rather than a dedicated causative morpheme (Kuginuki 1996). In this paper, however, I follow standard synchronic analyses that group -e with causative exponents for present purposes (e.g., Jacobsen 1992; Miyagawa 1994; Harley 2008). [^]
  4. While both directions of transitivity mismatch are grammatically possible in Japanese, there appears to be an asymmetry in naturalness: causative-to-inchoative mismatches (e.g., (7)) seem more natural than inchoative-to-causative mismatches (e.g., (8)) (Ryoichiro Kobayashi, p.c.). This asymmetry may reflect processing factors, though its investigation lies beyond the present scope. [^]
  5. The reasons why English does not permit inchoative-causative mismatches will be discussed in section 2.4. [^]
  6. For expository purposes, I adopt a simplified representation throughout this paper in which verbal stems like sim- are analyzed as heading V directly. This abstracts away from a more articulated root-based structure where sim would be an acategorial root verbalized by V: [VP  V]. This distinction becomes particularly relevant for change-of-state verbs derived from adjectival roots (e.g., tuyo-m-e/ar-u ‘make/become strong’), where V verbalizes the adjectival root : [vP [VP  V(m)] v(e/ar)]. [^]
  7. A number of previous studies argue that A-movement of unaccusative subjects in Japanese is optional (e.g., Nakayama & Koizumi 1991; Yatsushiro 1999; Fukuda 2017; Asami & Tomioka 2025). The movement derivation presented in (16) is not intended to imply that such movement is obligatory. I present this derivation both to maintain parallelism with Bruening (2019) and to illustrate how the A-Movement Rule applies when the movement option is taken. The rule is independently required in cases discussed below in which the copied VPRes contains a variable (see the discussion of (21) and fn. 8). [^]
  8. A reviewer observes that the acceptability of (19b) improves if the subject is replaced with ano mado ‘that window’:
      1. (i)
      1. Mary-wa
      2. Mary-top
      1. watasi-ga
      2. I-nom
      1. kono
      2. this
      1. doa-o
      2. door-acc
      1. sim-e-ta-to
      2. close-caus-pst-comp
      1. iw-ta-ga,
      2. say-pst-but
      1. zissai-wa
      2. actually-top
      1. ano mado-ga
      2. that window-nom
      1. katteni
      2. on.one’s.own
      1. soo
      2. so
      1. nar-ta
      2. become-pst
      1. (no-da).
      2. comp-cop
      1. ‘Mary claimed that I closed the door, but that window actually did so on its own.’
    I take this improvement to follow from the clearer contrast between kono doa ‘this door’ and ano mado ‘that window’. Following Tateishi (1994: 60–61), I assume that in such cases the contrastive object in the antecedent clause undergoes focus-driven A-like movement, leaving a variable within VPRes, which soo copies. The variable is then bound by the contrastive subject in the target clause, yielding the non-identity interpretation, parallel to the topicalization case discussed below in the main text. [^]
  9. Alternatively, one could adopt a pronominal-variable-binding approach in which topics are base-generated in the left periphery and bind pro variables in argument positions. Under this approach, the antecedent VP would contain a pro variable in object position, bound by the base-generated topic. The target VP anaphor would then copy this structure, including a pro variable that can be bound by an independent topic in the target clause. I thank a reviewer for this valuable suggestion. [^]
  10. Following Bruening (2019: 28), A-movement heads are those that do not introduce an external argument (e.g., unaccusative and passive Voice). [^]
  11. As noted in fn. 7, if A-movement of unaccusative subjects is optional, the antecedent internal argument may remain in situ. In that case, the copied VPRes contains the full NP rather than a trace, and the theme argument of the target is identified with ‘the door’ directly. The main text pursues the movement derivation. [^]
  12. The same two options are independently motivated for VP ellipsis; see Bruening (2019: 35). [^]
  13. Alternatively, following Fiengo & May (1994), the trace can be treated as a pronoun (vehicle change; Bruening 2019: 35–36, fn. 31): t1 is then a pronoun coreferent with ‘the door’, converging on the same interpretation as (23). I thank a reviewer for this alternative. [^]
  14. Sakai et al. (2004) also suggest that these instances of su/nar do not function as dummy/helping verbs that merely support T. [^]
  15. See also Kikuchi & Takahashi (1991) for a small-clause analysis of the stative predicate constructions. [^]
  16. Sugioka (2002) distinguishes direct-argument from adjunct deverbal compounds and argues that only the latter directly combine with the light verb su. The compound ne-age/ne-agari shows properties of both types: ne ‘price’ may serve as an internal argument of age/agari ‘raise/rise’, yet the compound also readily combines with su. As this distinction does not affect the present analysis, I remain neutral regarding its classification. [^]
  17. Some deverbal compounds also have an accusative-marked use parallel to Sino-Japanese verbs (e.g., tabako-no ne-age-o si-ta ‘raised the price of cigarettes’). As noted in the main text above, I set this use and its internal structure aside, restricting attention to the non-accusative verbal use relevant here. [^]
  18. The root ag- is glossed ‘rise’ throughout for expository convenience. On the present analysis the bare root ag is in itself neutral between its transitive (ag-e ‘raise’) and intransitive (ag-ar ‘rise’) realizations, its valency being fixed by the functional structure; the gloss is therefore a pre-theoretical label, not a claim about the inherent valency of the root. (Historically, ag- is indeed attested as a transitive-causative ag(-u) ‘raise’ in Old Japanese, but nothing here turns on its diachronic status.) I thank a reviewer for pointing this out. [^]
  19. Not all Japanese adjectives lack transitive properties; certain adjectivals like suki ‘like’ and kirai ‘dislike’ can govern an accusative object alongside a nominative subject: John-ga Mary-o suki/kirai-na riyuu ‘the reason why John likes/dislikes Mary’ (e.g., Fukuda 2020). [^]
  20. In section 3.2.3, I will demonstrate that VNs function essentially as Vs, parallel to native Japanese verbs, but are featurally specified as [−Native]. [^]
  21. Syntactic independence does not preclude VN and su from forming a single prosodic word. According to Kishimoto (2013), predicative units in Japanese may arise through syntactic head movement or PF merger post–Vocabulary Insertion. I assume the VN-su construction reflects the latter: the two heads remain distinct for the purpose of VI, but undergo PF merger to yield one phonological word. [^]
  22. The structural independence of VNs and su is further supported by coordination phenomena (Miyamoto & Kishimoto 2016: 434), where VNs can be coordinated to the exclusion of su:
      1. (i)
      1. Sensei-wa
      2. teacher-top
      1. [sansei-mo
      2. [approval-also
      1. hantai-mo]
      2. disapproval-also]
      1. si-nakat-ta.
      2. do-neg-pst
      1. ‘The teacher neither approved nor disapproved.’
    In this construction, two VNs—sansei ‘approval’ and hantai ‘disapproval’—are conjoined via the focus particle -mo ‘also’, with the light verb su appearing external to the coordinated structure. This distributional pattern strongly indicates that VNs constitute independent syntactic constituents distinct from the light verb. See also Kuroda (2003) for related observations on the scope of -(s)ase in coordination. Additional evidence comes from nominalization patterns with -kata ‘way’: while NJVs directly combine with this suffix as in (iia), SJVs require the genitive marker (iib), indicating that VNs and light verbs do not form a morphological unit.
      1. (ii)
      1. a.
      1. manabi-kata
      2. learn-way
      1. ‘the way of learning’
      1.  
      1. b.
      1. benkyoo-*(no)
      2. study-gen
      1. si-kata
      2. do-way
      1. ‘the way of studying’
    [^]
  23. I remain agnostic as to whether the (complex) Voice head undergoes further syntactic head movement to higher functional projections such as T and C in Japanese. The presence and nature of V-to-T-to-C movement in Japanese remain a matter of ongoing debate in the literature. See Kobayashi (2023) for recent discussion of this issue and references therein. [^]
  24. We will revisit the vP ellipsis assumption in section 4, where we present additional evidence based on the interaction between VP anaphora and VP ellipsis. [^]
  25. An alternative analysis questions VP ellipsis in Japanese, proposing instead that these phenomena derive from argument ellipsis (AE) (e.g., Oku 1998; Sakamoto 2016). Under this approach, the contrast would suggest that AE applies to SJVs but not NJVs—attributable to the dual nominal-verbal status of SJVs. However, evidence from modification patterns suggests that SJVs directly combined with su (without case markers) do not function as nominals. Compare:
      1. (i)
      1. a.
      1.   John-ga
      2.   John-nom
      1. zyuuyoona/takusan-no
      2. important/many
      1. hatugen-o
      2. speaking.out-acc
      1. si-ta.
      2. do-pst
      1.  
      1. b.
      1. *John-ga
      2.   John-nom
      1. zyuuyoona/takusan-no
      2. important/many
      1. hatugen-si-ta.
      2. speaking.out-do-pst
      1.   ‘John made an important remark/many remarks.’
    The ungrammaticality of (b) shows that SJVs directly combined with su cannot be modified by adjectivals, suggesting they do not function as nominals in this context. This makes VP ellipsis a more plausible analysis for the patterns observed. [^]
  26. A reviewer observes that the inchoative compound ne-agar can also surface as a bare simplex form ne-ag-ar-ta (for many speakers, especially in central and western Japan), whereas its causative counterpart cannot (*ne-age-ta). Under the intervener refinement just sketched, this follows if the inchoative compound may be lexicalized as a simplex verb, bleeding the non-head intervener configuration that otherwise forces su, while the causative resists such lexicalization. A more strongly lexicalized parallel is ne-duku ‘take root’ (from ne ‘root’ and tuku ‘attach (vi)’), which lacks a causative counterpart (*ne-duke-ru) despite the existence of tuke-ru ‘attach (vt)’. Why only the inchoative lexicalizes in this way I leave for future research. [^]
  27. The identification of -S/-R as abstract exponents of causative and inchoative morphology draws upon the seminal work of Jacobsen (1992: 59), who formulates the following morphophonological generalization regarding transitivity alternations in Japanese: “Every suffix involved in transitive vs. intransitive oppositions containing an s is transitive, and affixes containing r are preponderantly intransitive.” [^]
  28. Miyagawa (1994; 1998) (see also Harley 2008) analyzes -s, -as, -e, etc. as contextual allomorphs of v, with -sase (both lexical and productive) as the elsewhere exponent. The present proposal extends this approach to a tripartite structure, using abstract morphemes -S/-R to capture root-conditioned patterns in NJVs. Crucially, however, in the present framework, the light verb su serves as the elsewhere exponent for vcaus. How the elsewhere-like distribution of -sase should be captured in the present system is left for future research. For related discussion of the interaction between causative morphology and nominalization in Japanese and its implications for argument structure, see Volpe (2005). [^]
  29. A reviewer asks about the status of ar ‘be’ in relation to the light verb nar, noting that ar can co-occur with adjectival predicates (examples adapted from the reviewer):
      1. (i)
      1. a.
      1. kabe-ga
      2. wall-nom
      1. akaku
      2. red
      1. ari-tuzuke-ru
      2. be-continue-prs
      1. tame-ni-wa
      2. sake-for-top
      1.  
      1. ‘in order for the wall to stay red, …’
      1.  
      1. b.
      1. heya-ga
      2. room-nom
      1. kirei-de
      2. clean-cop
      1. ari-tuzuke-ru
      2. be-continue-prs
      1. tame-ni-wa
      2. sake-for-top
      1.  
      1. ‘in order for the room to stay clean, …’
    I take ar to lie outside the su/nar alternation discussed in this paper. Whereas nar realizes vinch that introduces a change-of-state component, ar behaves as a copular element with no such semantics (Nishiyama 1999). It may either spell out a dedicated copular v or function as a dummy element (see Yamada 2023). [^]
  30. I thank a reviewer for drawing my attention to the connection with the literature in the Hale & Keyser tradition (Hale & Keyser 1993; 2002) and for the reference to Folli et al. (2005). [^]
  31. I am deeply indebted to two reviewers for drawing my attention to these mismatch patterns, which significantly enhanced the theoretical scope of the paper. [^]
  32. In Asami’s (2024) analysis, the relevant heads are labeled Cause and semantically contentless expletive Voice, which roughly correspond to vcaus and Voiceun in the present framework, respectively. It should also be noted that Asami treats the possessive relationship between the subject and the object (Taro and house in this case) as a pragmatic inference; the present analysis remains neutral on this issue. [^]
  33. A reviewer asks whether the inchoative (56b) has the same structure as (56a) in (57), differing only in carrying vinch instead of vcaus. I assume that it is a plain (intransitive) inchoative without Affect. On Asami’s (2024) analysis, Affect selects vcausP (and PassiveP) but is not extended to vinchP (his BecomeP), so the analysis of (56a) does not carry over. The subject is the nominative theme ie; Taro, the possessor of that theme (underlyingly Taro-no ie ‘Taro’s house’), appears as a topic, the precise derivation of which I leave for future research. [^]
  34. The denotation of vcaus in (14b) (2.3.2) includes an activity predicate, reflecting the canonical case where vPcaus merges with Voicetr and introduces an agentive Initiator. In the mismatch configurations discussed here—where vcaus combines with Voiceun and no Initiator is introduced—it is controversial whether activity(e) is still present. One possibility is that the activity component arises only when Voicetr introduces an agent, making it a property of the Voice–v complex rather than of vcaus alone. Under this view, vcaus in experiencer-subject causatives and transitive anticausatives, discussed below in the main text, would contribute only caus(e,s), with the event type left underspecified. Alternatively, activity(e) may be inherent to vcaus but interpreted differently without an Initiator. I leave this issue for future research. [^]
  35. I am grateful to a reviewer for raising this issue. [^]
  36. The present analysis follows Asami (2024) in treating Affect as covert in the relevant construction; this is not intended to imply that Affect is always covert. Bosse et al. (2012), for example, argue that Affect in the Japanese adversity passive projects above VoiceP and is overtly realized as -(r)are. [^]
  37. A reviewer suggests two possibilities for vinch + Voicetr configurations. The first involves what Kageyama (1996) terms anticausatives:
      1. (i)
      1. kami-ga
      2. paper-nom
      1. katteni
      2. on.one’s.own
      1. yabur-e-ta.
      2. tear-inch-pst
      1. ‘The paper tore by itself.’ (Kageyama 1996: 189)
    In this example, the subject might be simultaneously construed as the theme and the Initiator of the event through modification by the by oneself-type adverb. However, compatibility with by oneself adverbs indicates the absence of an external causer or agent rather than the presence of an external-argument-introducing head (e.g., Alexiadou et al. 2015: 21–22). Crucially, Kageyama (1996) analyzes such constructions as having no external argument syntactically. The second possibility involves subject honorification (SH):
      1. (ii)
      1. sensei-ga
      2. teacher-nom
      1. (wazato)
      2. on.purpose
      1. John-o
      2. John-acc
      1. o-tasuke-ni
      2. hon-help-prt
      1. nar-ta.
      2. become-pst
      1. ‘The teacher (hon) helped John (on purpose).’
    SH contains nar, which in the present analysis realizes vinch; and, as the reviewer notes, its subject of deference can be agentive, as shown by its compatibility with the agent-oriented adverb wazato ‘on purpose’, which points to Voicetr. These properties might therefore appear to instantiate a vinch + Voicetr pattern. This is not the only possible analysis, however. The agentivity of the SH subject may instead be licensed within the lexical verb domain. Ivana & Sakai (2007: 184, 187) propose such an analysis, on which the subject position is licensed within the domain of the main verb rather than within that of nar; recast in the present tripartite structure, it yields roughly the following layered representation (I leave the precise treatment of -ni open):
      1. (iii)
      1. [VoiceP [vP [honP [VoiceP sensei [vP John V(tasuke)-vcaus ] Voicetr ] hon]-ni vinch(nar)] Voiceun ]
    In (iii) there are two Voice domains. The agentive subject is licensed by the lower Voicetr, while the upper layer consists of vinch and Voiceun, as in the stative predicate construction (e.g., kirei-ni nar ‘to become clean’). Since the upper layer contains vinch + Voiceun and the lower one vcaus + Voicetr, neither instantiates a vinch + Voicetr configuration. The co-occurrence of an agentive subject with nar in SH therefore does not provide evidence for such a pattern. (iii) also addresses the reviewer’s related concern about head movement: if the honorific head hon intervened between v and Voice in SH just as in object honorification (51), it would block v-to-Voice movement and predict su by rule (46b), contrary to fact. Under (iii), hon projects over VoicetrP and is not located between v and Voice, so su is correctly not expected in SH. Finally, this layered analysis is compatible with another reviewer’s suggestion that the basic idea of honorification in Japanese is to suppress the will of the actor: the higher inchoative layer in (iii) can be understood as suppressing the volitionality of the subject, even though that subject remains agentive with respect to the lexical verb. This offers a plausible motivation for the inchoative character of the higher layer, while a fuller account is left for future work. [^]
  38. To put the contrast in (62b) and (62c) on a firmer footing, I informally consulted several native speakers in addition to relying on my own judgment. Although speakers differ in how sharply they judge the dispreferred variants, they agree on the direction of the contrast: in VP anaphora (62b), soo nar is preferred to soo su, whereas in VP ellipsis (62c), the preference reverses, the elided variant favoring su over nar. [^]
  39. The grouping of DCs with OH here is intended only in terms of morphological realization, not full structural identity. [^]
  40. I do not assume that -tari always selects vP as its complement. Consider for instance:
      1. (i)
      1. Kono-kenkyusitu-de-wa
      2. this-lab-in-top
      1. [inseetati-ga
      2. grad.students-nom
      1. eigo-de
      2. English-in
      1. ronbun-o
      2. paper-acc
      1. kak-tari]
      2. write-tari
      1. [gakubuseetati-mo
      2. undergrads-also
      1. eigo-de
      2. English-in
      1. ronbun-o
      2. paper-acc
      1. yom-tari]
      2. read-tari
      1. si-te-imasu.
      2. do-ger-prs
      1. ‘In this lab, graduate students write papers in English and undergraduate students also read papers in English.’ (Ryoichiro Kobayashi, p.c.)
    In this example, each conjunct contains a nominative-marked external argument, indicating that -tari can coordinate larger structures that include both vP and Voice projections. The crucial point here is that when -tari intervenes between v and Voice, it blocks head movement, resulting in pattern (f) with the realization of dual light verbs. [^]
  41. I thank a reviewer for drawing my attention to the relation between active Voice and passive Voice. [^]

References

Akimoto, Takayuki. 2018. The morphosyntax of transitivity in Japanese. Tokyo: Chuo University dissertation.

Alexiadou, Artemis & Anagnostopoulou, Elena & Schäfer, Florian. 2006. The properties of anticausatives crosslinguistically. In Frascarelli, Mara (ed.), Phases of interpretation, 187–212. Berlin, New York: De Gruyter Mouton. DOI:  http://doi.org/10.1515/9783110197723.4.187

Alexiadou, Artemis & Anagnostopoulou, Elena & Schäfer, Florian. 2015. External arguments in transitivity alternations: A layering approach. Oxford: Oxford University Press. DOI:  http://doi.org/10.1093/acprof:oso/9780199571949.001.0001

Asami, Daiki. 2024. Deriving and processing experiencer subject causatives. Glossa: a journal of general linguistics 9(1). 1–57. DOI:  http://doi.org/10.16995/glossa.16435

Asami, Daiki. 2025. Passive head only selects for agentive Voice in Japanese: A reply to Jo and Seo (2023). Journal of East Asian Linguistics 34. 205–240. DOI:  http://doi.org/10.1007/s10831-025-09294-4

Asami, Daiki & Tomioka, Satoshi. 2025. Psycholinguistic evidence for the optional movement of unaccusative subjects in Japanese. Syntactic Theory and Research 1(1). 12. DOI:  http://doi.org/10.16995/star.17600

Bosse, Solveig & Bruening, Benjamin & Yamada, Masahiro. 2012. Affected experiencers. Natural Language & Linguistic Theory 30. 1185–1230. DOI:  http://doi.org/10.1007/s11049-012-9177-1

Bouton, Lawrence. 1969. Identity constraints on the do-so rule. Research on Language & Social Interaction 1(2). 231–247. DOI:  http://doi.org/10.1080/08351816909389118

Bruening, Benjamin. 2019. Passive do so. Natural Language & Linguistic Theory 37. 1–49. DOI:  http://doi.org/10.1007/s11049-018-9408-1

Chomsky, Noam. 1993. A minimalist program for linguistic theory. In Hale, Kenneth & Keyser, Samuel Jay (eds.), The view from building 20, 1–52. Cambridge, MA: MIT Press.

Chomsky, Noam. 1995. The minimalist program. Cambridge, MA: MIT Press.

Chomsky, Noam. 2001. Derivation by phase. In Kenstowicz, Michael (ed.), Ken Hale: A life in language, 1–52. Cambridge, MA: MIT Press. DOI:  http://doi.org/10.7551/mitpress/4056.003.0004

D’Alessandro, Roberta & Franco, Irene & Gallego, Ángel J. (eds.) 2017. The verbal domain. Oxford: Oxford University Press. DOI:  http://doi.org/10.1093/oso/9780198767886.001.0001

Fiengo, Robert & May, Robert. 1994. Indices and identity. Cambridge, MA: MIT Press.

Folli, Raffaella & Harley, Heidi & Karimi, Simin. 2005. Determinants of event type in Persian complex predicates. Lingua 115(10). 1365–1401. DOI:  http://doi.org/10.1016/j.lingua.2004.06.002

Fukuda, Shin. 2017. Split intransitivity in Japanese is syntactic: Evidence for the Unaccusative Hypothesis from sentence acceptability and truth value judgment experiments. Glossa: a journal of general linguistics 2(1). 83. 1–41. DOI:  http://doi.org/10.5334/gjgl.268

Fukuda, Shin. 2020. Transitive nominals in Japanese and the syntax of predication. In Iwasaki, Shoichi & Strauss, Susan & Fukuda, Shin & Jun, Sun-Ah & Sohn, Sung-Ock & Zuraw, Kie (eds.), Japanese/Korean Linguistics 26, 129–139. Stanford: CSLI Publications.

Fukui, Naoki & Sakai, Hiromu. 2003. The visibility guideline for functional categories: Verb raising in Japanese and related issues. Lingua 113(4). 321–375. DOI:  http://doi.org/10.1016/S0024-3841(02)00080-3

Funakoshi, Kenshi. 2016. Verb-stranding verb phrase ellipsis in Japanese. Journal of East Asian Linguistics 25(2). 113–142. DOI:  http://doi.org/10.1007/s10831-016-9140-4

Grimshaw, Jane & Mester, Armin. 1988. Light verbs and theta-marking. Linguistic Inquiry 19(2). 205–232.

Hale, Kenneth & Keyser, Samuel Jay. 1993. On argument structure and the lexical expression of syntactic relations. In Hale, Kenneth & Keyser, Samuel Jay (eds.), The view from building 20, 53–109. Cambridge, MA: MIT Press.

Hale, Kenneth & Keyser, Samuel Jay. 2002. Prolegomenon to a theory of argument structure. Cambridge, MA: MIT Press. DOI:  http://doi.org/10.7551/mitpress/5634.001.0001

Halle, Morris & Marantz, Alec. 1993. Distributed morphology and the pieces of inflection. In Hale, Kenneth & Keyser, Samuel Jay (eds.), The view from building 20, 111–176. Cambridge, MA: MIT Press.

Hankamer, Jorge & Sag, Ivan. 1976. Deep and surface anaphora. Linguistic Inquiry 7(3). 391–428.

Harley, Heidi. 2008. On the causative construction. In Miyagawa, Shigeru & Saito, Mamoru (eds.), The Oxford handbook of Japanese linguistics, 20–53. New York: Oxford University Press.

Harley, Heidi. 2013. External arguments and the mirror principle: On the distinctness of voice and v. Lingua 125. 34–57. DOI:  http://doi.org/10.1016/j.lingua.2012.09.010

Harley, Heidi. 2017. The “bundling” hypothesis and the disparate functions of little v. In D’Alessandro, Roberta & Franco, Irene & Gallego, Ángel J. (eds.), The verbal domain, 3–28. Oxford: Oxford University Press. DOI:  http://doi.org/10.1093/oso/9780198767886.003.0001

Hasegawa, Nobuko. 2001. Causatives and the role of v: Agent, Causer, and Experiencer. In Inoue, Kazuko & Hasegawa, Nobuko (eds.), Linguistics and interdisciplinary research: Proceedings of the COE international symposium, 1–35. Chiba: Kanda University of International Studies.

Hayashi, Shintaro. 2015. Head movement in an agglutinative SOV language. Yokohama: Yokohama National University dissertation.

Hinds, John V. 1973. Some remarks on soo su-. Journal of Japanese Linguistics 2(1). 18–30. DOI:  http://doi.org/10.1515/jjl-1973-0103

Houser, Michael. 2010. The syntax and semantics of do so anaphora. Berkeley, CA: University of California dissertation.

Ikawa, Shiori. 2022. Agree feeds interpretation: Evidence from Japanese object honorifics. Syntax 25(4). 508–544. DOI:  http://doi.org/10.1111/synt.12242

Inoue, Kazuko. 1976. Henkei bunpoo to nihongo ge [Transformational grammar and Japanese II]. Tokyo: Taishukan.

Ivana, Adrian & Sakai, Hiromu. 2007. Honorification and light verbs in Japanese. Journal of East Asian Linguistics 16(2). 171–191. DOI:  http://doi.org/10.1007/s10831-007-9011-7

Jacobsen, Wesley M. 1992. The transitive structure of events in Japanese. Tokyo: Kurosio Publishers.

Kageyama, Taro. 1991. Light verb constructions and the syntax-morphology interface. In Nakajima, Heizo (ed.), Current English linguistics in Japan, 169–203. Berlin: Mouton de Gruyter. DOI:  http://doi.org/10.1515/9783110854213.169

Kageyama, Taro. 1993. Bumpoo to gokeisei [Grammar and word formation]. Tokyo: Hituzi Syobo.

Kageyama, Taro. 1996. Doosi imiron [Verbal semantics]. Tokyo: Kurosio Publishers.

Kandybowicz, Jason. 2015. On prosodic vacuity and verbal resumption in Asante Twi. Linguistic Inquiry 46(2). 243–272. DOI:  http://doi.org/10.1162/LING_a_00181

Kikuchi, Akira & Takahashi, Daiko. 1991. Agreement and small clauses. In Nakajima, Heizo & Tonoike, Shigeo (eds.), Topics in small clauses: Proceedings of Tokyo Small Clause Festival, 75–105. Tokyo: Kurosio Publishers.

Kiparsky, Paul. 1973. ‘Elsewhere’ in phonology. In Anderson, Stephen R. & Kiparsky, Paul (eds.), A festschrift for Morris Halle, 93–106. New York: Holt, Rinehart and Winston.

Kishimoto, Hideki. 2013. Verbal complex formation and negation in Japanese. Lingua 135. 132–154. DOI:  http://doi.org/10.1016/j.lingua.2012.11.007

Kobayashi, Ryoichiro. 2023. On the verb-raising analysis of non-constituent coordination in Japanese. The Linguistic Review 40(3). 405–418. DOI:  http://doi.org/10.1515/tlr-2023-2008

Kratzer, Angelika. 1996. Severing the external argument from its verb. In Rooryck, Johan & Zaring, Laurie (eds.), Phrase structure and the lexicon, 109–137. Dordrecht: Springer. DOI:  http://doi.org/10.1007/978-94-015-8617-7_5

Kratzer, Angelika. 2005. Building resultatives. In Maienborn, Claudia & Wöllstein, Angelika (eds.), Event arguments: Foundations and applications, 177–212. Berlin, Boston: Max Niemeyer Verlag. DOI:  http://doi.org/10.1515/9783110913798.177

Kuginuki, Toru. 1996. Kodai Nihongo no keitaihenka [Morphological changes in Old Japanese]. Tokyo: Izumi Shyoin.

Kuroda, Shige-Yuki. 2003. Complex predicates and predicate raising. Lingua 113(4–6). 447–480. DOI:  http://doi.org/10.1016/S0024-3841(02)00082-7

Larson, Richard K. 1988. On the double object construction. Linguistic Inquiry 19(3). 335–391.

Legate, Julie Anne. 2014. Voice and v: Lessons from Acehnese. Cambridge, MA: MIT Press. DOI:  http://doi.org/10.7551/mitpress/9780262028141.001.0001

Merchant, Jason. 2001. The syntax of silence: Sluicing, islands, and the theory of ellipsis. Oxford: Oxford University Press. DOI:  http://doi.org/10.1093/oso/9780199243730.001.0001

Miyagawa, Shigeru. 1989. Light verbs and the ergative hypothesis. Linguistic Inquiry 20(4). 659–668.

Miyagawa, Shigeru. 1994. (S)ase as an elsewhere causative. In Program of the conference on theoretical linguistics and Japanese language teaching: Seventh symposium on Japanese language. 61–76. Tokyo: Tsuda University.

Miyagawa, Shigeru. 1998. (S)ase as an elsewhere causative and the syntactic nature of words. Journal of Japanese Linguistics 16. 67–110. DOI:  http://doi.org/10.1515/jjl-1998-0105

Miyamoto, Tadao & Kishimoto, Hideki. 2016. Light verb constructions with verbal nouns. In Kageyama, Taro & Kishimoto, Hideki (eds.), Handbook of Japanese lexicon and word formation, 425–458. Berlin: De Gruyter Mouton. DOI:  http://doi.org/10.1515/9781614512097-016

Nakau, Minoru. 1973. Sentential complementation in Japanese. Tokyo: Kaitakusha.

Nakayama, Mineharu & Koizumi, Masatoshi. 1991. Remarks on Japanese subjects. Lingua 85(4). 303–319. DOI:  http://doi.org/10.1016/0024-3841(91)90001-L

Nishiyama, Kunio. 1998. The morphosyntax and morphophonology of Japanese predicates. Ithaca, NY: Cornell University dissertation.

Nishiyama, Kunio. 1999. Adjectives and the copulas in Japanese. Journal of East Asian Linguistics 8(3). 183–222. DOI:  http://doi.org/10.1023/A:1008395915524

Oku, Satoshi. 1998. A theory of selection and reconstruction in the minimalist perspective. Storrs, CT: University of Connecticut dissertation.

Oseki, Yohei. 2021. Bunsankeitairon to Nihongo no tadoosei kootai [Distributed Morphology and transitivity alternation in Japanese]. In Kishimoto, Hideki (ed.), Rekisikon kenkyu no gendaiteki kadai [Contemporary issues in lexicon research], 1–23. Tokyo: Kurosio Publishers.

Otani, Kazuyo & Whitman, John. 1991. V-raising and VP-ellipsis. Linguistic Inquiry 22(2). 345–358.

Pylkkänen, Liina. 2002. Introducing arguments. Cambridge, MA: MIT dissertation.

Pylkkänen, Liina. 2008. Introducing arguments. Cambridge, MA: MIT Press.

Ramchand, Gillian. 2008. Verb meaning and the lexicon: A first-phase syntax. Cambridge: Cambridge University Press. DOI:  http://doi.org/10.1017/CBO9780511486319

Ross, John Robert. 1972. Act. In Davidson, Donald & Harman, Gilbert (eds.), Semantics of natural language, 70–126. Dordrecht: Reidel.

Saito, Mamoru. 1985. Some asymmetries in Japanese and their theoretical implications. Cambridge, MA: MIT dissertation.

Saito, Mamoru & Hoshi, Hiroto. 2000. The Japanese light verb construction and the minimalist program. In Martin, Roger & Michaels, David & Uriagereka, Juan (eds.), Step by step: Essays on minimalist syntax in honor of Howard Lasnik, 261–298. Cambridge, MA: MIT Press.

Sakai, Hiromu & Ivana, Adrian & Zhang, Chao. 2004. The role of light verb projection in transitivity alternation. English Linguistics 21(2). 348–375. DOI:  http://doi.org/10.9793/elsj1984.21.348

Sakamoto, Yuta. 2016. Phases and argument ellipsis in Japanese. Journal of East Asian Linguistics 25(3). 243–274. DOI:  http://doi.org/10.1007/s10831-016-9145-6

Sato, Yosuke. 2015. Argument ellipsis in Javanese and voice agreement. Studia Linguistica 69(1). 58–85. DOI:  http://doi.org/10.1111/stul.12029

Schäfer, Florian. 2025. Anticausatives in transitive guise. Natural Language & Linguistic Theory 43. 421–475. DOI:  http://doi.org/10.1007/s11049-024-09612-w

Si, Fuzhen. 2021. Towards a cartography of light verbs. In Si, Fuzhen & Rizzi, Luigi (eds.), Current issues in syntactic cartography: A crosslinguistic perspective, 217–242. Amsterdam: John Benjamins. DOI:  http://doi.org/10.1075/la.267.10si

Smith, Ryan Walter & Kobayashi, Ryoichiro. 2017. Focusing on coordination: The case of Japanese -toka and -tari. In Erlewine, Michael Yoshitaka (ed.), Proceedings of GLOW in Asia XI, vol. 2. 204–215. Cambridge, MA: MIT Working Papers in Linguistics.

Sugioka, Yoko. 2002. Incorporation vs. modification in deverbal compounds. In Akatsuka, Noriko M. & Strauss, Susan (eds.), Japanese/Korean Linguistics 10, 495–508. Stanford, CA: CSLI Publications.

Tanaka, Hideharu. 2016. The derivation of soo-su: Some implications for the architecture of Japanese VP. In Kenstowicz, Michael & Levin, Ted & Masuda, Ryo (eds.), Japanese/Korean Linguistics 23, 265–279. Stanford, CA: CSLI Publications.

Tateishi, Koichi. 1994. The syntax of ‘subjects’. Stanford, CA & Tokyo: CSLI Publications & Kurosio Publishers.

Volpe, Mark J. 2005. Japanese morphology and its theoretical consequences: Derivational morphology in Distributed Morphology. Stony Brook, NY: Stony Brook University dissertation.

Yamada, Akitaka. 2023. Looking for default vocabulary insertion rules: Diachronic morphosyntax of the Japanese addressee-honorification system. Glossa: a journal of general linguistics 8(1). 1–47. DOI:  http://doi.org/10.16995/glossa.8338

Yatsushiro, Kazuko. 1999. Case licensing and VP structure. Storrs, CT: University of Connecticut dissertation.