You are on page 1of 14

Bilingualism: Language and When left is right: The role of typological

Cognition
similarity in multilinguals’ inhibitory
cambridge.org/bil control performance
Sarah Von Grebmer Zu Wolfsthurn1,2 , Anna Gupta3, Leticia Pablos1,2
Research Article and Niels O. Schiller1,2
1
Cite this article: Von Grebmer Zu Wolfsthurn Leiden University Center for Linguistics (LUCL), Reuvensplaats 3-4, 2311 BE Leiden, The Netherlands; 2Leiden
S, Gupta A, Pablos L, Schiller NO (2023). When Institute for Brain and Cognition (LIBC), LUMC, PO Box 9600, 2300 RC Leiden, The Netherlands and 3University of
left is right: The role of typological similarity in Konstanz, Universitätsstraße 10, 78464 Konstanz, Germany
multilinguals’ inhibitory control performance.
Bilingualism: Language and Cognition 26,
165–178. https://doi.org/10.1017/ Abstract
S1366728922000426 Both inhibitory control and typological similarity between two languages feature frequently in
Received: 12 October 2021
current research on multilingual cognitive processing mechanisms. Yet, the modulatory effect
Revised: 4 May 2022 of speaking two typologically highly similar languages on inhibitory control performance
Accepted: 27 May 2022 remains largely unexplored. However, this is a critical issue because it speaks directly to the
First published online: 29 July 2022 organisation of the multilingual’s cognitive architecture. In this study, we examined the influ-
ence of typological similarity on inhibitory control performance via a spatial Stroop paradigm
Keywords:
inhibitory control performance; typological in native Italian and native Dutch late learners of Spanish. Contrary to our hypothesis, we did
similarity; spatial Stroop task; late language not find evidence for a differential Stroop effect size for the typologically similar group
learners (Italian–Spanish) compared to the typologically dissimilar group (Dutch–Spanish). Our
results therefore suggest a limited influence of typological similarity on inhibitory control per-
Address for correspondence:
Sarah Von Grebmer Zu Wolfsthurn formance. The study has critical implications for characterising inhibitory control processes in
Leiden University Center for Linguistics multilinguals.
Reuvensplaats 3-4, 2311 BE Leiden, The
Netherlands
s.von.grebmer.zu.wolfsthurn@hum.leidenuniv.nl
1. Introduction
A remarkable feature of multilingual speakers is the ability to engage with several acquired
languages, seemingly without effort. In this paper, we will broadly refer to multilinguals as
those language users who have acquired one or more non-native language(s) in addition to
their native language, L1 (Cenoz, 2013; De Groot, 2017). Over the past decades, numerous
studies have attempted to capture the complexity of the multilingual experience. In particular,
they focused on the cognitive, structural and functional consequences of managing several lan-
guages in the brain (Abutalebi & Green, 2007; Bialystok, Craik, & Luk, 2012; Green, 1998;
Kroll, Dussias, Bice, & Perrotti, 2015; Mosca, 2017; Pliatsikas, 2020; Schwieter, 2016;
Sebastián-Gallés & Kroll, 2003).
A well-established aspect of the cognitive architecture of multilingualism is the parallel acti-
vation of languages across a range of proficiency levels, language combinations and linguistic
domains (Blumenfeld & Marian, 2013; Colomé, 2001; Costa, Caramazza, & Sebastián-Gallés,
2000; Dijkstra, Van Heuven, & Grainger, 1998; Guo & Peng, 2006; Hoshino & Thierry, 2011).
In order to successfully mitigate parallel activation and to ultimately select the appropriate
target language, multilinguals must employ a language control mechanism on the non-target
language (Abutalebi & Green, 2007; Christoffels, Firk, & Schiller, 2007; Costa & Santesteban,
2004; Declerck, Koch, Duñabeitia, Grainger, & Stephan, 2019; Green, 1998). Here, language
control is conceptualised as a collection of control mechanisms applied to multilingual speech
production and comprehension (Abutalebi, 2008; Green & Abutalebi, 2013). From a theoret-
ical point of view, this notion is featured in Green’s (1998) Inhibitory Control (IC) model of
language control, which postulates that the non-target language needs to be suppressed prior
to the linguistic output.
© The Author(s), 2022. Published by The exact nature of the mechanisms underlying language control is yet to be established.
Cambridge University Press. This is an Open There is a substantial amount of evidence suggesting that language control is strongly asso-
Access article, distributed under the terms of
ciated with domain-general inhibitory control, also termed cognitive control or executive con-
the Creative Commons Attribution licence
(http://creativecommons.org/licenses/by/4.0/), trol (Bialystok et al., 2012; Declerck, Meade, Midgley, Holcomb, Roelofs, & Emmorey, 2021;
which permits unrestricted re-use, distribution Festman, Rodriguez-Fornells, & Münte, 2010). Inhibitory control is an executive function
and reproduction, provided the original article used to regulate and inhibit irrelevant information with respect to thoughts or behaviour, as
is properly cited. well as switching attention (Diamond, 2013; Miyake, Friedman, Emerson, Witzki, Howerter,
& Wager, 2000). Some studies indicate that language control impacts executive functions –
for example, inhibitory control (Bialystok, 2010; Bialystok, Craik, Klein, & Viswanathan,
2004a; Green & Abutalebi, 2013; Kroll & Bialystok, 2013; Miyake et al., 2000; Wiseheart,
Viswanathan, & Bialystok, 2016). Critically, evidence further suggests that language control

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


166 Sarah Von Grebmer Zu Wolfsthurn et al.

may share some underlying processing mechanisms with inhibi- inhibitory control performance is reflected in the spatial Stroop
tory control (Declerck et al., 2021; Green, 1998; Linck, Hoshino, effect, which describes the quantitative difference in RTs between
& Kroll, 2005; Weissberger, Gollan, Bondi, Clark, & Wierenga, congruent and incongruent trials (Hilbert et al., 2014; La Heij, Van
2015), although this notion is still debated (Branzi, Della Rosa, der Heijden, & Plooij, 2001; Marian, Blumenfeld, Mizrahi, Kania,
Canini, Costa, & Abutalebi, 2016; Calabria, Hernandez, Branzi, & Cordes, 2013; Roelofs, 2021; Van Heuven et al., 2011). Drawing
& Costa, 2012). parallels between the Simon task, a smaller Stroop effect is reported
In the current study, we investigated the impact of multilin- to indicate better inhibitory control performance (Bialystok &
gualism on inhibitory control performance. More specifically, Martin, 2004; Costa, Hernández, & Sebastián-Gallés, 2008;
we examined whether the typological similarity between lan- Heidlmayr, Moutier, Hemforth, Courtin, Tanzmeister, & Isel,
guages of a multilingual plays a role in modulating inhibitory con- 2014; Pardo, Pardo, Janer, & Raichle, 1990).
trol performance. Typological similarity, also termed typological In the current study, the critical question we sought to answer
distance or language similarity, refers to linguistic and structural was the following: does typological similarity between the two lan-
(dis)similarities across different languages spoken by multilinguals guages significantly modulate inhibitory control performance in
(Foote, 2009; Putnam, Carlson, & Reitter, 2018; Westergaard, multilinguals? A relevant theoretical framework for this particular
Mitrofanova, Mykhaylyk, & Rodina, 2017). For example, Italian question is the Conditional Routing Model (CRM) by Stocco,
and Spanish may be considered as more typologically similar lan- Yamasaki, Natalenko, and Prat (2014). The model is based
guages compared to language pairs such as Dutch and Spanish upon the notion that the multilingual experience dynamically
because of the larger degree of overlap in morphosyntax, gender impacts domain-general executive functions, including inhibitory
systems and cognates (Paolieri, Padilla, Koreneva, Morales, & control, as a result of the parallel activation of the languages
Macizo, 2019; Schepens, Dijkstra, & Grootjen, 2012; Serratrice, (Bialystok & Martin, 2004; Festman et al., 2010). Here, the
Sorace, Filiaci, & Baldo, 2012). model postulates that executive functions are effectively trained
Several studies have focused on the modulating effects of typo- over time (Kroll & Bialystok, 2013; Yamasaki et al., 2018),
logical similarity on language control – for example, within the which results in a strengthening of the neural circuits underlying
context of a classical Stroop paradigm (Brauer, 1998; Coderre, these executive functions. When the languages within a multilin-
Van Heuven, & Conklin, 2013; Van Heuven, Conklin, Coderre, gual system are highly typologically similar, one may predict a
Guo, & Dijkstra, 2011). However, studies directly investigating higher degree of cross-language interference (Cenoz, 2001;
the effect of typological similarity on domain-general inhibitory Chen, Zhao, Zhaxi, & Liu, 2020; De Bot, 2004). In turn, this
control performance are scarce (but see Bialystok, Craik, Grady, implies that speakers of these languages develop better inhibitory
Chau, Ishii, Gunji, & Pantev, 2005; Linck et al., 2005; Yamasaki, control skills compared to speakers of typologically less similar
Stocco, & Prat, 2018). Typical experimental paradigms to explore languages (Yamasaki et al., 2018). Therefore, the CRM provides
domain-general inhibitory control are the Simon task (Bialystok us with a testable prediction for the effect of typological similarity
et al., 2004a, 2005; Simon & Small, 1969), and the spatial on inhibitory control performance: speakers of typologically simi-
Stroop task (Hilbert, Nakagawa, Bindl, & Bühner, 2014; Lu & lar languages should exhibit a better inhibitory control perform-
Proctor, 1995; Luo & Proctor, 2013). The core feature of the ance compared to speakers of typologically less similar
Simon task is a conflict between the physical location of a stimu- languages. Applied to the context of a spatial Stroop task used
lus and the response, e.g., a stimulus appearing on the right side in this study, speakers of typologically similar languages (e.g.,
of a screen while the corresponding response button is located on Italian–Spanish) should therefore show a smaller Stroop effect
the left side. The Simon effect quantifies the difference in compared to speakers of typologically less similar languages
response times (RTs) between trials in which stimulus and (e.g., Dutch–Spanish speakers).
response location match and trials in which stimulus and
response location mismatch. Typically, longer RTs are linked to
1.1. The current study
the mismatch trials. Accordingly, a smaller Simon effect reflects
better inhibitory control performance, whereas a larger Simon We explored the modulatory role of typological similarity on
effect reflects lower inhibitory control performance (Bialystok, inhibitory control performance in a spatial Stroop task (hereafter
Craik, Klein, & Viswanathan, 2004b). simply Stroop task) in two groups of speakers with differing
In this study, we used the spatial Stroop task (Hilbert et al., degrees of typological similarity. Participants were native Italian
2014; Lu & Proctor, 1995), which is a combination of the learners of Spanish, and native Dutch learners of Spanish. On
Simon task and the classical colour-word Stroop task the basis of typological work by Schepens et al. (2012) and Van
(MacLeod, 1992; Stroop, 1935). While the classical Stroop task der Slik (2010), we defined our Italian–Spanish group as our typo-
involves the naming of a colour-word written in either the match- logically similar group, and our Dutch–Spanish group as our
ing ink colour (congruent trial), e.g., the word RED written in red typologically dissimilar group. All participants had a Spanish pro-
ink, or the mismatching ink colour (incongruent trial), e.g., the ficiency level in the B1/B2 range within the CEFR framework
word RED written in blue ink, the spatial Stroop task focuses (Council of Europe, 2001). We followed a spatial Stroop paradigm
on spatial stimulus-stimulus conflicts. The basic feature of the inspired by Hilbert et al. (2014), who used the location words
spatial Stroop task is that a target word (“left”, “right”, “up”, “left”, “right”, “up” and “down” to study the Stroop effect in
“down”) either matches its location on the screen, e.g., LEFT native speakers of German (see also Hodgson, Parris, Gregory,
shown on the left side of the screen (congruent trial), or it does & Jarvis, 2009; Lu & Proctor, 1995; Shor, 1970). In our paradigm,
not match its location on the screen, e.g., LEFT shown on the we exploited the conflict between the target word and the location
right side of the screen (incongruent trial). The key to success of the target word on the screen – for example, the Spanish loca-
in this task is to inhibit the irrelevant spatial stimulus property tion word [izquierda] “left” displayed on the right side of the
(e.g., the location of the word) and to instead focus on the relevant screen, or the Spanish word [derecha] “right” displayed on the
target stimulus property (the target word itself). In this task, left side of the screen. The translation equivalents for [izquierda]

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


Bilingualism: Language and Cognition 167

“left” and [derecha] “right” are “sinistra” and “destra” in Italian, level of Spanish at Pompeu Fabra University (Barcelona, Spain).
and “links” and “rechts” in Dutch, respectively. In the congruent Mean age of the Italian–Spanish group was 27.12 years (SDage =
condition, the target word and the target word location matched. 4.08). Our recruitment criteria for this group were the following:
In contrast, in the incongruent condition, the target word and the no additional language learnt before the age of three, age of acqui-
target word location did not match. We measured accuracy and sition of Spanish from fourteen years onwards, a maximum time
RTs during this task. Post-experiment, we calculated the Stroop spent in a Spanish-speaking country of no longer than one year,
effect by subtracting the RTs for congruent trials from RTs for no psychological, neurological, visual, auditory, or language-related
incongruent trials. Importantly, we employed an equiprobable impairments; and finally, an age range between 18 and 35 years.
Stroop task design, whereby the probability of each condition For the Dutch–Spanish group, we recruited and tested 25 healthy,
occurring in the subsequent trial is identical. Within the frame- right-handed native speakers of Dutch (16 females) with a B1/B2
work of the Dual Mechanisms of Control (DMC) model level of Spanish at Leiden University (Leiden, The Netherlands).
(Braver, 2012), an equiprobable Stroop task design is linked to a Mean age of the Dutch–Spanish group was 22.84 years (SDage =
proactive control strategy. At the core of this particular strategy 3.05). Our recruitment criteria were identical to the Italian–
is the maintenance of goal-relevant information over time to suc- Spanish group, with the cap on maximum time spent in a
ceed at the task (Braver, 2012; Gonthier, Braver, & Bugg, 2016). Spanish-speaking country less stringent due to the testing location.
Therefore, our Stroop task taps not only into inhibitory control Data from the LEAP-Q was analysed to establish a detailed linguis-
performance per se, but also into the cognitive mechanisms of tic profile of each participant. See Appendix A and Appendix B for
monitoring the task. an overview of the profiles for the Italian–Spanish group and the
Dutch–Spanish group, respectively.
Research questions
Our research questions were the following: first, is there a differ- LEAP-Q: Italian–Spanish group
ence in terms of RTs as a function of typological similarity (typo- With respect to their linguistic profile in Spanish, two participants
logically similar vs. typologically dissimilar)? Secondly, connected acquired Spanish as first foreign language, whereas eighteen parti-
to this first question, is the Stroop effect larger for one group com- cipants acquired Spanish as second foreign language. Spanish was
pared to the other, thereby reflecting an effect of typological simi- the third foreign language for ten participants, and three partici-
larity on inhibitory control performance? pants acquired Spanish as fourth foreign language (Appendix A).
The mean age of acquisition (AoA) of Spanish was 23.93 years
Hypotheses (SD = 5.07). On average, participants reported to be fluent in
Based on the literature outlined above, we first predicted a Stroop Spanish at the age of 24.88 years (SD = 4.48), to have started read-
effect for both the Italian–Spanish group and the Dutch–Spanish ing in Spanish at the age of 24.36 years (SD = 4.91) and to be fluent
group. Behaviourally speaking, this would be reflected in higher readers by the age of 24.24 (SD = 4.82). On average, participants
accuracy and shorter RTs for congruent trials compared to incon- spent 0.46 years (SD = 0.343) in a Spanish-speaking country and
gruent trials. Next, in line with the CRM (Stocco et al., 2014) we had learnt Spanish for 0.93 years (SD = 1.17) either at school as a
hypothesised overall shorter RTs for the Italian–Spanish group foreign language, or as a language course in Spain. Twenty-five par-
compared to the Dutch–Spanish group. Finally, we expected a dif- ticipants were completing or had completed a formal Spanish lan-
ference in inhibitory control performance as a function of typo- guage course that was not part of the school curriculum shortly
logical similarity: we expected an interaction effect of condition before or upon their arrival in Spain (mean length of course:
(congruent vs. incongruent) and typological similarity (typologi- 0.53 years, SD = 0.889 years). Finally, participants quantified their
cally similar vs. typologically dissimilar) on Stroop effect sizes. A current daily exposure to Spanish as 40% (SD = 18.37%) of the
smaller Stroop effect for the Italian–Spanish group would imply time with respect to the other languages spoken. In terms of dom-
that the overall inhibitory control performance is better for the inance, thirteen participants classified Spanish as their most dom-
typologically similar languages compared to the less typologically inant language after Italian, fourteen participants as their second
similar Dutch–Spanish group. In turn, this would support the most dominant language after Italian, five participants as their
CRM (Stocco et al., 2014). third most dominant language after Italian, and one participant
as their fourth most dominant language after Italian. On a ten-
point scale, ten being maximally proficient, participants rated
2. Methods
their speaking proficiency at 5.95 (SD = 2.02), comprehension pro-
In addition to the spatial Stroop task, we asked participants to com- ficiency at 7.20 (SD = 1.71) and their reading proficiency at 7.33
plete the Language Experience and Proficiency Questionnaire, (SD = 1.47).
LEAP-Q (Marian, Blumenfeld, & Kaushanskaya, 2007). The
LEAP-Q is a questionnaire designed to obtain a measure for the LEAP-Q: Dutch–Spanish group
linguistic profile of our participants in terms of their proficiency In this group, nine participants stated that they acquired Spanish
levels and experiences with the languages within their multilingual as their second foreign language, nine participants as their third
system (Marian et al., 2007). Finally, participants also completed foreign language and seven participants as their fourth foreign
the Lextale-Esp (Izura, Cuetos, & Brysbaert, 2014), a lexical deci- language (Appendix B). Mean AoA of Spanish was 17.84 years
sion task to establish vocabulary size in Spanish, for descriptive (SD = 3.16). People stated to be fluent in Spanish on average at
purposes. the age of 19.6 years (SD = 2.52), that they started reading in
Spanish at the age of 18.44 years (SD = 3.24), and that they were
on average fluent in reading by 19.76 years (SD = 3.41). Eighteen
2.1. Participants
out of the twenty-five participants spent on average 0.57 years
For the Italian–Spanish group, we recruited and tested 33 healthy, (SD = 0.66) in a Spanish-speaking country (e.g., Spain, Argentina,
right-handed native speakers of Italian (26 females) with a B1/B2 Colombia, Mexico). Compared to the other languages, participants

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


168 Sarah Von Grebmer Zu Wolfsthurn et al.

quantified their daily exposure to Spanish with 12.96% (SD = Stroop task
10.07). Critically, two participants reported Spanish as their second The procedure for the Stroop task was the same for both groups.
most dominant language, nineteen as their third most dominant, Participants were asked to focus on the target word while ignoring
three as their fourth most dominant and one participant as their the location of the target word on the screen and to respond to the
fifth most dominant language following Dutch. On a ten-point target word via button presses. The trial procedure was as follows:
scale (ten being maximally proficient), participants reported an first, participants saw a black fixation cross for 500 ms in the cen-
average speaking proficiency in Spanish of 6.4 (SD = 1.47), a com- tre of a white screen. Next, they saw a target word appear on either
prehension proficiency of 7.08 (SD = 1.32) and a reading profi- the left or right side of the screen along the horizontal midline in
ciency of 7.08 (SD = 1.22). These ratings are highly comparable Spanish. This target word was either [izquierda] “left” or [dere-
with the Italian–Spanish group. cha] “right”. The target word was visible on the screen until par-
ticipants responded or for a maximum display time of 1,000 ms
(Figure 1). The next trial was initiated after participants’ response,
2.2. Materials and design or if the response time limit was reached. Half of the trials were
Prior to the experiment, participants completed the LEAP-Q congruent trials, where the target word matched the location on
(Marian et al., 2007) at home to reduce self-report biases fre- the screen. The other half of the trials were incongruent trials,
quently induced in laboratory settings (Rosenman, Tennekoon, where the target word and the location on the screen did not
& Hill, 2011). During the experiment, we first asked participants match. There were 24 trials for each target word (izquierda/ dere-
to complete the Lextale-Esp (Izura et al., 2014), followed by the cha) and target location on the screen (left side/right side),
Stroop task. amounting to 48 trials for the congruent condition and 48 trials
for the incongruent condition. Prior to the start of the main
experimental round, there was a short practise round to familiar-
Tasks and stimuli: Lextale-Esp ise participants with the task procedure. Trial order was rando-
We administered the Lextale-Esp to establish vocabulary size in mized in the practice round and in the main experimental round.
Spanish. The task was programmed in E-prime2 (Schneider,
Eschman, & Zuccolotto, 2002), using the exact same stimuli as
in the original version by Izura et al. (2014). 3. Results
3.1. Data exclusion
Tasks and stimuli: Stroop task.
We administered the Stroop task to measure inhibitory control For the Italian–Spanish group, data from one participant were lost
performance in our Italian–Spanish speakers and Dutch– due to a technical failure. Therefore, we included 32 datasets in
Spanish speakers. We again generated an E-prime2 (Schneider the analysis. In contrast, for the Dutch–Spanish group we
et al., 2002) script for this task. The target words were the written included all 25 datasets in the analysis, adding to a total of 57
Spanish words [izquierda] “left” and [derecha] “right”. datasets.

3.2. Data analysis


2.3. Procedure
We analysed our behavioural data using R, Version 4.0.3 (R Core
Prior to initiating the experiment, participants were provided with
Team, 2020) in RStudio, Version 1.4.1106. We employed a single
an information sheet and the opportunity to ask clarification
trial generalised linear mixed effects modelling approach using
questions. Then, participants signed the consent form in compli-
the lme4 package (Bates, Mächler, Bolker, & Walker, 2014;
ance with the ethics code for linguistic research at the Faculty of
Bates, Mächler, Bolker, Walker, Christensen, Singmann, Dai,
Humanities at Leiden University. Before each task, we provided
Scheipl, Grothendieck, Green, Fox, Bauer, & Krivitsky, 2020).
participants with written task instructions in Spanish. Upon ter-
We first modelled the outcome variables accuracy and RTS separ-
mination of all tasks, participants were provided with a debrief
ately for each individual group. Next, we pooled our data from
sheet, they signed the final consent form and received a monetary
both groups for a group comparison analysis to study potential
compensation for their participation.

Lextale-Esp
The Lextale-Esp procedure was identical for both groups. We
asked participants to indicate via a button press whether the string
corresponded to a Spanish word (e.g., [secuestro] “kidnapping”)
or a pseudoword (e.g., plaudir). Participants were instructed
that incorrectly assigning a word status to a pseudoword and
vice versa would lead to a deduction in the score. The trial pro-
cedure was as follows: first, a black fixation cross was displayed
for 1,000 ms on a white screen. Then, a letter string correspond-
ing to either a word or a pseudoword was displayed in the centre
of the screen. The letter string remained on the screen until the
participants’ response. After the participants’ response, the next
trial was initiated. Sixty trials were Spanish word trials, whereas
thirty were pseudoword trials. Trial order was randomised so Figure 1. Example trial procedure for a congruent trial followed by an incongruent
that each participant was presented with a unique trial order. trial.

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


Bilingualism: Language and Cognition 169

effects of typological similarity on Stroop effect sizes. For both the from the percentage of correctly identified words (Izura et al.,
individual group analyses and the group comparison analysis, we 2014). For the Italian–Spanish group, the mean Lextale-Esp score
applied the following model fitting procedure: first, we con- was 26.30 (SD = 14.04). Large individual differences were evident
structed a theoretically plausible model with a maximal random from the range of scores, which was between −7.37 to 49.30. In
effects structure as supported by our data (Barr, contrast, the mean Lextale-Esp score for the Dutch–Spanish
Levy, Scheepers, & Tily, 2013; Matuschek, Kliegl, Vasishth, group was 22.69 (SD = 17.19). The range was from −11.92 to
Baayen, & Bates, 2017). In our case, the maximal model was a 54.73, which yielded similar large individual differences between
random-intercept and random-slope model for both accuracy participants. A two-sample t-test yielded no significant statistical
and RTs. In the case of non-convergence or singular fit, we sim- difference in LexTALE-Esp scores between the two groups with t
plified our random effects structure. Next, we generated the model (45.90) = 0.851, p = 0.399. According to calculations provided by
of best fit in a top-down procedure, whereby we simplified the Lemhöfer and Broersma (2012), all of our speakers were at or
fixed effects structure in a stepwise fashion. After fitting each below the B2 level for Spanish according to CEFR standards
model, we performed model diagnostics to establish the goodness (Council of Europe, 2001), in line with our recruitment criteria.
of fit using the DHARMa package (Hartig, 2020). This involved
the plotting of the model residuals against the predicted values,
and closely investigating the distribution of the residuals and 3.4. Stroop task
the presence of influential data points to identify issues in We first computed descriptive statistics for accuracy and RTs for
terms of the model fit. Then, we compared models with different both groups. See Table 1 for descriptive mean accuracy, mean RTs
fixed effects structures to establish the model of best fit using the and Stroop effects for the Italian–Spanish group and the Dutch–
anova() function, which is based on the Akaike’s Information Spanish group. Descriptively speaking, results yielded overall longer
Criterion, AIC (Akaike, 1974), the Bayesian Information RTs for the Italian–Spanish group compared to the Dutch–Spanish
Criterion, BIC (Neath & Cavanaugh, 2012) and the log-likelihood group. Moreover, the Stroop effect was descriptively larger for the
ratio. To test for the significance of the terms in the fixed effects typologically similar languages compared to the typologically dis-
structure, absolute t-values greater than 1.96 were interpreted as similar languages. We first discuss the individual analysis for the
statistically significant at α = 0.05 (Alday, Schlesewsky, & Italian–Spanish group and the Dutch–Spanish group, respectively.
Bornkessel-Schlesewsky, 2017; Matuschek et al., 2017). Finally, Then, we discuss the group comparison for the Stroop effect size.
the models of best fit for RTs were re-fitted using the REML cri-
terion (Bates et al., 2014; Verbyla, 2019). All best-fitting models
Italian–Spanish group: Accuracy
and model parameters are reported in Appendix C, D and E.
For the Italian–Spanish group, the model of best fit included con-
To model accuracy, we used the glmer() function with a bino-
dition as fixed effect, as well as subject and item as random effects.
mial distribution. This particular function from the lme4 package
The by-condition random slopes for subjects led to singular fit
uses maximum likelihood estimation via the Laplace approxima-
and were therefore dropped from the model fitting procedure.
tion (Bates et al., 2020). In contrast, we used the lmer() function
The fixed effects LexTALE-Esp score and order of acquisition of
with a normal distribution to model RTs for correct trials. For the
Spanish did not significantly improve the model fit. Participants
individual group analysis, our fixed effect of interest was condition
were significantly more accurate for congruent trials compared to
(congruent vs. incongruent), whereas subject and item were
incongruent trials with β = 0.560, SE = 0.171, z = -3.38, p = .001
included as random effects. For the group comparison analysis,
(see Appendix C for the full model parameters). See Figure 2 for
we used the lmer() function to model the interaction effect of con-
mean accuracy for the Italian–Spanish group.
dition (congruent vs. incongruent) and typological similarity
(typologically similar vs. typologically dissimilar) as well as their
main effects on RTs. Subject and item were again included as ran- Italian–Spanish group: Response times
dom effects. To control for potential co-variates, we included For the Italian–Spanish group, the model of best fit yielded an
Lextale-Esp score and order of acquisition of Spanish as fixed effect of condition, a random effect for subject and item and a
effects in all analyses. by-subject random slope for condition (Figure 2). Neither
Lextale-Esp score nor order of acquisition of Spanish significantly
modulated the outcome variable or improved the model fit. These
two fixed effects were therefore excluded from the model fitting
3.3. Lextale-Esp
procedure. Participants were statistically faster in responding in
Post-experiment, we calculated Lextale-Esp vocabulary size scores. the congruent condition compared to the incongruent condition
We subtracted the percentage of incorrectly identified pseudowords with β = 29.46, SE = 5.79, t = 5.09, p < .001 (see Appendix C).

Table 1. Mean accuracy and RTs for the Stroop task for the Italian–Spanish group (n = 32) and the Dutch–Spanish group (n = 25).

Italian Dutch

Condition Accuracy (%) RTs (ms) Accuracy (%) RTs (ms)

Mean SD Mean SD Mean SD Mean SD


Congruent 94.7 22.4 579 113 95.7 20.4 560 123
Incongruent 91.2 28.3 607 112 90.7 29.1 576 112
Stroop effect 28 16

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


170 Sarah Von Grebmer Zu Wolfsthurn et al.

Figure 2. Mean accuracy (A) for each participant (n = 32)


and response times (B) for each condition for the
Italian–Spanish group.

Dutch–Spanish group: Accuracy the model fit and were subsequently dropped from the model
For the Dutch–Spanish group, the model of best fit included a selection procedure. Participants were significantly faster in
fixed effect of condition and Lextale-Esp score, as well as responding in the congruent condition compared to the incongru-
by-subject random slopes for condition and subject as random ent condition, with β = 15.31, SE = 5.56, t = 2.75, p = .006 (see
effect. Item led to singular fit and was excluded from the model Appendix D).
fitting procedure. Further, the fixed effect of order of acquisition In sum, data from both the Italian–Spanish group and the
of Spanish did not significantly improve the model fit. Dutch–Spanish groups suggest that participants were significantly
Participants were significantly more accurate in the congruent more accurate and faster in the congruent condition compared to
compared to the incongruent condition with β = 0.466, SE = the incongruent condition. Therefore, both groups displayed the
0.243, z = −3.13, p = .002. Despite being included in the final Stroop effect.
model, the fixed effect of Lextale-Esp score was not statistically sig-
nificant with β = 0.986, SE = 0.009, z = −1.67, p = .096 (see
Appendix D for the full model parameters). See Figure 3 for Stroop effect: group comparison
mean accuracy for the Dutch–Spanish group. Finally, we compared the Stroop effect (RT incongruent trials
minus RTs congruent trials) between the Italian–Spanish and
Dutch–Spanish group to explore the possible impact of typo-
Dutch–Spanish group: Response times logical similarity. Here, we explored the interaction effect between
For the Dutch–Spanish group, we found that the model of best fit condition and typological similarity on the size of the Stroop
included condition as fixed effect, subject as random effect and effect. Descriptively speaking, the Stroop effect was larger for
by-subject random slopes for condition (Figure 3). The random the Italian–Spanish group compared to the Dutch–Spanish
effect for item was not supported by our data and was therefore group. However, the model of best fit yielded a main effect of con-
excluded from the random effects structure. Neither Lextale-Esp dition with participants being faster for congruent trials compared
score nor order of acquisition of Spanish significantly improved to incongruent trials with β = 23.20, SE = 4.40, t = 5.27, p < .001.

Figure 3. Mean accuracy (A) for each participant (n = 25)


and response times (B) for each condition for the
Dutch–Spanish group.

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


Bilingualism: Language and Cognition 171

Figure 4. Mean response times for the Italian–Spanish group (left) and the Dutch–Spanish group (right) for each condition for Spanish Stroop targets (n = 57).

The model also included a main effect of typological similarity, information (location of the target) and instead focus on the tar-
with participants from the typologically similar group (Italian– get word itself to provide a correct response. Further, as discussed
Spanish) being significantly slower compared to the typologically in the introduction, participants had to employ a proactive con-
dissimilar group (Dutch–Spanish) with β = 30.70, SE = 11.72, t = trol strategy (Braver, 2012; Gonthier et al., 2016) and monitor
2.62, p = .009 (see Appendix E). There was no evidence for an the goal-relevant information during the task, as described in
interaction effect between condition and typological similarity. the DMC model (Braver, 2012; Gonthier et al., 2016).
See Appendix E for full model specification details, as well as a Therefore, the presence of a Stroop effect in both groups reflects
comparison between the model that included the interaction not only a measure for inhibitory control performance, but also a
term and the best-fitting model that did not include the inter- monitoring strategy to solve this task.
action term. Further, the model of best fit also included a With respect to the first research question, the group compari-
by-subject random slope for condition as well as item as random son analysis showed that the typologically dissimilar (Dutch–
effect. The fixed effects Lextale-Esp score and order of acquisition Spanish) group was comparatively faster than the typologically
of Spanish did not significantly contribute to improving the model similar (Italian–Spanish) group in this task. This finding contrasts
fit and were therefore not included in the final model. See Figure 4 with our predictions. The original prediction on the basis of
for the comparison of the Stroop effect across the Italian–Spanish the CRM (Stocco et al., 2014) was a processing advantage for
and Dutch–Spanish group. the typologically similar Italian–Spanish group compared to the
Dutch–Spanish group due to continuous training of executive
functions and inhibitory control skills over time. In contrast,
4. Discussion
our findings suggest that typologically dissimilar Dutch–Spanish
In this study, we explored the effect of typological similarity on group had a processing advantage in terms of RTs over the
inhibitory control performance in a group of Italian–Spanish Italian–Spanish group. In the literature, similar findings were
speakers and a group of Dutch–Spanish speakers via a spatial reported by Bialystok et al. (2005), who investigated the role of
Stroop task. The goal of this study was twofold: first, we examined typological similarity on the performance during a Simon task
whether or not the typologically similar (Italian–Spanish) group in highly proficient Cantonese–English speakers (typologically
showed a general processing advantage over the typologically dis- dissimilar group) and highly proficient French–English speakers
similar (Dutch–Spanish) group in terms of RTs. Secondly, we (typologically more similar group). Results showed a processing
studied whether typological similarity yielded a difference advantage for Cantonese–English speakers compared to the
between the two groups in terms of Stroop effect sizes (difference French–English speakers in the form of faster RTs on the
in RTs between incongruent and congruent trials). Here, a smaller Simon task for Cantonese–English speakers (see also Linck
Stroop effect would be indicative of better inhibitory control et al., 2005). Our results are comparable to Bialystok et al.
performance. On the basis of the CRM (Stocco et al., 2014), we (2005), and suggest that in this particular task, typological dis-
expected shorter RTs and a smaller Stroop effect for the similarity was advantageous over typological similarity.
Italian–Spanish group compared to the Dutch–Spanish group. Moreover, these results suggest a qualitative difference between
Stroop data from both the Italian–Spanish and the Dutch– the Italian–Spanish and the Dutch–Spanish group – namely, a
Spanish group showed that participants were sensitive to the more efficient inhibitory control strategy for the speakers of the
inherent task conflict. More specifically, results demonstrated less typologically similar languages. Within the framework of
higher accuracy and shorter RTs for congruent compared to the DMC model (Braver, 2012) and the application of proactive
incongruent trials. This yields the typical Stroop effect, which is control strategies during this task (Braver, 2012; Gonthier et al.,
a measure of inhibitory control performance in this task. To suc- 2016), this implies that Dutch–Spanish speakers were more effect-
ceed at this task, participants had to ignore the irrelevant ive at employing a proactive control strategy, as reflected in overall

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


172 Sarah Von Grebmer Zu Wolfsthurn et al.

shorter RTs. In other words, speakers of typologically more dis- On the other hand, the between-language Stroop effect quantifies
similar languages were better at monitoring and actively main- the differences in RTs between the congruent and incongruent
taining goal-related information compared to speakers of condition when the stimulus and response languages are different
typologically similar languages. This has critical implications for (Brauer, 1998; Marian et al., 2013; Van Heuven et al., 2011).
the conceptualisation of the underlying cognitive mechanisms Critically, Brauer (1998) included low- and high-proficient speak-
for typologically similar vs. dissimilar language combinations. ers to also explore the effect of proficiency on inhibitory control
With respect to our second research question, there was a performance. All three groups showed a within-language and a
descriptive trend of a smaller Stroop effect for the Dutch– between-language Stroop effect. On the one hand, low proficiency
Spanish group compared to the Italian–Spanish group. However, in the non-native language was linked to larger differences
the overall processing advantage of the Dutch–Spanish group between the within-language and the between-language Stroop
over the Italian–Spanish group was not reflected in the size of effect across the native and non-native language, irrespective of
the Stroop effect. More concretely, we did not find a statistical dif- typological similarity. On the other hand, highly proficient speak-
ference between the Stroop effect size for the Italian–Spanish group ers in the typologically dissimilar group were linked to larger
compared to the Dutch–Spanish group. This finding was somewhat within-language compared to between-language Stroop effects
surprising and contrasts with our original predictions. Our result in both the native and non-native language. Importantly, highly
suggested, first, that the Stroop effect was unaffected by typological proficient speakers in the typologically similar group showed no
similarity, and second, that speakers of both groups demonstrated a difference between the within-language and the between-language
highly comparable inhibitory control performance in this task. Stroop effect. Therefore, these results suggest that when the differ-
Importantly, the CRM framework proposed by Stocco et al. ence in proficiency levels is considerable (i.e., low proficiency in
(2014) does not fully account for these specific findings. Instead, the non-native language), the effect of typological similarity on
our findings strongly suggest a limited modulatory role of typo- language control performance may be limited, potentially because
logical similarity on inhibitory control performance in this study. the amount of “training” of the inhibitory skills has not yet been
One arising question here is the following: why were the Dutch– sufficient to elicit any typological similarity effects.
Spanish speakers faster, but not better, compared to the Italian– Given the strong link between language control and domain-
Spanish speakers at performing the Stroop task? general inhibitory control (Bialystok et al., 2012; Declerck et al.,
One interpretation of our findings could be that factors other 2021; Festman et al., 2010), this argument could be applied to
than typological similarity influence inhibitory control perform- our study: our Italian–Spanish and Dutch–Spanish speakers
ance in this task. These other potentially modulating factors were late language learners of Spanish who had a B1/B2 profi-
exert their influence such that one group had an inhibitory con- ciency level in Spanish. We therefore postulate that the difference
trol advantage in terms of processing speed, but not in terms of in proficiency between the native language (i.e., Italian or Dutch)
overall performance. A well-established modulatory factor in and the non-native language Spanish was too substantial to elicit a
language control, but less in inhibitory control, is language profi- typological similarity effect on inhibitory control performance,
ciency, as postulated in the IC model (Green, 1998). Previous even at intermediate B1/B2 proficiency levels. However, we antici-
studies have shown that multilingual children with a low non- pate that with increasing non-native proficiency levels, a typo-
native proficiency display unilateral cross-language interactions logical similarity effect on inhibitory control may be more
from the L1 into the L2 compared to multilingual children with pronounced. In view of this, it may not be surprising that inhibi-
high non-native proficiency (Brenders, Van Hell, & Dijkstra, tory control performance (i.e., the size of the Stroop effect) was
2011; Poarch & Van Hell, 2012a). As outlined in Poarch and statistically equal given that our groups had highly comparable
Van Hell (2012b), this could indicate that less language control proficiency levels in their non-native language Spanish. Thus,
effort is needed to manage the native and the non-native while our findings are not fully compatible with the CRM frame-
languages. In turn, this implies less training of more general work proposed in the introduction (Stocco et al., 2014), they sug-
executive control functions such as inhibitory control if the differ- gest that at intermediate non-native proficiency levels, the
ence in proficiency levels between the native and non-native lan- modulating role of typological similarity is not yet traceable at
guage is considerable. the behavioural level.
More specific to our intermediate late learners of Spanish, one A second interpretation of our findings could be that man-
could argue that our participants have not yet sufficiently trained aging cross-language interference between two typologically simi-
their inhibitory control skills given their intermediate level of lar languages does not directly transfer to strengthening the
non-native proficiency, in turn accounting for a limited effect of networks underlying inhibitory control. While we know that
typological similarity in this study. Therefore, one possibility is speaking multiple languages has a direct impact on language con-
that there is an interaction effect between typological similarity trol (Coderre & Van Heuven, 2014; Coderre et al., 2013; Green,
and non-native proficiency, and only a particular degree of typo- 1998; Green & Abutalebi, 2013; Mosca, 2019), this may not gen-
logical similarity paired with a specific proficiency level leads to eralise to broader executive functions such as inhibitory control.
training of the inhibitory control skills. This tentative hypothesis Contrary to the predictions by the CRM (Stocco et al., 2014), it
is partially in line with language control research by Brauer may be the case that speaking typologically similar languages
(1998). This study explored the effect of typological similarity does not result in a quantitative difference in the amount of train-
on language control via the within-language Stroop effect and ing of executive functions over time compared to typologically
the between-language Stroop effect in speakers of typologically dissimilar languages. Therefore, the link between speaking typolo-
similar languages (German–English) and typologically dissimilar gically similar languages, language control and inhibitory control
languages (English–Greek and English–Chinese) in the classical needs to be more closely inspected in future studies, specifically,
Stroop paradigm. The within-language Stroop effect refers to the association between language control and inhibitory control.
the differences in RTs between the congruent and incongruent Considering our compelling findings, the current study takes an
condition when the stimulus and response languages are identical. important step towards understanding the relative contribution of

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


Bilingualism: Language and Cognition 173

typological similarity to inhibitory control performance. Taken Olusayo Oloruntuyi for their support during the data collection at Leiden
together, our results suggest that typological similarity only plays a University and their feedback on the first draft. We also thank Michal
limited role in modulating inhibitory control performance, already Korenar for his helpful comments. Finally, we would like to thank our
anonymous reviewers and all our participants for their time.
at the stage when there is a moderate difference in proficiency levels
between the native and the non-native language. However, typo- CRediT Author Contribution Statement. Sarah Von Grebmer Zu
logical similarity may start to play a role only when non-native profi- Wolfsthurn: Conceptualisation, Methodology, Software, Investigation, Formal
ciency becomes more native-like. Second, our findings further suggest Analysis, Data Curation, Writing-Original Draft, Writing-Review & Editing,
a more complex link between managing multiple languages and more Visualisation. Anna Gupta: Conceptualisation, Investigation, Formal Analysis,
general inhibitory control skills. This could imply that multilingual- Data Curation, Writing-Original Draft, Writing-Review & Editing. Leticia Pablos:
ism primarily influences language control, but that it has only limited Conceptualisation, Methodology, Writing-Review & Editing, Supervision. Niels
effect on domain-general inhibitory control mechanisms. Therefore, O. Schiller: Conceptualisation, Writing-Review & Editing, Supervision, Funding
Acquisition.
our results have important implications for the conceptualisation of
the underlying processes of inhibitory control and add novel evidence Declaration of Competing Interests. No known competing interests.
to the debate around the role of typological similarity in inhibitory
control performance. Funding Statement. This project has received funding from the European
Union’s Horizon2020 research and innovation programme under the Marie
Skłodowska Curie grant agreement No 765556 – The Multilingual Mind.
5. Conclusions
Citation Diversity Statement. Within academia, research is witnessing a
In this study, we used a spatial Stroop task to examine whether systematic underrepresentation of female researchers and members of minor-
and how inhibitory control performance measured via the ities in published articles (Dworkin, Linn, Teich, Zurn, Shinohara, & Bassett,
Stroop effect was modulated by typological similarity. We found 2020; Rust & Mehrpour, 2020; Torres, Blevins, Bassett, & Eliassi-Rad, 2020;
that the typologically dissimilar (Dutch–Spanish) group was faster Zurn, Bassett, & Rust, 2020). With this Citation Diversity Statement, we aim
to raise awareness about this issue. We classified the first and last author
in performing the task compared to the typologically similar
based on their preferred gender for each reference in our reference list (wher-
(Italian–Spanish) group. This implied that the Dutch–Spanish
ever this information was available). Our reference list contained 25% woman/
group was better at monitoring goal-related information through- woman authors, 42% man/man, 13% woman/man and finally, 16% man/
out the task compared to the Italian–Spanish group. Critically, woman authors. For lack of direct comparison in the psycholinguistic field,
this did not impact the overall Stroop task performance. we compared this to 6.7% for woman/woman, 58.4% for man/man, 25.5%
Instead, the size of the Stroop effect, and in turn inhibitory con- woman/man, and lastly, 9.4% for man/woman authored references for the
trol performance, were similar across both groups, irrespective of field of neuroscience (Dworkin et al., 2020). Note that the limitations of this
typological similarity. Therefore, our results suggest that typo- classification are twofold: first, we need to develop adequate comparison
logical similarity plays a limited role in modulating inhibitory metrics for every research field; and second, we need to improve this rudimen-
control performance, particularly in intermediate proficient mul- tary binary gender classification system. However, we are confident that future
work will address both issues.
tilinguals with considerable differences in proficiency between
their L1 and non-native language(s). Data availability statement. The data that support the findings of this
study are openly available in Open Science Framework at https://osf.io/
e5ba9/?view_only=e8c7dbdf07984cbebb0fce50a132a605 [View-Only link]
5.1. Future directions
Our findings open new avenues to expand on current theoretical References
frameworks describing the impact of typological similarity on
inhibitory control. An emerging line of research could focus on Abutalebi, J (2008) Neural aspects of second language representation and lan-
guage control. Acta Psychologica 128, 466–478. https://doi.org/10.1016/j.
quantifying the degree of interference between typologically simi-
actpsy.2008.03.014
lar vs. dissimilar languages and the consequences for language Abutalebi, J, Della Rosa, PA, Green, DW, Hernandez, M, Scifo, P, Keim, R,
control and/or inhibitory control. For this, future studies should Cappa, SF and Costa, A (2012) Bilingualism tunes the anterior cingulate
first: investigate language pairs with varying degrees of typological cortex for conflict monitoring. Cerebral Cortex 22, 2076–2086. https://doi.
similarity; second, include separate measures for both language org/10.1093/cercor/bhr287
control and inhibitory control performance; and, finally, recruit Abutalebi, J and Green, DW (2007) Bilingual language production: The neuro-
speakers of different proficiency levels to tease apart the poten- cognition of language representation and control. Journal of Neurolinguistics 20,
tially critical effects of proficiency in modulating inhibitory con- 242–275. https://doi.org/10.1016/j.jneuroling.2006.10.003
trol performance. Recent years have also seen an increase in Akaike, H (1974) A new look at the statistical model identification. IEEE
Transactions on Automatic Control 19, 716–723.
research on the neurocognition of inhibitory control which com-
Alday, PM, Schlesewsky, M and Bornkessel-Schlesewsky, I (2017)
bines behavioural measures with electrophysiological and neuroi-
Electrophysiology reveals the neural dynamics of naturalistic auditory lan-
maging methods (Abutalebi et al., 2012; Christoffels et al., 2007; guage processing: Event-related potentials reflect continuous model
Constantinidis & Luna, 2019; Grundy, Anderson, & Bialystok, updates. ENeuro 4, 1–19. https://doi.org/10.1523/ENEURO.0311-16.2017
2017). Future studies in this area of research should also incorp- Barr, DJ, Levy, R, Scheepers, C and Tily, HJ (2013) Random effects structure
orate both offline and online measures such as electroencephalog- for confirmatory hypothesis testing: Keep it maximal. Journal of Memory
raphy or fMRI measures to model the cognitive and neural and Language 68, 255–278. https://doi.org/10.1016/j.jml.2012.11.001
mechanisms underlying inhibitory control performance in multi- Bates, DM, Mächler, M, Bolker, B and Walker, S (2014) Fitting linear
lingual language processing. mixed-effects models using lme4. ArXiv:1406.5823 [Stat]. http://arxiv.org/
abs/1406.5823
Acknowledgments. We thank Núria Sebastián-Gallés and the CBC labora- Bates, DM, Mächler, M, Bolker, B, Walker, S, Christensen, RHB,
tory team from Pompeu Fabra University for providing the facilities for the Singmann, H, Dai, B, Scheipl, F, Grothendieck, G, Green, P, Fox, J,
data collection. A special thanks goes to Philipp Flieger and Ifeoluwa Bauer, A and Krivitsky, PN (2020) Package ‘lme4.

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


174 Sarah Von Grebmer Zu Wolfsthurn et al.

Bialystok, E (2010) Global–local and trail-making tasks by monolingual and Constantinidis, C and Luna, B (2019) Neural substrates of inhibitory control
bilingual children: Beyond inhibition. Developmental Psychology 46, 93–105. maturation in adolescence. Trends in Neurosciences 42, 604–616. https://doi.
https://doi.org/10.1037/a0015466 org/10.1016/j.tins.2019.07.004
Bialystok, E, Craik, FIM, Grady, C, Chau, W, Ishii, R, Gunji, A and Pantev, Costa, A, Caramazza, A and Sebastián-Gallés, N (2000) The cognate facilita-
C (2005) Effect of bilingualism on cognitive control in the Simon task: tion effect: Implications for models of lexical access. Journal of Experimental
Evidence from MEG. NeuroImage 24, 40–49. https://doi.org/10.1016/j.neu- Psychology: Learning, Memory, and Cognition 26, 1283–1296. https://doi.
roimage.2004.09.044 org/10.1037/0278-7393.26.5.1283
Bialystok, E, Craik, FIM, Klein, R and Viswanathan, M (2004a) Bilingualism, Costa, A, Hernández, M and Sebastián-Gallés, N (2008) Bilingualism aids
aging, and cognitive control: Evidence from the Simon Task. Psychology and conflict resolution: Evidence from the ANT task. Cognition 106, 59–86.
Aging 19, 290–303. https://doi.org/10.1037/0882-7974.19.2.290 https://doi.org/10.1016/j.cognition.2006.12.013
Bialystok, E, Craik, FIM, Klein, R and Viswanathan, M (2004b) Costa, A and Santesteban, M (2004) Lexical access in bilingual speech pro-
Bilingualism, aging, and cognitive control: Evidence From the Simon duction: Evidence from language switching in highly proficient bilinguals
Task. Psychology and Aging 19, 290–303. https://doi.org/10.1037/0882- and L2 learners. Journal of Memory and Language 50, 491–511. https://
7974.19.2.290 doi.org/10.1016/j.jml.2004.02.002
Bialystok, E, Craik, FIM and Luk, G (2012) Bilingualism: Consequences for Council of Europe. (2001) Common European Framework of Reference for
mind and brain. Trends in Cognitive Sciences 16, 240–250. https://doi.org/ Languages: Learning, Teaching, Assessment. Cambridge University Press.
10.1016/j.tics.2012.03.001 De Bot, K (2004) The multilingual lexicon: Modelling selection and control.
Bialystok, E and Martin, MM (2004) Attention and inhibition in bilingual International Journal of Multilingualism 1, 17–32. https://doi.org/10.1080/
children: Evidence from the dimensional change card sort task. 14790710408668176
Developmental Science 7, 325–339. https://doi.org/10.1111/j.1467-7687. Declerck, M, Koch, I, Duñabeitia, JA, Grainger, J and Stephan, DN (2019)
2004.00351.x What absent switch costs and mixing costs during bilingual language com-
Blumenfeld, HK and Marian, V (2013) Parallel language activation and cog- prehension can tell us about language control. Journal of Experimental
nitive control during spoken word recognition in bilinguals. Journal of Psychology: Human Perception and Performance 45, 771–789. https://doi.
Cognitive Psychology 25, 547–567. https://doi.org/10.1080/20445911.2013. org/10.1037/xhp0000627
812093 Declerck, M, Meade, G, Midgley, KJ, Holcomb, PJ, Roelofs, A and
Branzi, FM, Della Rosa, PA, Canini, M, Costa, A and Abutalebi, J (2016) Emmorey, K (2021) On the connection between language control and
Language control in bilinguals: Monitoring and response selection. executive control—An ERP study. Neurobiology of Language 2, 628–646.
Cerebral Cortex 26, 2367–2380. https://doi.org/10.1093/cercor/bhv052 https://doi.org/10.1162/nol_a_00032
Brauer, M (1998) Stroop interference in bilinguals: The role of similarity De Groot, AMB (2017) Bi-and Multilingualism. In YY Kim and
between the two languages. In AF Healy and LE Bourne, Jr. (Eds.), K McKay-Semmler (Eds.), The International Encyclopedia of Intercultural
Foreign language learning: Psycholinguistic studies on training and retention Communication (1st ed., Vol. 2). John Wiley & Sons, pp. 1–10 https://doi.
(1st ed.). Lawrence Erlbaum Associates Publishers, pp. 317–337. https://psy- org/10.1002/9781118783665
cnet.apa.org/record/1998-06629-013 Diamond, A (2013) Executive functions. Annual Review of Psychology 64,
Braver, TS (2012) The variable nature of cognitive control: A dual mechan- 135–168. https://doi.org/10.1146/annurev-psych-113011-143750
isms framework. Trends in Cognitive Sciences 16, 106–113. https://doi.org/ Dijkstra, T, Van Heuven, WJB and Grainger, J (1998) Simulating cross-
10.1016/j.tics.2011.12.010 language competition with the bilingual interactive activation model.
Brenders, PEA, Van Hell, JG and Dijkstra, T (2011) Word recognition in Psychologica Belgica 38, 177–196.
child second language learners: Evidence from cognates and false friends. Dworkin, JD, Linn, KA, Teich, EG, Zurn, P, Shinohara, RT and Bassett, DS
Journal of Experimental Child Psychology 109, 383–396. https://doi.org/10. (2020) The extent and drivers of gender imbalance in neuroscience refer-
1016/j.jecp.2011.03.012 ence lists. BioRxiv, 2020.01.03.894378. https://doi.org/10.1101/2020.01.03.
Calabria, M, Hernandez, M, Branzi, F and Costa, A (2012) Qualitative dif- 894378
ferences between bilingual language control and executive control: Evidence Festman, J, Rodriguez-Fornells, A and Münte, TF (2010) Individual differ-
from task-switching. Frontiers in Psychology 2. https://doi.org/10.3389/ ences in control of language interference in late bilinguals are mainly related
fpsyg.2011.00399 to general executive abilities. Behavioral and Brain Functions 6, 5. https://
Cenoz, J (2001) The effect of linguistic distance, L2 status and age on cross- doi.org/10.1186/1744-9081-6-5
linguistic influence in third language acquisition. In J Cenoz, B Hufeisen Foote, R (2009) Transfer in L3 acquisition: The role of typology. In Y.-K. I. Leung
and U Jessner (Eds.), Cross-linguistic influence in third language acquisition: (Ed.), Third language acquisition and universal grammer. Multilingual Matters,
Psycholinguistic perspectives. Multilingual Matters, pp. 8–20. https://doi.org/ pp. 89–114. https://doi.org/10.21832/9781847691323-008
10.21832/9781853595509-002 Gonthier, C, Braver, TS and Bugg, JM (2016) Dissociating proactive and
Cenoz, J (2013) Defining multilingualism. Annual Review of Applied reactive control in the Stroop task. Memory & Cognition 44, 778–788.
Linguistics 33, 3–18. https://doi.org/10.1017/S026719051300007X https://doi.org/10.3758/s13421-016-0591-1
Chen, J, Zhao, Y, Zhaxi, C and Liu, X (2020) How does parallel language acti- Green, DW (1998) Mental control of the bilingual lexico-semantic system.
vation affect switch costs during trilingual language comprehension? Bilingualism: Language and Cognition 1, 67–81. https://doi.org/10.1017/
Journal of Cognitive Psychology 32, 526–542. https://doi.org/10.1080/ S1366728998000133
20445911.2020.1788036 Green, DW and Abutalebi, J (2013) Language control in bilinguals: The adap-
Christoffels, IK, Firk, C and Schiller, NO (2007) Bilingual language control: tive control hypothesis. Journal of Cognitive Psychology 25, 515–530. https://
An event-related brain potential study. Brain Research 1147, 192–208. doi.org/10.1080/20445911.2013.796377
https://doi.org/10.1016/j.brainres.2007.01.137 Grundy, J, Anderson, J and Bialystok, E (2017) Neural correlates of cognitive
Coderre, EL and Van Heuven, WJB (2014) The effect of script similarity on processing in monolinguals and bilinguals. Annals of the New York
executive control in bilinguals. Frontiers in Psychology 5, 1–16. https://doi. Academy of Sciences 1396. https://doi.org/10.1111/nyas.13333
org/10.3389/fpsyg.2014.01070 Guo, T and Peng, D (2006) Event-related potential evidence for parallel acti-
Coderre, EL, Van Heuven, WJB and Conklin, K (2013) The timing and mag- vation of two languages in bilingual speech production: NeuroReport 17,
nitude of Stroop interference and facilitation in monolinguals and bilin- 1757–1760. https://doi.org/10.1097/01.wnr.0000246327.89308.a5
guals. Bilingualism: Language and Cognition 16, 420–441. https://doi.org/ Hartig, F (2020) DHARMa: Residual diagnostics for hierarchical (multi-level/
10.1017/S1366728912000405 mixed) regression models [English]. https://cran.r-project.org/web/packages/
Colomé, À. (2001) Lexical activation in bilinguals’ speech production: DHARMa/vignettes/DHARMa.html
Language-specific or language-independent? Journal of Memory and Heidlmayr, K, Moutier, S, Hemforth, B, Courtin†, C, Tanzmeister, R and
Language 45, 721–736. https://doi.org/10.1006/jmla.2001.2793 Isel, F (2014) Successive bilingualism and executive functions: The effect

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


Bilingualism: Language and Cognition 175

of second language use on inhibitory control in a behavioural Stroop Colour Paolieri, D, Padilla, F, Koreneva, O, Morales, L and Macizo, P (2019)
Word task. Bilingualism: Language and Cognition 17, 630–645. https://doi. Gender congruency effects in Russian–Spanish and Italian–Spanish bilin-
org/10.1017/S1366728913000539 guals: The role of language proximity and concreteness of words.
Hilbert, S, Nakagawa, TT, Bindl, M and Bühner, M (2014) The spatial Bilingualism: Language and Cognition 22, 112–129. https://doi.org/10.
Stroop effect: A comparison of color-word and position-word interference. 1017/S1366728917000591
Psychonomic Bulletin & Review 21, 1509–1515. https://doi.org/10.3758/ Pardo, JV, Pardo, PJ, Janer, KW and Raichle, ME (1990) The anterior cin-
s13423-014-0631-4 gulate cortex mediates processing selection in the Stroop attentional conflict
Hodgson, TL, Parris, BA, Gregory, NJ and Jarvis, T (2009) The saccadic paradigm. Proceedings of the National Academy of Sciences 87, 256–259.
Stroop effect: Evidence for involuntary programming of eye movements https://doi.org/10.1073/pnas.87.1.256
by linguistic cues. Vision Research 49, 569–574. https://doi.org/10.1016/j. Pliatsikas, C (2020) Understanding structural plasticity in the bilingual brain:
visres.2009.01.001 The Dynamic Restructuring Model. Bilingualism: Language and Cognition
Hoshino, N and Thierry, G (2011) Language selection in bilingual word pro- 23, 459–471. https://doi.org/10.1017/S1366728919000130
duction: Electrophysiological evidence for cross-language competition. Poarch, GJ and Van Hell, JG (2012a) Cross-language activation in children’s
Brain Research 1371, 100–109. https://doi.org/10.1016/j.brainres.2010.11. speech production: Evidence from second language learners, bilinguals, and
053 trilinguals. Journal of Experimental Child Psychology 111, 419–438. https://
Izura, C, Cuetos, F and Brysbaert, M (2014) Lextale-Esp: A test to rapidly doi.org/10.1016/j.jecp.2011.09.008
and efficiently assess the Spanish vocabulary size. Psicologica 35, 49–67. Poarch, GJ and Van Hell, JG (2012b) Executive functions and inhibitory con-
Kroll, JF and Bialystok, E (2013) Understanding the consequences of bilin- trol in multilingual children: Evidence from second-language learners,
gualism for language processing and cognition. Journal of Cognitive bilinguals, and trilinguals. Journal of Experimental Child Psychology 113,
Psychology 25, 497–514. https://doi.org/10.1080/20445911.2013.799170 535–551. https://doi.org/10.1016/j.jecp.2012.06.013
Kroll, JF, Dussias, PE, Bice, K and Perrotti, L (2015) Bilingualism, Mind, and Putnam, MT, Carlson, M and Reitter, D (2018) Integrated, not isolated:
Brain. Annual Review of Linguistics 1, 377–394. https://doi.org/10.1146/ Defining typological proximity in an integrated multilingual architecture.
annurev-linguist-030514-124937 Frontiers in Psychology 8, 1–16. https://doi.org/10.3389/fpsyg.2017.02212
La Heij, W, Van der Heijden, AHC and Plooij, P (2001) A paradoxical R Core Team. (2020) R: A language and environment for statistical computing
exposure-duration effect in the Stroop task: Temporal segregation between (4.1.2) [Computer software]. R Foundation for Statistical Computing.
stimulus attributes facilitates selection. Journal of Experimental Psychology: https://www.R-project.org/
Human Perception and Performance 27, 622–632. https://doi.org/10.1037/ Roelofs, A (2021) Response competition better explains Stroop interference
0096-1523.27.3.622 than does response exclusion. Psychonomic Bulletin & Review 28, 487–
Lemhöfer, K and Broersma, M (2012) Introducing LexTALE: A quick and 493. https://doi.org/10.3758/s13423-020-01846-0
valid lexical test for advanced learners of English. Behavior Research Rosenman, R, Tennekoon, V and Hill, LG (2011) Measuring bias in self-
Methods 44, 325–343. https://doi.org/10.3758/s13428-011-0146-0 reported data. International Journal of Behavioural & Healthcare Research
Linck, JA, Hoshino, N and Kroll, JF (2005) Cross-language lexical processes 2, 320–332. https://doi.org/10.1504/IJBHR.2011.043414
and inhibitory control. The Mental Lexicon 3, 349–374. https://doi.org/10. Rust, NC and Mehrpour, V (2020) Understanding image memorability. Trends
1075/ml.3.3.06lin in Cognitive Sciences 24, 557–568. https://doi.org/10.1016/j.tics.2020.04.001
Lu, C and Proctor, RW (1995) The influence of irrelevant location informa- Schepens, J, Dijkstra, T and Grootjen, F (2012) Distributions of cognates in
tion on performance: A review of the Simon and spatial Stroop effects. Europe as based on Levenshtein distance. Bilingualism: Language and
Psychonomic Bulletin & Review 2, 174–207. https://doi.org/10.3758/ Cognition 15, 157–166. https://doi.org/10.1017/S1366728910000623
BF03210959 Schneider, W, Eschman, A and Zuccolotto, A (2002) E-prime User’s Guide.
Luo, C and Proctor, RW (2013) Asymmetry of congruency effects in spatial https://www.researchgate.net/publication/
Stroop tasks can be eliminated. Acta Psychologica 143, 7–13. https://doi.org/ 260296789_E-prime_User’s_Guide
10.1016/j.actpsy.2013.01.016 Schwieter, JW (2016) Cognitive control and consequences of multilingualism.
MacLeod, CM (1992) The Stroop task: The “gold standard” of attentional John Benjamins Publishing Company.
measures. Journal of Experimental Psychology: General 121, 12–14. Sebastián-Gallés, N and Kroll, JF (2003) Phonology in bilingual language
https://doi.org/10.1037/0096-3445.121.1.12 processing: Acquisition, perception, and production. In NO Schiller and
Marian, V, Blumenfeld, HK and Kaushanskaya, M (2007) The Language AS Meyer (Eds.), Phonetics and phonology in language comprehension and
Experience and Proficiency Questionnaire (LEAP-Q): Assessing language production: Differences and similarities. Walter de Gruyter, pp. 279–318.
profiles in bilinguals and multilinguals. Journal of Speech, Language, and Serratrice, L, Sorace, A, Filiaci, F and Baldo, M (2012) Pronominal objects in
Hearing Research 50, 940–967. https://doi.org/10.1044/1092-4388(2007/067) English–Italian and Spanish–Italian bilingual children. Applied
Marian, V, Blumenfeld, HK, Mizrahi, E, Kania, U and Cordes, A.-K. (2013) Psycholinguistics 33, 725–751. https://doi.org/10.1017/S0142716411000543
Multilingual Stroop performance: Effects of trilingualism and proficiency Shor, RE (1970) The processing of conceptual information on spatial direc-
on inhibitory control. International Journal of Multilingualism 10, tions from pictorial and linguistic symbols. Acta Psychologica 32, 346–
82–104. https://doi.org/10.1080/14790718.2012.708037 365. https://doi.org/10.1016/0001-6918(70)90109-5
Matuschek, H, Kliegl, R, Vasishth, S, Baayen, HR and Bates, DM (2017) Simon, JR and Small, AM (1969) Processing auditory information:
Balancing Type I error and power in linear mixed models. Journal of Memory Interference from an irrelevant cue. Journal of Applied Psychology 53,
and Language 94, 305–315. https://doi.org/10.1016/j.jml.2017.01.001 433–435. https://doi.org/10.1037/h0028034
Miyake, A, Friedman, NP, Emerson, MJ, Witzki, AH, Howerter, A and Stocco, A, Yamasaki, B, Natalenko, R and Prat, CS (2014) Bilingual brain
Wager, TD (2000) The unity and diversity of executive functions and training: A neurobiological framework of how bilingual experience
their contributions to complex “frontal lobe” tasks: A latent variable ana- improves executive function. International Journal of Bilingualism 18, 67–
lysis. Cognitive Psychology 41, 49–100. https://doi.org/10.1006/cogp.1999. 92. https://doi.org/10.1177/1367006912456617
0734 Stroop, JR (1935) Studies of interference in serial verbal reactions. Journal of
Mosca, M (2017) Multilinguals’ language control [Doctoral Dissertation]. Experimental Psychology 8, 643–662. https://doi.org/10.1037/h0054651
University of Potsdam. Torres, L, Blevins, AS, Bassett, DS and Eliassi-Rad, T (2020) The why, how,
Mosca, M (2019) Trilinguals’ language switching: A strategic and flexible and when of representations for complex systems. ArXiv:2006.02870 [Cs,
account. Quarterly Journal of Experimental Psychology 72, 693–716. q-Bio]. http://arxiv.org/abs/2006.02870
https://doi.org/10.1177/1747021818763537 Van der Slik, FWP (2010) Acquisition of Dutch as a second language: The
Neath, AA and Cavanaugh, JE (2012) The Bayesian Information Criterion: explanative power of cognate and genetic linguistic distance measures for
Background, derivation, and applications. WIREs Computational Statistics 11 West European first languages. Studies in Second Language Acquisition
4, 199–203. https://doi.org/10.1002/wics.199 32, 401–432. https://doi.org/10.1017/S0272263110000021

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


176 Sarah Von Grebmer Zu Wolfsthurn et al.

Van Heuven, WJB, Conklin, K, Coderre, EL, Guo, T and Dijkstra, T (2011) Linguistic Proximity Model. International Journal of Bilingualism 21,
The influence of cross-language similarity on within- and between-language 666–682. https://doi.org/10.1177/1367006916648859
Stroop effects in trilinguals. Frontiers in Psychology 2, 1–15. https://doi.org/ Wiseheart, M, Viswanathan, M and Bialystok, E (2016) Flexibility
10.3389/fpsyg.2011.00374 in task switching by monolinguals and bilinguals*. Bilingualism:
Verbyla, AP (2019) A note on model selection using information criteria for Language and Cognition 19, 141–146. https://doi.org/10.1017/
general linear models estimated using REML. Australian & New Zealand S1366728914000273
Journal of Statistics 61, 39–50. https://doi.org/10.1111/anzs.12254 Yamasaki, BL, Stocco, A and Prat, CS (2018) Relating individual differences
Weissberger, GH, Gollan, TH, Bondi, MW, Clark, LR and Wierenga, CE in bilingual language experiences to executive attention. Language,
(2015) Language and task switching in the bilingual brain: Bilinguals are Cognition and Neuroscience 33, 1128–1151. https://doi.org/10.1080/
staying, not switching, experts. Neuropsychologia 66, 193–203. https://doi. 23273798.2018.1448092
org/10.1016/j.neuropsychologia.2014.10.037 Zurn, P, Bassett, DS and Rust, NC (2020) The Citation Diversity Statement:
Westergaard, M, Mitrofanova, N, Mykhaylyk, R and Rodina, Y (2017) A practice of transparency, a way of life. Trends in Cognitive Sciences 24,
Crosslinguistic influence in the acquisition of a third language: The 669–672. https://doi.org/10.1016/j.tics.2020.06.009

Appendix
Appendix A. Linguistic profile of the Italian–Spanish group (N = 33) according to the LEAP-Q (Marian et al., 2007)

Native First foreign Second foreign Third foreign Fourth foreign


Language language language language language language Total

Italian n = 33 33
Spanish n=2 n = 18 n = 10 n=3 33
English n = 27 n=5 32
French n=4 n=8 n=3 15
German n=1 n=2 3
Catalan n=1 n=1 2
Portuguese n=3 3
Total 33 33 32 16 7

Appendix B. Linguistic profile of the Dutch–Spanish group (N = 25) according to the LEAP-Q (Marian et al., 2007)

Native First foreign Second foreign Third foreign Fourth foreign


Language language language language language language Total

Dutch n = 25 25
Spanish n=9 n=9 n=7 25
English n = 23 n=2 25
German n=7 n=7 14
French n=1 n=7 8
Portuguese n=2 n=2 4
Frisian n=1 1
Japanese n=1 n=1 2
Chinese n=1 1
Italian n=1 1
Total 25 25 25 19 12

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


Bilingualism: Language and Cognition 177

Appendix C. Models of best fit for accuracy and RTs, including odd ratios/estimates, confidence intervals, test statistics and
p-values for the Italian–Spanish group

Formula: accuracy ∼ condition + Formula: RTs ∼ condition +


(1|subject) + (1|item) (condition|subject) + (1|item)

95% CI 95% CI
Fixed Effect Odds Ratio [low – high] Statistic p Estimate [low – high] Statistic p

(Intercept) 24.86 16.56 – 37.32 15.50 <.001 579.21 561.94 – 596.48 65.77 <.001
Condition [incongruent] 0.580 0.400 – 0.784 −3.38 .001*** 29.46 18.10 – 40.82 5.09 <.001***

Random Effects
σ2 3.29 10783.88
τ00 Subject 0.72 2102.95
τ00 Item 0.01 8.82
τ11 Subject [incong] 307.02
ρ01 Subject −0.28
ICC 0.18 0.16
N Subject 32 32
N Item 4 4
Observations 3072 2856
2 2
Marginal R /Conditional R 0.020/0.198 0.017/0.173

Appendix D. Models of best fit for accuracy and RTs, including odd ratios/estimates, confidence intervals, test statistics
and p-values for the Dutch–Spanish group

Formula: accuracy ∼ condition + LexTALE-Esp + (condition|


subject) Formula: RTs ∼ condition + (condition|subject)

95% CI 95% CI
Fixed Effect Odds Ratio [low – high] Statistic p Estimate [low – high] Statistic p

(Intercept) 36.32 20.23 – 65.21 12.03 <.001 559.80 539.98 – 579.63 55.37 <.001
Condition [incongruent] 0.466 0.29 – 0.75 −3.13 .002** 15.31 4.40 – 26.21 2.75 .006**
LexTALE-Esp 0.986 0.97 – 1.00 −1.67 0.096

Random Effects
σ2 3.29 12180.85
τ00 Subject 0.32 2289.48
τ11 Subject [incong] 0.43 226.95
ρ01 Subject −0.31 −0.78
ICC 0.11 0.13
N Subject 25 25
Observations 2400 2236
Marginal R2 /Conditional R2 0.053/0.159 0.004/0.136

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press


178 Sarah Von Grebmer Zu Wolfsthurn et al.

Appendix E. Comparison between the model with the interaction effect of condition and typological similarity (left) and
the best-fitting model (right) with a main effect for condition and typological similarity

Formula: RTs ∼ typological similarity (TS) * condition + Formula: RTs ∼ typological similarity (TS) + condition +
(condition|subject) + (1|item) (condition|subject) + (1|item)

95% CI 95% CI
Fixed Effect Estimate [low – high] Statistic p Estimate [low – high] Statistic p

(Intercept) 559.79 540.73 – 578.85 57.57 <.001 553.48 534.99 – 571.96 58.68 <.001
TS [high] 19.44 −5.93 – 44.82 1.50 .133 30.70 7.72 – 53.68 2.62 .009**
Condition [incongruent] 15.33 4.36 – 26.29 2.74 .006** 23.20 14.56 – 31.83 5.27 <.001***
TS [high] *Condition 13.98 −0.40 – 28.36 1.91 .057
[incongruent]
Random Effects
σ2 11399.52 11397.84
τ00 Subject 2101.67 2210.02
τ00 Item 1.10 5.03
τ11 Subject [incong] 243.08 307.02
ρ01 Subject −0.50 −0.50
ICC 0.14 0.15
N Subject 57 57
N Item 4 4
Observations 5092 5092
2 2
Marginal R /Cond. R 0.023/ 0.027/
0.161 0.170

https://doi.org/10.1017/S1366728922000426 Published online by Cambridge University Press

You might also like