Professional Documents
Culture Documents
Co-Evolution of Language and of The Language Acquisition Device
Co-Evolution of Language and of The Language Acquisition Device
Ted Briscoe
ejb@cl.cam.ac.uk
Computer Laboratory
University of Cambridge
Pembroke Street
Cambridge CB2 3QG, UK
arXiv:cmp-lg/9705001v1 1 May 1997
Backward Application:
Y X\Y ⇒ X λ y [X(y)] (y) ⇒ X(y)
Forward Composition:
X/Y Y/Z ⇒ X/Z λ y [X(y)] λ z [Y(z)] ⇒ λ z [X(Y(z))]
Backward Composition:
Y\Z X\Y ⇒ X\Z λ z [Y(z)] λ y [X(y)] ⇒ λ z [X(Y(z))]
heritance hierarchy provides a partial ordering on postpositions in which specifiers and modifiers follow
parameters, which is exploited in the learning pro- heads. There are other languages in the SOV family
cedure. For example, the atomic categories, N, with less consistent left-branching syntax in which
NP and S are each represented by a parameter en- specifiers and/or modifiers precede phrasal heads,
coding the presence/absence or lack of specification some of which are attested. “German” is a more
(T/F/?) of the category in the (U)G. Since they will complex SOV language in which the parameter verb-
be unordered in the lattice their ordering in the se- second (v2) ensures that the surface order in main
quential coding is arbitrary. However, the ordering clauses is usually SVO.4
of the directional types gendir and subjdir (with There are 20 p-settings which determine the rule
values L/R) is significant as the latter is a more spe- schemata available, the atomic category set, and so
cific type. The distinctions between absolute, de- forth. In all, this CGUG defines just under 300
fault or unset specifications also form part of the grammars. Not all of the resulting languages are
encoding (A/D/?). Figure 3 shows several equiva- (stringset) distinct and some are proper subsets of
lent and equally correct sequential encodings of the other languages. “English” without the rule of per-
fragment of the English type system outlined above. mutation results in a stringset-identical language,
A set of grammars based on typological distinc- but the grammar assigns different derivations to
tions defined by basic constituent order (e.g. Green- some strings, though the associated logical forms are
berg, 1966; Hawkins, 1994) was constructed as a identical. “English” without composition results in
(partial) GCUG with independently varying binary- a subset language. Some combinations of p-settings
valued parameters. The eight basic language fami- result in ‘impossible’ grammars (or UGs). Others
lies are defined in terms of the unmarked order of yield equivalent grammars, for example, different
verb (V), subject (S) and objects (O) in clauses. combinations of default settings (for types and their
Languages within families further specify the order subtypes) can define an identical category set.
of modifiers and specifiers in phrases, the order of ad- The grammars defined generate (usually infinite)
positions and further phrasal-level ordering param- stringsets of lexical syntactic categories. These
eters. Figure 4 list the language-specific ordering strings are sentence types since each is equivalent
parameters used to define the full set of grammars to a finite set of grammatical sentences formed by
in (partial) order of generality, and gives examples selecting a lexical instance of each lexical category.
of settings based on familiar languages such as “En- Languages are represented as a finite subset of sen-
glish”, “German” and “Japanese”.3 “English” de- tence types generated by the associated grammar.
fines an SVO language, with prepositions in which These represent a sample of degree-1 learning trig-
specifiers, complementizers and some modifiers pre- gers for the language (e.g. Lightfoot, 1991). Subset
cede heads of phrases. There are other grammars in languages are represented by 3-9 sentence types and
the SVO family in which all modifers follow heads, ‘full’ languages by 12 sentence types. The construc-
there are postpositions, and so forth. Not all combi- tions exemplified by each sentence type and their
nations of parameter settings correspond to attested length are equivalent across all the languages defined
languages and one entire language family (OVS) is by the grammar set, but the sequences of lexical cat-
unattested. “Japanese” is an SOV language with egories can differ. For example, two SOV language
3
Throughout double quotes around language names
renditions of The man who Bill likes gave Fred a
are used as convenient mnemonics for familiar combina-
4
tions of parameters. Since not all aspects of these actual Representation of the v1/v2 parameter(s) in terms
languages are represented in the grammars, conclusions of a type constraint determining allowable functor cate-
about actual languages must be made with care. gories is discussed in more detail in Briscoe (1997b).
gen v1 n subj obj v2 mod spec relcl adpos compl
Engl R F R L R F R R R R R
Ger R F R L L T R R R R R
Jap L F L L L F L L L L ?
Bill likes who the-man a-present Fred gave 1. The Reduce Step: if the top 2 cells of the
NPs (S\NPs )\NPo Rc\(S\NPo ) NPs \Rc NPo2 stack are occupied,
NPo1 ((S\NPs )\NPo2 )\NPo1 then try
The-man Bill likes who a-present Fred gave a) Application, if match, then apply and goto
NPs /Rc NPs (S\NPs )\NPo Rc\(S\NPo ) NPo2 1), else b),
NPo1 ((S\NPs )\NPo2 )\NPo1 b) Combination, if match then apply and goto
1), else c),
c) Permutation, if match then apply and goto
1), else goto 2)
2.2 The Parser
2. The Shift Step: if the first cell of the Input
The parser is a deterministic, bounded-context
Buffer is occupied,
stack-based shift-reduce algorithm. The parser op-
then pop it and move it onto the Stack to-
erates on two data structures, an input buffer or
gether with its associated lexical syntactic cat-
queue, and a stack or push down store. The algo-
egory and goto 1),
rithm for the parser working with a GCG which in-
else goto 3)
cludes application, composition and permutation is
given in Figure 5. This algorithm finds the most left- 3. The Halt Step: if only the top cell of the Stack
branching derivation for a sentence type because Re- is occupied by a constituent of category S,
duce is ordered before Shift. The category sequences then return Success,
representing the sentence types in the data for the else return Fail
entire language set are designed to be unambiguous
relative to this ‘greedy’, deterministic algorithm, so The Match and Apply operation: if a binary
it will always assign the appropriate logical form to rule schema matches the categories of the top 2 cells
each sentence type. However, there are frequently al- of the Stack, then they are popped from the Stack
ternative less left-branching derivations of the same and the new category formed by applying the rule
logical form. schema is pushed onto the Stack.
The parser is augmented with an algorithm which
computes working memory load during an analy- The Permutation operation: each time step 1c)
sis (e.g. Baddeley, 1992). Limitations of working is visited during the Reduce step, permutation is ap-
memory are modelled in the parser by associating a plied to one of the categories in the top 2 cells of the
cost with each stack cell occupied during each step Stack until all possible permutations of the 2 cate-
of a derivation, and recency and depth of process- gories have been tried using the binary rules. The
ing effects are modelled by resetting this cost each number of possible permutation operations is finite
time a reduction occurs: the working memory load and bounded by the maximum number of arguments
(WML) algorithm is given in Figure 6. Figure 7 gives of any functor category in the grammar.
the right-branching derivation for Kim loves Sandy,
Figure 5: The Parsing Algorithm
found by the parser utilising a grammar without per-
mutation. The WML at each step is shown for this
derivation. The overall WML (16) is higher than for
the left-branching derivation (9).
The WML algorithm ranks sentence types, and
Stack Input Buffer Operation Step WML
Kim loves Sandy 0 0
Kim:NP:kim′ loves Sandy Shift 1 1
loves:(S\NP)/NP:λ y,x(love′ x, y) Sandy Shift 2 3
Kim:NP:kim′
Sandy:NP:sandy′ Shift 3 6
loves:(S\NP)/NP:λ y,x(love′ x, y)
Kim:NP:kim′
loves Sandy:S/NP:λ x(love′ x, sandy′ ) Reduce (A) 4 5
Kim:NP:kim′
Kim loves Sandy:S:(love′ kim′ , sandy′ ) Reduce (A) 5 1
After each parse step (Shift, Reduce, Halt (see its structure or context of use, is determinately un-
Fig 5): parsable with the correct interpretation (e.g. Light-
foot, 1991). In this model, the issue of ambigu-
1. Assign any new Stack entry in the top cell (in-
ity and triggers does not arise because all sentence
troduced by Shift or Reduce) a WML value of
types are treated as triggers represented by p-setting
0
schemata. The TLA is memoryless in the sense that
2. Increment every Stack cell’s WML value by 1 a history of parameter (re)settings is not maintained,
in principle, allowing the learner to revisit previous
3. Push the sum of the WML values of each Stack hypotheses. This is what allows Niyogi and Berwick
cell onto the WML-record (1995) to formalize parameter setting as a Markov
process. However, as Brent (1996) argues, the psy-
When the parser halts, return the sum of the WML-
chological plausibility of this algorithm is doubt-
record gives the total WML for a derivation
ful – there is no evidence that children (randomly)
Figure 6: The WML Algorithm move between neighbouring grammars along paths
that revisit previous hypotheses. Therefore, each
parameter can only be reset once during the learn-
thus indirectly languages, by parsing each sentence ing process. Each step for a learner can be defined
type from the exemplifying data with the associ- in terms of three functions: p-setting, grammar
ated grammar and then taking the mean of the and parser, as:
WML obtained for these sentence types. “En-
glish” with Permutation has a lower mean WML parseri (grammari (p-settingi (Sentencej )))
than “English” without Permutation, though they
A p-setting defines a grammar which in turn defines
are stringset-identical, whilst a hypothetical mix-
a parser (where the subscripts indicate the output of
ture of “Japanese” SOV clausal order with “En-
each function given the previous trigger). A param-
glish” phrasal syntax has a mean WML which is 25%
eter is updated on parse failure and, if this results
worse than that for “English”. The WML algorithm
in a parse, the new setting is retained. The algo-
is in accord with existing (psycholinguistically-
rithm is summarized in Figure 8. Working mem-
motivated) theories of parsing complexity (e.g. Gib-
ory grows through childhood (e.g. Baddeley, 1992),
son, 1991; Hawkins, 1994; Rambow and Joshi, 1994).
and this may assist learning by ensuring that trigger
2.3 The Parameter Setting Algorithm sentences gradually increase in complexity through
the acquisition period (e.g. Elman, 1993) by forcing
The parameter setting algorithm is an extension of
the learner to ignore more complex potential triggers
Gibson and Wexler’s (1994) Trigger Learning Al-
that occur early in the learning process. The WML
gorithm (TLA) to take account of the inheritance-
of a sentence type can be used to determine whether
based partial ordering and the role of memory in
it can function as a trigger at a particular stage in
learning. The TLA is error-driven – parameter set-
learning.
tings are altered in constrained ways when a learner
cannot parse trigger input. Trigger input is de-
fined as primary linguistic data which, because of
Data: {S1 , S2 , . . . Sn } Variables Typical Values
Population Size 32
unless Interaction Cycle 2K Interactions
parseri (grammari (p-settingi (Sj ))) = Success Simulation Run 50 Cycles
then Crossover Probability 0.9
p-settingj = update(p-settingi ) Mutation Probability 0
unless Learning memory limited yes
parserj (grammarj (p-settingj (Sj ))) = Success critical period yes
then
RETURN p-settingsi Figure 9: The Simulation Options
else
RETURN p-settingsj (Cost/Benefits per sentence (1–6); summed for each
LAgt at end of an interaction cycle and used to cal-
Update: culate fitness functions (7–8)):
Reset the first (most general) default or unset pa-
rameter in a left-to-right search of the p-set accord- 1. Generate cost: 1 (GC)
ing to the following table:
2. Parse cost: 1 (PC)
Input: D1 D0 ? ?
(where 1 3. Generate subset language cost: 1 (GSC)
Output: R 0 R 1 ? 1/0 (random)
= T/L and 0 = F/R)
4. Parse failure cost: 1 (PF)
Figure 8: The Learning Algorithm 5. Parse memory cost: WML(st)
6. Interaction success benefit: 1 (SI)
3 The Simulation Model SI GC 1
7. Fitness(WML): GC+P C × GC+GSC × ( PW ML
C−P F )
The computational simulation supports the evolu-
SI GC
tion of a population of Language Agents (LAgts), 8. Fitness(¬WML): GC+P C × GC+GSC
similar to Holland’s (1993) Echo agents. LAgts gen-
erate and parse sentences compatible with their cur-
Figure 10: Fitness Functions
rent p-setting. They participate in linguistic inter-
actions which are successful if their p-settings are
compatible. The relative fitness of a LAgt is a func- eters which either were unset or had default settings
tion of the proportion of its linguistic interactions ‘at birth’. However, p-settings with an absolute
which have been successful, the expressivity of the value (principles) cannot be altered during the life-
language(s) spoken, and, optionally, of the mean time of an LAgt. Successful LAgts reproduce at the
WML for parsing during a cycle of interactions. An end of interaction cycles by one-point crossover of
interaction cycle consists of a prespecified number (and, optionally, single point mutation of) their ini-
of individual random interactions between LAgts, tial p-settings, ensuring neo-Darwinian rather than
with generating and parsing agents also selected ran- Lamarckian inheritance. The encoding of p-settings
domly. LAgts which have a history of mutually suc- allows the deterministic recovery of the initial set-
cessful interaction and high fitness can ‘reproduce’. ting. Fitness-based reproduction ensures that suc-
A LAgt can ‘live’ for up to ten interaction cycles, cessful and somewhat compatible p-settings are pre-
but may ‘die’ earlier if its fitness is relatively low. It served in the population and randomly sampled in
is possible for a population to become extinct (for the search for better versions of universal grammar,
example, if all the initial LAgts go through ten in- including better initial settings of genuine parame-
teraction cycles without any successful interaction ters. Thus, although the learning algorithm per se
occurring), and successful populations tend to grow is fixed, a range of alternative learning procedures
at a modest rate (to ensure a reasonable proportion can be explored based on the definition of the inital
of adult speakers is always present). LAgts learn set of parameters and their initial settings. Figure 9
during a critical period from ages 1-3 and reproduce summarizes crucial options in the simulation giving
from 4-10, parsing and/or generating any language the values used in the experiments reported in §4
learnt throughout their life. and Figure 10 shows the fitness functions.
During learning a LAgt can reset genuine param-
4 Experimental Results Unset Default None
4.1 Effectiveness of Learning Procedures WML 15 39 26
¬WML 34 17 29
Two learning procedures were predefined – a default
learner and an unset learner. These LAgts were ini- Table 2: Overall preferences for parameter types
tialized with p-settings consistent with a minimal in-
herited CGUG consisting of application with NP and
S atomic categories. All the remaining p-settings 4.2 Evolution of Learning Procedures
were genuine parameters for both learners. The un-
set learner was initialized with all unset, whilst the In order to test the preference for default versus un-
default learner had default settings for the parame- set parameters under different conditions, the five
ters gendir and subjdir and argorder which spec- parameters which define the difference between the
ify a minimal SVO right-branching grammar, as well two learning procedures were tracked through an-
as default (off) settings for comp and perm which other series of 50 cycle runs initialized with either 16
determine the availability of Composition and Per- default learning adult speakers and 16 unset learning
mutation, respectively. The unset learner represents adult speakers, with or without memory-limitations
a ‘pure’ principles-and-parameters learner. The de- during learning and parsing, speaking one of the
fault learner is modelled on Bickerton’s bioprogram eight languages described above. Each condition was
learner. run ten times. In the memory limited runs, default
Each learner was tested against an adult LAgt parameters came to dominate some but not all pop-
initialized to generate one of seven full lan- ulations. In a few runs all unset parameters dis-
guages in the set which are close to an at- appeared altogether. In all runs with populations
tested language; namely, “English” (SVO, predom- initialized to speak “English” (SVO) or “Malagasy”
inantly right-branching), “Welsh” (SVOv1, mixed (VOS) the preference for default settings was 100%.
order), “Malagasy” (VOS, right-branching), “Taga- In 8 runs with “Tagalog” (VSO) the same preference
log” (VSO, right-branching), “Japanese” (SOV, emerged, in one there was a preference for unset pa-
left-branching), “German” (SOVv2, predominantly rameters and in the other no clear preference. How-
right-branching), “Hixkaryana” (OVS, mixed or- ever, for the remaining five languages there was no
der), and an unattested full OSV language with left- strong preference.
branching syntax. In these tests, a single learner in- The results for the runs without memory limita-
teracted with a single adult. After every ten interac- tions are different, with an increased preference for
tions, in which the adult randomly generated a sen- unset parameters across all languages but no clear
tence type and the learner attempted to parse and 100% preference for any individual language. Ta-
learn from it, the state of the learner’s p-settings was ble 2 shows the pattern of preferences which emerged
examined to determine whether the learner had con- across 160 runs and how this was affected by the
verged on the same grammar as the adult. Table 1 presence or absence of memory limitations.
shows the number of such interaction cycles (i.e. the To test whether it was memory limitations during
number of input sentences to within ten) required by learning or during parsing which were affecting the
each type of learner to converge on each of the eight results, another series of runs for “English” was per-
languages. These figures are each calculated from formed with either memory limitations during learn-
100 trials to a 1% error rate; they suggest that, in ing but not parsing enabled, or vice versa. Memory
general, the default learner is more effective than limitations during learning are creating the bulk of
the unset learner. However, for the OVS language the preference for a default learner, though there
(OVS languages represent 1.24% of the world’s lan- appears to be an additive effect. In seven of the
guages, Tomlin, 1986), and for the unattested OSV ten runs with memory limitations only in learning, a
language, the default (SVO) learner is less effective. clear preference for default learners emerged. In five
So, there are at least two learning procedures in the of the runs with memory limitations only in parsing
space defined by the model which can converge with there appeared to be a slight preference for defaults
some presentation orders on some of the grammars emerging. Default learners may have a fitness ad-
in this set. Stronger conclusions require either ex- vantage when the number of interactions required to
haustive experimentation or theoretical analysis of learn successfully is greater because they will tend to
the model of the type undertaken by Gibson and converge faster, at least to a subset language. This
Wexler (1994) and Niyogi and Berwick (1995). will tend to increase their fitness over unset learners
who do not speak any language until further into the
Learner Language
SVO SVOv1 VOS VSO SOV SOVv2 OVS OSV
Unset 60 80 70 80 70 70 70 70
Default 60 60 60 60 60 60 80 70
30
60
No. of Speakers
25
%
20
40
15
10
20
0 0
0 10 20 30 40 50 60 70 0 5 10 15 20 25 30 35 40 45 50
Interaction Cycles Interaction Cycles