PhonBank Manual A Guide to the PhonBank Database

Transcription

PhonBank Manual A Guide to the PhonBank Database
PhonBank Manual
A Guide to the PhonBank Database
Web access:
PhonBank CHAT data: http://childes.psy.cmu.edu/data/PhonBank/
PhonBank Phon data: http://childes.psy.cmu.edu/data/PhonBank-Phon/
PhonBank media: http://childes.psy.cmu.edu/media/media/PhonBank/
CLAN Application: http://childes.psy.cmu.edu/clan/
Phon Application: http://phon.ling.mun.ca/phontrac/wiki/Downloads
Document prepared by Carla Peddle
In collaboration with Brian MacWhinney & Yvan Rose
Last revised: August 25, 2014
Table of Contents
Dutch-Levelt & Fikkert ................................................................................................3
Dutch-Zink ..................................................................................................................4
English-Chiat ...............................................................................................................5
English-Compton & Streeter ........................................................................................6
English-Davis ..............................................................................................................7
English-Inkelas .......................................................................................................... 10
English-Smith ............................................................................................................ 11
English-Stanford ....................................................................................................... 12
French-Goad & Rose ................................................................................................. 13
French-Kern .............................................................................................................. 14
French-Stanford ........................................................................................................ 15
German-PAIDUS........................................................................................................ 16
German-Stuttgart/TAKI ............................................................................................. 18
Japanese-Ota ............................................................................................................ 21
Japanese-Stanford .................................................................................................... 22
PAIDOLOGOS-Cantonese/English/Greek/Japanese .................................................... 23
Portuguese-Lisbon .................................................................................................... 24
Portuguese-Freitas .................................................................................................... 26
Romanian-Kern ......................................................................................................... 27
Spanish-PAIDUS ........................................................................................................ 28
Swedish-Stanford...................................................................................................... 30
Tunisian-Kern ........................................................................................................... 32
2
Dutch-Levelt & Fikkert
Project Name: Dutch-CLPF
Investigators: Claartje Levelt
Paula Fikkert
Contact:
[email protected]
[email protected]
Location:
Leiden & Groningen, NL
Number of Participants:
12
Nature of the Study:
babbling, full word, naturalistic, longitudinal
Media Type (if any):
wave files
Citation Information:
Fikkert, Paula (1994). On the Acquisition of Prosodic Structure. HIL Dissertations in Linguistics 6.
The Hague: Holland Academic Graphics.
Levelt, Clara (1994). On the Acquisition of Place. Leiden: HIL Dissertations in Linguistics 8. The
Hague: Holland Academic Graphics.
Participant Name
Date of Birth
Age range
Catootje
David
Elke
Enzo
Eva
Jarmo
Leon
Leonie
Noortje
Robin
Tirza
Tom
1988-10-24
1987-11-23
1988-04-06
1987-11-30
1988-06-06
1988-05-03
1987-12-10
1986-05-07
1988-02-10
1988-05-27
1988-04-13
1988-10-05
1;10.10 - 2;07.04
1;11.07 - 2;03.27
1;06.24 - 2;04.28
1;11.08 - 2;06.11
1;04.12 - 1;11.08
1;04.18 - 2;04.01
1;10.01 - 2;08.19
1;09.15 - 1;11.18
1;07.14 - 2;11.00
1;05.10 - 2;04.29
1;07.09 - 2;06.12
1;00.24 - 2;03.02
Number of
Sessions
16
6
19
16
12
23
23
7
21
23
20
25
Sex
F
M
F
M
F
M
M
F
F
F
F
M
Recordings were made in the child’s home during natural, spontaneous, interactive sessions with
one or both of the experimenters and occasionally with one of the parents. Typically a researcher
or parent would interact with the child by reading books or playing with toys and occasionally
asking the child what s/he saw in books or what s/he was doing.
During transcription, unintelligible utterances were left out, however, false starts, errors,
breakdowns, etc. were transcribed since they may provide valuable information about the child’s
competence. If the child’s utterances contained more than one word, boundaries were placed
between the words. No boundaries were placed if it was unclear whether the child knew the
different words of which the utterance consisted, such as “what is that?”.
This corpus was converted from the original CLPF database. We extend many thanks to Claartje
Levelt and Paula Fikkert for their assistance with the conversion process. The transcripts have
been linked to the available audio clips for five children: Catootje, Enzo, Eva, Robin, and Tirza.
We hope to add the audio for the other children at a later date.
3
At this stage, the data have not been entirely rechecked from the original transcriptions. Careful
spot-checking however suggests that the conversion is accurate. Please feel free to contact us or
the original authors in case you have questions.
The following lists illustrate some child-specific forms and onomatopoeia present in the childrens’
speech.
Child-specific forms
Bertje
Bumba
Dribbel
Jules
Willy
Zina
Lola
Tiny
Bambi
Musti
Sofie
Samson
Sjoewie
Referent
piglet
clown
dog
boy
animal
dog
chicken
girl
deer
cat
crocodile
dog
dog
Onomatopoeia
mei
pok
hi
kwak
ia
huuu
woef
miauw
kukeleku
Referent
Sheep, goat
chicken
horse
duck
donkey
horse
dog
cat
rooster
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
Dutch-Zink
Project Name: Dutch-Zink
Investigators: Inge Zink
Contact:
[email protected]
Location:
Leuven, Brabant and Antwerp, BE
Number of Participants:
4
Nature of the Study:
naturalistic, longitudinal
Media Type (if any):
Mini DV Cassettes (later digitized to DVDs)
Citation Information:
Zink, I. (2005). De verwerving van de klankproductie tijdens de brabbelperiode bij vier Vlaamse
kinderen. Logopedie, 18,4, 13-20.
Kern, S., Davis, B., Zink, I. (2009). From babbling to first words in four languages: Common
trends across languages and individual differences. In: d'Errico F., Hombert J. (Eds.),
Becoming Eloquent. Advances in the emergence of language, human cognition, and
modern cultures (pp. 205-232). Philadelphia:. John Benjamins Publishing Company.
4
Participant Name
Date of Birth
Age range
Number of
Sessions
Sex
David
2002-09-03
08.04 – 2;02.14
38
M
Judith
2003-06-13
08.03 – 2;00.26
39
F
Laurien
2002-03-08
08.07 – 2;01.26
37
F
Meinder
2002-05-07
08.14 – 2;02.05
38
M
The recordings for this corpus were made in Leuven, Brabant, Belgium (3 children: Meinder,
Judith, Laurien) and Antwerp, Belgium (1 child: David). The participants were recorded every two
weeks from 8 months to 25 months of age. Each recording session lasted approximately 60
minutes. The team made recordings using a Panasonic Digital Video Camera (NV-GS180) and a
Sennheiser ew100 microphone on 60 minute fuji Mini DVCassettes. The video files were later
digitized to DVDs. During the recording sessions, the child interacted with his/her mother (or both
parents) and a speech therapist and/or a master’s student of speech-language pathology was
present.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
English-Chiat
Project Name:
English-Chiat
Investigator(s):
Shula Chiat
Contact:
[email protected]
Location:
UK
Number of Participants:
3
Nature of study:
clinical, naturalistic, longitudinal
Media Type (if any):
Audio files (wave files)
Citation Information:
Chiat, S. (1983). Why Mikey’s right and My key’s wrong: The significance of stress and word
boundaries in a child’s output system. Cognition, 14, 275-300.
Chiat, S. (1989). The relation between prosodic structure, syllabification and segmental
realization: Evidence from a child with fricative stopping. Clinical Linguistics and
Phonetics, 3, 223-42
Chiat, S. (1994). From lexical access to lexical output: What is the problem for children with
impaired phonology? In M. Yavas (Ed.), First and second language phonology. San
Diego: Singular Publishing Group.
5
Participant Name
Date of Birth
Age range
DI
1983-04-06
5;05.16 – 5;05.28
Number of
Sessions
4
SB
1982-11-27
5;05-26 – 5;07.09
4
M
SR
1976-01-30
5;07.21 – 5;08.12
3
M
Sex
M
Recordings for this project were carried out in a quiet room at school or a clinic. Each child can be
considered its own single case study. Utterances were spontaneous in nature through play
sessions with the researcher. There were also some elicited repetitions of real words, non-words
and phrases to sample the realization of segmental targets in different prosodic positions at the
word- and phrase-level.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
_____________________________________________________________________________
English-Compton & Streeter
Incomplete
Project Name:
English-Compton & Streeter
Investigator(s):
A. J. Compton & M. Streeter
Contact:
Joe Pater
[email protected]
Location:
U.S.A.
Number of Participants:
4
Nature of study:
naturalistic, longitudinal, monolingual
Media Type (if any):
TBA
Citation Information:
Compton, A. J., and M. Streeter. (1977). Child Phonology: Data Collection and Preliminary
Analyses, in Papers and Reports on Child Language Development 7. Stanford University,
Stanford, California.
Pater, Joe. (1997). Minimal Violation and Phonological Development. Language Acquisition, 6:3,
201-253.
Participant Name
Date of Birth
Age range
Derek
Julia
Sean
Trevor
TBA
TBA
TBA
TBA
1;06 – 3;02.01
1;02.21 – 3;01.03
1;01.25 – 3;02.20
0;8 – 3;01.08
Number of
Sessions
TBA
TBA
TBA
TBA
Sex
M
F
M
M
A team, under the direction of A. J. Compton, in the 1970s originally collected these data. The
method of data collection and some preliminary analyses, were presented in Compton & Streeter
6
(1977). The project was undertaken to map out, as precisely as possible, the development of
children’s sound systems. With this goal in mind, a diary method of data collection was chosen,
with parents keep track of their children’s utterances by recording them in notebooks at least four
days a week and scattered throughout the child’s waking hours, covering about four hours a day.
The parents were speech pathologists and received additional training in the phonetic
transcription of child speech prior to the study.
All of the children were learning American English as spoken in California. None of the children
had any language or learning-related impairments.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
_____________________________________________________________________________
English-Davis
Project Name:
English-Davis
Investigator(s):
Barbara Davis
Contact:
[email protected]
Location:
Austin, TX
Number of Participants:
21
Nature of study:
naturalistic, longitudinal
Media Type (if any):
Reel-to-reel and DAT Audio tape
Citation Information:
Davis, Barbara L., Peter F. MacNeilage & Christine L. Matyear. (In Press). Acquisition of Serial
Complexity in Speech Production: A Comparison of Phonetic and Phonological
Approaches to First Word Productions. Phonetica.
Dolata, Jill K., Davis, B L. & MacNeilage, P F. (2008). Characteristics of the Rhythmic
Organization of Babbling: Implications for an Amodal Linguistic Rhythm, Infant Behavior
and Development, 31, 422-431.
Serkhane, J.E., Schwartz, J.L., Boë, L J., Davis, B. L., Matyear, C.L. (2007). Infants’ vocalizations
analyzed with an articulatory model: A preliminary report, Journal of Phonetics.
MacNeilage, P.F. & Davis, B.L. (2005). Invited submission, Functional Organization of Speech
Across the Life Span: A Critique of Generative Phonology, Linguistics Review, 22 161181.
Davis, B. L., MacNeilage P.F. & Matyear, C. (2002). Acquisition of serial complexity in speech
production: A Comparison of Phonetic and Phonological Approaches to First Word
Production. Phonetica, 59, 75-107.
Serkhane, J., Schwartz, J-L., Boe, L.J., Davis, B.L., Bessiere, P., & Matyear, C. (2002). Premiere
etude comparative de vocalizations de bebes humains et de bebes robots, Journées
d'Etudes sur la parole, 22-28.
Davis, B.L. & MacNeilage, P.F. (2002). The Internal Structure of the Syllable: An Ontogenetic
Perspective on Origins in Givon, T. & Malle, B (Eds.) The Rise of Language out of Prelanguage. John Benjamin Publ.
MacNeilage, P.F & Davis, B.L. (2002). On The Origins of Intersyllabic Complexity. In Malle, B.
(Ed.) The Rise of Language out of Pre-language, Benjamin
MacNeilage, P.F. & Davis, B.L. (2001). Commentary: The role of rhythmic cyclicities in infant
7
action development, Developmental Science. 4 (1) 79-83.
MacNeilage, P.F. & Davis, B.L. (2000). Origin of the Internal Structure of Words. Science. 288,
527-531.
Davis, B.L. & MacNeilage, P.F. (2000). An embodiment perspective on the acquisition of speech
perception. Special Issue. Phonetica. 57, 229-241.
MacNeilage, P.F. & Davis, B.L (2000). Deriving speech from non-speech: A view from ontogeny,
Special Issue, Phonetica. 57,284-296.
Davis, B.L., MacNeilage, P.F., Matyear, C.L. & Powell, J., (2000). Prosodic correlates of stress in
babbling: An acoustical study. Child Development. Vol. 71, (5) 1258-1270.
Gildersleeve-Neumann, C.E., Davis, B.L. & MacNeilage, P.F. (2000). Contingencies governing
production of for fricatives, affricates and liquids in babbling. Applied Psycholinguistics.
21-341-363.
MacNeilage, P.F., Davis, B.L., Kinney, A. & Matyear, C.M. (2000). The motor core of speech: A
comparison of serial organization patterns in infants and languages. Invited submission
to Special Millenium Issue. Child Development. 71 (1), 153-163.
MacNeilage, P.F., Davis, B.L., Matyear, C.M. & Kinney, A. (2000). Origin of Speech Output
Complexity in Infants and in Languages. Psychological Science . 10 (5) 459-460.
Vihman, M.M., DePaolis, R. & Davis, B.L. (1998). Prosody in early word productions: A study of
Trochees and Iambs, Child Development.
Redford, M.L., MacNeilage, P.F. & Davis, B.L. (1997). Perceptual and motor influences on final
consonant inventories in babbling. Phonetica. 54, 172-186.
Davis, B.L. & MacNeilage, P.F. (1995). The articulatory basis of babbling. Journal of Speech and
Hearing Research, 38, 1199-1211.
Davis, B.L. & MacNeilage, P.F. (1994). Organization of Babbling: A case study. Language &
Speech, 37 (4), 341-355.
MacNeilage, P.F. & Davis, B.L. (1990). Acquisition of speech production: Achievement of
segmental independence. In Hardcastle, W.I. & Marchal A. (eds.) Speech Production and
Speech Modeling, Dordrecht: Kluwer, 55-68.
MacNeilage, P.F. & Davis B.L. (1990). Acquisition of speech production: Frames, then content. In
Jeannerod, M. (ed.) Attention and Performance XIII; Motor Representation and Control,
Hillsdale, N.J.: LEA, 452-468.
Davis, B.L. & MacNeilage, P.F. (1990). The acquisition of vowels: a case study. Journal of
Speech and Hearing Research, 33, 16-27.
8
Participant Name
Date of Birth
Age range
Aaron
1981-01-14
1;08.01 – 2;01.03
Number of
Sessions
11
Anthony
1981-05-03
1;08.05 – 2;01.15
13
M
Ben
2002-04-06
0;11.21 – 2;04.02
21
M
Cameron
1992-02-15
0;07.11 – 2;11.24
52
F
Charlotte
2002-10-29
0;10.12 – 2;11.22
44
F
Georgia
2002-08-27
0;08.25 – 2;11.05
45
F
Hannah
2002-08-23
0;11.14 – 2;04.24
29
F
Jodi
1981-09-07
1;00.29 – 1;07.09
14
F
Kaeley
2002-11-21
1;00.24 – 2;01.23
13
F
Kate
1981-05-29
1;03.02 – 1;07.25
10
F
Martin
1993-02-22
1;05.19 – 2;02.09
13
M
Micah
1988-09-13
0;08.01 – 1;06.19
29
M
Nate
2001-08-06
0;10.07 – 2;09.07
26
M
Nick
1991-12-09
0;10.20 – 3;01.02
46
M
Paxton
1992-03-19
0;08.02 – 2;00.02
47
M
Rachel
1992-01-30
0;08.04 – 1;10.02
40
F
Rebecca
1983-09-17
1;01.24 – 1;07.17
41
F
Rowan
2001-11-11
0;10.23 – 2;10.19
41
M
Sadie
1992-04-14
0;07.11 – 1;07.03
33
F
Sam
2001-08-24
0;10.07 – 2;01.03
26
M
Will
1992-03-27
0;07.29 – 1;04.13
24
M
Sex
M
Data analyzed for this study were collected as a part of a longitudinal project tracking early
normal speech development from the onset of canonical babbling through age 3.5 years.
Subjects were located by informal referral from the surrounding community. Normal development
was established through parent case history report. In addition, each infant was administered the
Battelle Developmental Screening Inventory (Guidubaldi, Newborg, Stock, Svinicki, & Wneck,
1984) and hearing screening using sound field techniques.
Video and audio recordings were gathered in the children’s homes during natural interactions and
situations that occurred in their daily lives. The children were interacting with their parent or
others present in their home setting. The experimenter sometimes interacted with the child and
parent and sometimes observed as seemed natural in the setting.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
_____________________________________________________________________________
9
English-Inkelas
Project Name:
English-Inkelas
Investigator(s):
Sharon Inkelas
Contact:
[email protected]
Location:
Northern California
Number of Participants:
1
Nature of study:
naturalistic, diary, longitudinal
Media Type (if any):
none
Citation Information:
Inkelas, Sharon & Yvan Rose. 2003. Velar Fronting Revisited. In Barbara Beachley, Amanda
Brown & Fran Conlin (eds.), Proceedings of the 27th Annual Boston University
Conference on Language Development, 334-345. Somerville, MA: Cascadilla Press.
Inkelas, Sharon & Yvan Rose. 2008. Positional Neutralization: A Case Study from Child
Language. Language 83, 707-736.
Participant Name
Date of Birth
Age range
Eli
1997-12-22
0;06.08 - 3;09.28
Number of
Sessions
200
Sex
M
This study is based primarily on a new longitudinal corpus from E, a typically developing
monolingual learner of English with no history of hearing problems or language delays. E has one
sibling, a brother who is two years and five months older and who also exhibited normal
phonological development. E’s mother is a native speaker of American English. E’s father is a
native speaker of Turkish, but fluent in English; his slight Turkish accent is unlikely to have been a
factor in E’s phonological patterns. During the course of language development, E was raised in a
monolingual English-speaking environment. He was exposed to minimal amounts of Turkish but
was virtually never addressed directly in this language. As of age 2;7, due to his sibling’s entry
into a Spanish immersion school program, E began to be exposed to some Spanish, including the
word hola, but not by native Spanish speakers.
The data from E were gathered in a naturalistic, diary setting primarily by his mother, a trained
phonologist. E’s father, also a phonologist, participated in the data collection as well. Data
collection was performed as often as possible, during unstructured sessions and regular family
activities. The period of data gathering spans E’s ages 0;06.09 to 3;09.29, with most of the data
collected prior to the last year of the study. Between ages 0;06.09 and 2;09.09, a span that
corresponds to E’s phonological development peak (i.e. during which most aspects of the target
grammar were acquired), a total of 3,267 words (distributed over 1,713 utterances) were
transcribed phonetically. An additional 113 words were transcribed over the last year of the study,
during which E’s mother was attending mainly to the development of /l/.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
_____________________________________________________________________________
10
English-Smith
Project Name:
English-Smith
Investigator(s):
Neilson Smith
Contact:
Heather Goad
[email protected]
Location:
U.K.
Number of Participants:
Tania Zamuner
[email protected]
1
Nature of study:
naturalistic, longitudinal, monolingual, diary.
Media Type (if any):
occasional tape recordings
Citation Information:
Smith, Neilson V. (1973). The Acquisition of Phonology: A Case Study. Cambridge: Cambridge
University Press.
Participant Name
Date of Birth
Age range
Amahl
1967-06-04
2;2.00 – 3;09.13
Number of
Sessions
29
Sex
M
The child in the case study, Amahl, is the son of an English father with some derivations from
“received pronunciations” and Indian mother who speaks English only as a fourth language
(following Hindi, Bengali, Marathi). Amahl was born in Boston, Massachusetts and moved to the
U.K. at approximately one year of age, which may have influenced his late speaking. The child
lived with his paternal aunt from 19 months to 2 years, and started attending day nurseries and
play groups in Bedfordshire and Hertfordshire at 30 months.
The data discussed were recorded in IPA transcriptions on index cards. Some tape recordings
were made occasionally, however, most of the description is based on non-recorded material.
There were instances of spontaneous speech, however the majority of the words were elicited
with phrases such as “can you say “zink” for me?”.
This corpus was retyped and adapted from the original diary study. A significant portion of this
work was undertaken by Tania Zamuner; additional work was performed under an FCAR grant to
Heather Goad. The recording dates in the corpus correspond to the first date of Smith's (1973)
descriptive stages. Please note that the data have not been entirely rechecked from Smith's
original transcriptions. Careful spot-checking however suggests that the conversion is accurate.
Please feel free to contact us in case you have questions or if you notice any errors.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by
the above reference.
11
English-Stanford
Incomplete
Project Name:
English-Stanford
Investigator(s):
Marilyn Vihman
Contact:
Location:
[email protected]
Stanford area, CA
Number of Participants:
5
Nature of study:
naturalistic, longitudinal
Media Type (if any):
audio cassettes and/or reel-to-reel/VHS video
Citation Information:
Vihman, M. M., Macken, M.A., Miller, R., Simmons, H. & Miller, J. (1985). From babbling to
speech: a reassessment of the continuity issue. Language, 61, 395-443.
Vihman, M. M., C.A. Ferguson & M. Elbert. (1986). Phonological development from babbling to
speech: Common tenden-cies and individual differences. Applied Psycholinguistics, 7, 340.
Boysson-Bardies, B. de & Vihman, M. M. (1991). Adaptation to language: Evidence from
babbling and first words in four languages. Language, 67, 297-319.
Vihman, M. M., Kay, E., Boysson-Bardies, B. de, Durand, C. & Sundberg, U. (1994). External
sources of individual differences? A cross-linguistic analysis of the phonetics of mothers'
speech to one-year-old children. Developmental Psychology, 30, 652-663.
Vihman, M. M., DePaolis, R.A. & Davis, B.L. (1998). Is there a “trochaic bias” in early word
learning? Evidence from English and French. Child Development, 69, 933-947.
Participant Name
Date of Birth
Age range
Deb
TBA
0;09.17 – 1;03.24
Emi
Mol
Sea
Tim
TBA
TBA
TBA
TBA
0;10.07 – 1;02.39
0;08.17 – 1;02.20
0;10.00 – 1;03.23
0;10.08 – 1;04.22
Number of
Sessions
6
6
6
6
6
Sex
F
F
F
M
M
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
12
French-Goad & Rose
Project Name: French-GoadRose
Investigators: Heather Goad
Yvan Rose
Contact:
[email protected]
[email protected]
Location:
Québec City & just outside Montréal, QC
Number of Participants:
2
Nature of the Study:
babbling to full word, naturalistic, longitudinal
Media Type (if any):
wave files
Citation Information:
Rose, Yvan (2000). Headedness and Prosodic Licensing in the L1 Acquisition of Phonology.
Ph.D. Dissertation. McGill University.
Participant Name
Date of Birth
Age range
Clara
Théo
1996-03-26
1994-11-25
1;00.27-2;07.19
1;10.26-4;00.00
Number of
Sessions
34
45
Sex
F
M
The participants in this project were recorded in their homes in a naturalistic setting, generally in
the absence of siblings, approximately every second week. Clara was recorded by her mother, a
sociolinguist and Théo was recorded by a female French-speaking family member who was a
linguistics graduate student from McGill.
During the recordings the children typically looked at picture books and played with toys. The
interviewer encouraged spontaneous word productions by the child to avoid speech sample
repetition, and repeated the words produced by the child to facilitate identification of data for
extraction and transcription.
The recordings were made using an analog recording machine and a multidirectional
microphone, which was set up near the child on a foamy cushion the floor to reduce interfering
noises from movements and toys. The tapes were digitized using SoundEdit 16v2 in 16 bit
sample size at a rate of 22050 kHz.
This corpus was converted from the original database of Québec French used by Rose (2000).
The data were collected under an FCAR grant to Heather Goad. Please note that the data have
not been entirely rechecked from the original transcriptions. Careful spot-checking however
suggests that the conversion is accurate. Please feel free to contact the contributor in case you
have questions.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by
the above reference.
13
French-Kern
Project Name:
French-Kern
Investigator(s):
Sophie Kern
Contact:
[email protected]
Location:
Lyon, FR
Number of Participants:
4
Nature of study:
naturalistic, longitudinal, monolingual
Media Type (if any):
digital video (.mov, .mp4)
Citation Information:
Kern, S, B.L. Davis, & I. Zink. (2009). From babbling to first words in four languages: Common
trends, cross language and individual differences. In Francesco d’Errico & Jean-Marie
Hombert (eds) Becoming eloquent: Advances in the Emergence of language, human
cognition and modern culture. John Benjamins' Publishing Company.
Kern, S. & B.L. Davis .(2009). Emergent complexity in early vocal acquisition: Cross-linguistic
comparisons of canonical babbling. In F. Pellegrino, E. Marsico, I. Chitoran & C. Coupé
(eds), Approaches to phonological complexity. Berlin, Mouton de Gruyter.
Kern, S. (2005). De l'universalité et des spécificités du développement langagier précoce. In
Hombert, J.M. (ed) Aux origines du langage et des langues. Fayard
Davis, B., S. Kern, A. Vilain & C. Lalevée (2008). Des babils à Babel: les premiers pas de la
parole. Revue Française de Linguistique Appliquée. 13(2), 81-91.
Participant Name
Date of Birth
Age range
Baptiste
Emma
Esteban
Jules
2000-09-11
2000-10-12
2001-10-23
2001-03-28
0;09.07 - 02;00.12
0;08.16 - 2;01.08
0;07.15 – 2;00.24
0;09.13 - 2;00.28
Number of
Sessions
35
36
30
28
Sex
M
F
M
M
Five types of data were collected. First, one hour of spontaneous vocalization data was audio and
video recorded every two weeks from 8 months of age through 25 months of age. Recording took
place in children’s homes. The parents were told to follow their normal types of activities with their
child. Second, minimally 1,000 dictionary entries from the ambient language employed by the
parents of each child participant were analyzed for comparison with the child data for that
language. Parental reports were administered using adaptations of the MacArthur Development
Inventories (Fenson et al., 1993) respectively elaborated for Dutch for French participants.
Mothers filled out the questionnaire once in a month. For the remaining languages there is no
adaptation yet, but one could imagine using the spontaneous data to elaborate the same
instrument. An object manipulation categorization task was administered every two months. This
task was conceived to evaluate the children’s spontaneous nonverbal categorization abilities.
Several toys, which were consistent across the language groups served as stimuli. Each task
involved a contrast of objects from two different categories (animal, means of locomotion,
furniture).
Children were developing normally by community standards as well as reports from parents and
physicians regarding developmental milestones were observed in the normal daily environment.
14
These children were becoming monolingual speakers of French, Romanian and Tunisian Arabic.
Languages were chosen to include different language families as well as for the contrasts they
show in phonetic and phonological features of interest for early language development. These
characteristics include word length, syllabic types as well as phonemic inventory diversity
One hour of spontaneous vocalization data was audio and video recorded every two weeks from
8 months of age through 25 months of age. Recording took place in children’s homes. The
parents were told to follow their normal types of activities with their child. After collection, data
was phonetically transcribed using the International Phonetic Alphabet (IPA). Broad phonetic
transcriptions were used, supplemented by some diacritics (mainly for palatalized,
pharyngealized, nasalized sounds and duration of sounds). Tokens considered as single
utterance strings were bounded by one second of silence, noise or adult speech.
This work was supported by the EUROCORE Program “The Origin of Man, Language and
Languages” (OMLL) and the French CNRS program “Origine de l’Homme, du Langage et des
Langues” (OHLL)
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
French-Stanford
Incomplete
Project Name:
French-Stanford
Investigator(s):
Marilyn Vihman
Contact:
[email protected]
Location:
Paris, FR
Number of Participants:
5 (or 6)
Nature of study:
naturalistic, longitudinal
Media Type (if any):
audio cassettes and/or reel-to-reel/VHS video
Citation Information:
Vihman, M. M., Macken, M.A., Miller, R., Simmons, H. & Miller, J. (1985). From babbling to
speech: a reassessment of the continuity issue. Language, 61, 395-443.
Vihman, M. M., C.A. Ferguson & M. Elbert. (1986). Phonological development from babbling to
speech: Common tenden-cies and individual differences. Applied Psycholinguistics, 7, 340.
Boysson-Bardies, B. de & Vihman, M. M. (1991). Adaptation to language: Evidence from
babbling and first words in four languages. Language, 67, 297-319.
Vihman, M. M., Kay, E., Boysson-Bardies, B. de, Durand, C. & Sundberg, U. (1994). External
sources of individual differences? A cross-linguistic analysis of the phonetics of mothers'
speech to one-year-old children. Developmental Psychology, 30, 652-663.
Vihman, M. M., DePaolis, R.A. & Davis, B.L. (1998). Is there a “trochaic bias” in early word
learning? Evidence from English and French. Child Development, 69, 933-947.
Participant Name
Date of Birth
Age range
Number of
Sessions
Sex
15
Car
Cha
Kur
Lau
Mar
TBA
TBA
TBA
TBA
TBA
0;10.15 – 1;02.05
0;10.16 – 1;03.19
Noe
1985-12-28
0;09.26 – 1;05.15
0;11.21 – 1;07.24
6
7
2
6
7
F
M
M
M
F
0;10.06 – 1;05.23
6
F
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
German-PAIDUS
Project Name:
German-PAIDUS
Investigator(s):
Conxita Lleó
Antonio Maldonado
Contact:
[email protected]
Universidad Autónoma de Madrid
Location:
Hamburg, Germany
Number of Participants:
5
Nature of study:
naturalistic, babbling, first words, spontaneous, informal testing
Media Type (if any):
Wave files
Citation Information:
Lleó, Conxita, Christliebe El Mogharbel & Michael Prinz (1994). Babbling und FrühwortProduktion im Deutschen und Spanischen. Linguistische Berichte 151, 191-217.
Lleó, Conxita & Michael Prinz (1996). Consonant clusters in child phonology and the
directionality of syllable structure assignment. Journal of Child Language 23, 31-56.
Lleó, Conxita, Michael Prinz, Christliebe El Mogharbel & Antonio Maldonado (1996). Early
phonological acquisition of German and Spanish: A reinterpretation of the continuity
issue within the Principles and Parameters model. In C.E. Johnson & J.H.V. Gilbert
(eds.), Children's Language, Volume 9, 11-31. Hillsdale: Lawrence Erlbaum
Associates.
Lleó, Conxita & Michael Prinz (1997). Syllable Structure parameters and the acquisition of
affricates. In S.J. Hannahs & M. Young-Scholten (eds.), Focus on Phonological
Acquisition, 143-163. Amsterdam/Philadelphia: John Benjamins Publishing Co.
Kehoe, Margaret & Conxita Lleó (2003b). The acquisition of nuclei: a longitudinal analysis of
phonological vowel length in three German-speaking children. Journal of Child
Language 30 (3), 527-556.
Kehoe, Margaret & Conxita Lleó (2003c). A Phonological Analysis of Schwa in German First
Language Acquisition. Canadian Journal of Linguistics 48 (3/4), 289-327.
Lleó, Conxita, Imme Kuchenbrandt, Margaret Kehoe & Cristina Trujillo (2003). Syllable final
consonants in Spanish and German monolingual and bilingual acquisition. In N. Müller
(ed.), (In)vulnerable Domains in Multilingualism, 191-220. Amsterdam, Philadelphia:
John Benjamins.
Kehoe, Margaret, Conxita Lleó & Margaret Rakow (2004). Voice Onset Time in Bilingual
German- Spanish Children. Bilingualism: Language and Cognition 7, 71-88.
Lleó, Conxita (2006). The Acquisition of Prosodic Word Structures in Spanish by Monolingual
16
and Spanish-German Bilingual Children. Language and Speech 49 (2), 207-231.
Lleó, Conxita & Javier Arias (2006). Foot, word and phrase constraints in first language
acquisition of Spanish stress. In Fernando Martínez-Gil & Sonia Colina (eds.),
Optimality-Theoretic Studies in Spanish Phonology, 470-496. John Benjamins.
Lleó, Conxita & Martin Rakow (2006). The prosody of early two-word utterances by German
and Spanish monolingual and bilingual children. In C. Lleó (ed.), Interfaces in
Multilingualism: Acquisition and representation, 1-26. Amsterdam/Philadelphia: John
Benjamins.
Lleó, Conxita, Martin Rakow & Margaret Kehoe (2007). Acquiring rhythmically different
languages in a bilingual context. In Proceedings of the 16th International Congress of
Phonetic Sciences (ICPhS XVI), 1545-1548.
Participant Name
Date of Birth
Age range
Bernd
1989-12-15
0;08.22 – 3;02.24
Number of
Sessions
34
Britta
Johannes
Marion
Thomas
1989-08-10
1989-10-31
1989-08-14
1989-08-06
0;09.01 – 3;03.15
0;09.00 – 2;10.10
0;08.27 – 3;03.11
0;09.26 – 3;05.12
38
30
35
38
Sex
M
F
M
F
M
The project BIDS/PAIDUS was carried out at the University of Hamburg (Department of Romance
Languages) with the support of the Deutsche Forschungsgemeinschaft under the direction of
Conxita Lleó. The project began in 1990 first as BIDS, with the goal of comparing the phonetic
characteristics of babbling in German and comparing it with babbling in Spanish. With this
purpose, a collaboration with Antonio Maldonado of the Universidad Autónoma de Madrid (Dept.
of Psychology) was started and we applied for a joint grant from “Acciones Integradas”, so that
the two teams could visit each other. The German data were collected in Hamburg, Germany,
and the Spanish data in Madrid, Spain.
All five German children had been born in Hamburg and lived there. Their parents were all
German and spoke the Northern variety of German. The children were visited from very early on,
before they began to produce their first words, as we wanted to study their babbling. Some
informal tests, in order to figure their (passive) vocabularies were conducted. After the children
began saying words, we continued recording them within the project PAIDUS, which had as its
goal to finding out how phonological parameters were possibly fixed for German as opposed to
Spanish. The German children were visited every 2 weeks at first, and later on once a month.
Recordings were done at the children’s homes by two project research assistants. At first, the
mother used to be present, but later on, we made most of the recordings with the child alone
interacting with one research assistant. We tried to record the child under spontaneous
conditions, playing, chatting, etc. We also brought some objects (little animals or toys) in order to
elicit certain words from all children, making their phonetic and phonological characteristics
comparable. The mother tried to keep a record of possible new words that were new in the child’s
vocabularies, and we tried to elicit them, as well. We also used books with pictures, which
belonged to the child and others that we brought. The recordings were audio recordings, and in
order not to miss the ambient setting and context, one of the researchers was writing down all
possible details, including what the child was doing, while the other assistant was interacting with
the child.
Both projects, BIDS and PAIDUS, were financed by the DFG with grant to Conxita Lleó, BIDS
from 1989 until 1991 and PAIDUS from 1991 until 1994. Research Assistants in the project BIDS
were Christliebe El Mogharbel and Kerstin Feuge. Research Assistants in the project PAIDUS
were Christliebe El Mogharbel and Michael Prinz.
17
Reserachers using the BIDS/PAIDUS data should send a copy of their publication based on the
data to the contributor of the data. In accordance with CHILDES rules, any use of data from this
corpus must be accompanied by at least one of the above references.
German-Stuttgart/TAKI
Project Name:
German-Stuttgart/TAKI
Investigator(s):
Bernd Möbius
Contact:
Britta Lintfert
[email protected]
Location:
[email protected]
Stuttgart area, Southern Germany
Number of Participants:
11
Nature of study:
naturalistic, longitudinal, spontaneous, parental interaction, naming
Media Type (if any):
DAT tapes
Citation Information:
Allen, G. D. (1980): Development of Prosodic Phonology in Children’s Speech: Further Evidence
from the TAKI Task. In: Proceedings of the Fourth International Phonology Meeting.
Grimm, H. and H. Doil (2004): ELFRA, Elternfragebögen für die Früherke nnung von Risikokindern. Hogrefe, Verlag für Psychologie.
Stoel-Gammon, C. (1989): Prespeech and early speech development of two late talkers. First
Language 9, 207–224.
Participant Name
Date of Birth
Age range
Maarten (MS)
Lena (LL)
Tjark (TS)
1998-09-08
1999-03-23
2000-11-19
3;9 – 8;0
3;3 – 7;6
1;6 – 2;4
Number of
Sessions
40
46
42
Evelyn (ED)
2002-02-21
2;3 – 4;7
105
F
Robin (RL)
Emma (EL)
Hannah (HH)
Olli (OZ)
Bennie (BW)
Nils (NB)
Rike (FZ)
2002-02-21
2002-06-22
2003-08-02
2004-03-01
2004-04-22
2004-08-20
2005-09-05
3;2 – 4;7
0;5 - 4;3
0;4 – 3;1
0;8 – 2;9
0;11 – 2;11
0;8 – 2;8
1;0 – 2;6
58
M
F
F
M
M
M
F
7
11
10
32
32
14
Sex
M
F
M
The Stuttgart Corpus was constructed as part of a study on the acquisition of stress in German
founded by the German Research Foundation. In this project the researchers recorded and
analyzed acoustic data of 11 children from 6 months up to 15 years to develop an exemplarbased model of stress acquisition in German and to build up a prosodic annotated speech corpus
of babbling, first words and meaningful speech from German speaking children. The following
table summarizes information about the children and the recordings.
18
Child
MS
LL
TS
ED
RL
EL
HH
OZ
BW
NB
FZ
First Recording
2002-05
2002-06
2002-05
2004-05
2005-04
2002-11
2003-11
2004-11
2005-03
2005-04
2005-09
Babbling
Mixing
1;6-2;4
0;5-1;6
0;4-1;6
0;8-1;6
0;11-1;6
0;8-1;6
1;0-1;6
1;6-3;0
1;6-2;9
1;6-2;9
1;6-2;11
1;6-2;8
1;6-2;6
Taki-task
3;9-8;0
3;3-7;6
2;4-7;6
2;3-4;7
3;2-4;7
3;0-4;3
2;9-3;1
The recordings were made at the children’s homes in familiar play situations with their parents.
We recorded using a Sony DAT TCD-D100 and a high-quality wireless microphone NADY LT-4
(Lavalier) E-701 (600 Ohm). We tried to keep the distance between the microphone and the
mouth as constant as possible during the entire session to obtain high-quality recordings for
acoustic analysis. The recorded data were transferred to a computer and down sampled to 16
kHz.
1.
Babbling and first words: 5–18 months of age
Between 5 and 18 months of age, the speech data from six children (3boys, 3 girls) were
collected. The infants were audio-recorded every 6–8 weeks, starting between five and seven
months, when first CV-syllable productions occur, at their homes in familiar play situations with
their parents. No person unfamiliar to the child was present during the recordings. During a
recording session a parent (normally the mother) played with the child and later on, at the oneword stage, looked and talked about a picture book and picture cards. Before the recordings
started, the parents were instructed on how to use the recording equipment. The parents’
utterances were recorded and analyzed as well. All of the infants lived in monolingual Germanspeaking families and had no unusual prenatal, sensory or developmental concerns or hearing
problems. At the age of 12 months, the speech development of all infants was tested using a
parental questionnaire for early recognition of children at risk (Grimm & Doil, 2004).
2.
Mixing-phase: 18–36 months of age
Between 18 and 36 months of age, we collected speech data every 6–8 weeks. Because we
have different recording tasks developed for this age we called this phase mixing. The recordings
were done until the third birthday in the way described for the babbling phase. The children born
in 2004 (OZ, BW, NB, FZ) were also tested at each recording session for their use of stress in
multisyllabic words.
3.
Card Naming (from 18 months of age)
We created picture cards representing two- and three-syllable German words with stress on the
first, second and third syllable with the vowels /a/, /i/, and /o/ in stressed and unstressed position.
From the age of 18 months, children were tested with these cards. With this task, the
development of different stress schemes can be tested and described. The words were: Akrobat,
Anorak, Banane, Bikini, Buchstabe, Eisenbahn, Elefant, Eskimo, Fliege, Flughafen, Fotograf,
Gardine, Giraffe, Gitarre, Gorilla, Kamel, Kanone, Kobolde, Kokosnuss, Korkodil, Lastwagen,
Lawine, Malerin, Matrose, Mikrofon, Müllwagen, Paket, Papagei, Pinguin, Pistole, Polizei,
Postauto, Postbote, Prinzessin, Pullover, Radfahrer, Rennfahrer, Sandale, Saxophon,
Schmetterling, Skifahrer, Spiegel, Stadion, Teddybär, Tomate, Trampolin, Trompete, Vulkan(e),
Wohnwagen, and Zitrone.
19
4.
TAKI task: from 36 months of age
Beginning at 36 months, recordings were made each 10–12 weeks using the TAKI task proposed
by Allen (1980). We created five pairs of animal toys and assigned nonce names to each. Within
each pair, the nonce names differed only in terms of the position of main stress. For example, in a
pair with a brown bear and a polar bear. Both were called “bimo”, but the stress was on the first
syllable for the brown bear and the second syllable for the polar bear. The nonce names were all
bi-syllabic or tri-syllabic and consisted of consonant-vowel (CV) syllables formed from vowels /a i
o/ and consonants /b d m n/. All of these stress patterns are possible in German, although stress
on the second or third syllable is much less common than stress on the first syllable.
Annotation labels
1. CV coding (.cv). To study the development of syllable structure, we marked the beginning
and ending of each vowel and consonant.
2. Stable Vowel (.marks). The beginning and end of the stable phase of the vowel is marked.
No influence of the surrounding context should be hearable. The stable phase of a vowel is
characterized by parallel formants (observable in the spectrogram) and a constant waveform
(observable in the time signal). VA (Anfang) marks the beginning of the vowel and VE
(Ende) the ending. The indixes “u” are used for unstressed (unbetont), “b” for stressed
(betont) vowels.
3. Stress (.stress): Perceptual prominence for each syllable: no prominence (0), most prominent
(1), prominent but not most (2).
4. Orthographic transcription (.trans). Each utterance was annotated on the syllabic level by two
trained transcribers.
5. Cover Symbol Transcription (.phones–cover). Manner and place of articulation describes the
consonantal structure: L=labial, A=Alveolar, V=velar, G=glottal, O=other; P=plosive, N=nasal,
F=fricative, G=glide O=other. Tongue height and manner describes the vocalic structure:
H=high, M=medium, T=low, tief. V=front, vorne, Z=central, zentrale, H=back, ?
6. Narrow transcription in XSAMPA (.phones).
7. Syllable structure coding (.sylstr): Syllable structure, words and syllable marker.
 Aw = beginning of utterance (Anfang)
 Ew = end of utterance (Ende)
 S = syllable border
 1 = precanonical babbling
 2 = reduplicated babbling, CV-syllable with real consonants ([mamma], [dae])
 3 = one or more CV-syllables with real consonant with different articulation place
and/or manner ([min], [dajaep])
 W = meaningful utterance
 M/V = utterance of mother/father
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
20
Japanese-Ota
Project Name:
Japanese-Ota
Investigator(s):
Mitsuhiko Ota
Contact:
[email protected]
Location:
Washington, DC
Number of Participants:
3
Nature of study:
naturalistic, longitudinal
Media Type (if any):
Audio cassette tapes, and video super 8 cassette tapes
Citation Information:
Ota, Mitsuhiko (2003). The development of prosodic structure in early words: Continuity,
divergence and change. Amsterdam/ Philadelphia: John Benjamins.
Participant Name
Date of Birth
Age range
Hiromi
Kenta
Takeru
1982-12-09
1982-05-15
1982-03-26
1;00.21 – 2;00.07
1;05.18 - 2;06.07
1;04.23 – 2;00.19
Number of
Sessions
22
26
15
Sex
F
M
M
The participants were from a Japanese community in the Washington, D.C. area. All children
were at the beginning of the one-word period and living in a family that spoke only Japanese
among themselves. The children had varying degrees of exposure to English.
The data were collected through a series of recordings, in two to three week intervals, of
spontaneous utterances rather than by controlled elicitation techniques. Audio recordings were
made on standard cassette tapes with a Sony WM-D6C analog tape recorder and a Sony ECM909A microphone. There were also video-recordings made on super 8 cassette tapes with a
Samsung SCX915 camcorder
The data were transcribed using a stantard Romanization system for Japanese and IPA. The
target adult form was determined from the context of the utterance. Indeterminable utterances
were removed from the analysis. Spontaneous speech and imitations of the adult’s utterances
were coded for. In the case of imitations, the full adult utterance is provided.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by
the above reference.
21
Japanese-Stanford
Incomplete
Project Name:
Japanese-Stanford
Investigator(s):
Marilyn Vihman
Contact:
[email protected]
Location:
Stanford, CA
Number of Participants:
6
Nature of study:
naturalistic, longitudinal
Media Type (if any):
audio cassettes and/or reel-to-reel/VHS video
Citation Information:
Vihman, M. M., Macken, M.A., Miller, R., Simmons, H. & Miller, J. (1985). From babbling to
speech: a reassessment of the continuity issue. Language, 61, 395-443.
Vihman, M. M., C.A. Ferguson & M. Elbert. (1986). Phonological development from babbling to
speech: Common tenden-cies and individual differences. Applied Psycholinguistics, 7, 340.
Boysson-Bardies, B. de & Vihman, M. M. (1991). Adaptation to language: Evidence from
babbling and first words in four languages. Language, 67, 297-319.
Vihman, M. M., Kay, E., Boysson-Bardies, B. de, Durand, C. & Sundberg, U. (1994). External
sources of individual differences? A cross-linguistic analysis of the phonetics of mothers'
speech to one-year-old children. Developmental Psychology, 30, 652-663.
Vihman, M. M., DePaolis, R.A. & Davis, B.L. (1998). Is there a “trochaic bias” in early word
learning? Evidence from English and French. Child Development, 69, 933-947.
Participant Name
Date of Birth
e
TBA
Number of
Sessions
2
TBA
0;10.14 – 1;04.07
4
F
1;01.14 – 1;06.07
5
M
F
emi
har
TBA
Age range
Sex
TBA
kaz
TBA
1;00.17 – 1;03.15
5
ken
TBA
1;00.04 – 1;05.26
6
M
tar
TBA
1;02.07 – 1;11.02
1
M
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
22
PAIDOLOGOS-Cantonese/English/Greek/Japanese
Incomplete
Project Name:
PAIDOLOGOS
Investigator(s):
TBA
Contact:
TBA
Location:
TBA
Number of Participants:
TBA
Nature of study:
naturalistic/experimental, longitudinal/cross-sectional
Media Type (if any):
TBA
Citation Information: TBA
Participant Name
Date of Birth
Age range
TBA
TBA
TBA
Number of
Sessions
TBA
Sex
TBA
The characters for the target vowel in the WorldBet and the WorldBetb.seg tiers differ because
the representation in the WorldBet tier denotes the vowel phoneme whereas the representation in
the WorldBet.seg tier designates the (potentially larger) "vowel class" that we set up to be able to
compare the effects of coarticulatory context on initial consonant accuracy across the four
languages.
Here are the five groups that we set up for English:
"I" includes both long tense [i:] and short lax [I] (so, e.g., "ti" is elicited in <teepee> and <teacher>
and also in <tickle>)
"E" includes both long tense [eI] and short lax [E] (e.g., "te" is elicited in <taste> and <tail> and
also in <tent>)
"A" includes both (or all three of) low back [A] (and low-mid back rounded [O], which does not
contrast with [A] in most Ohio dialects) and low-mid [^] (e.g., "ta" is elicited in <taco> and <tall>
and <tongue>)
"O" is just the long tense [oU]
"U" is both [u:] and [U]
Now, for the vowel classes, in Cantonese we had seven, because we wanted to be able to
compare the contrastively front rounded vowels of Cantonese to the analogous classes in French
(e.g., to contrast the effect of "u" versus "y" on accuracy of the preceding dental consonants,
where the high back vowel is phonotactically illegal in Cantonese). Also, in Cantonese, the tones
have spaces after them. The following are the seven vowel classes that were set up for
Cantonese:
"i" includes long [i:] both as a simple vowel and as the nucleus of a diphthongs, so, e.g., "khi" is
elicited in the morpheme "khi:u23" as well as in morphemes "khi:n21" and "khi:t3" (I'm writing just
23
the first zi in all three words)
"e" includes both long [E:] and short [e] and the latter both as a simple vowel and as the nucleus
of the diphthong [ei]
"a" includes both the short low central vowel "ax" (in WorldBet" and the long low back rounded
vowel that should be "5:" in WorldBet, but which we're writing as the easier to read "a:", and it
includes both of these vowels as both as simple vowels and as the nuclei of diphthongs
"o" includes both the short high-mid [o] (as the simple vowel or as the nucleus of the diphthong
[ou]) and the long low-mid [>:] (This ">" is how I should have designated the analogous vowel in
English which is included in the English "a" class instead, since in most Ohio dialects it doesn't
contrast with [A], so here's a place where it was impossible to set up optimal vowel class
comparisons.)
"u" includes long [u:] both as a simple vowel and as the nucleus of [u:i]
"oe" includes both the long low-mid rounded front vowel that is called "8" in WorldBet and the
mid-central rounded vowel (IPA "squashed theta") as the nucleus of the non-back rounded
diphthong "oxy"
"y" includes only the long high front rounded vowel [y:]
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
Portuguese-Lisbon
Project Name:
EP_Mono
Investigator(s):
Susana Correia
Teresa da Costa
Contact:
[email protected]
[email protected]
Location:
Lisbon and area, Portugal
Number of Participants:
5
Nature of study:
naturalistic, spontaneous, longitudinal, monolingual
Media Type (if any):
QuickTime movie files (.mov)
Citation Information:
Correia, S. (2010). The Acquisition of Primary Word Stress in European Portuguese. Ph.D.
Thesis. University of Lisbon.
Costa, T. (2010). The Acquisition of the Consonantal System in European Portuguese. Focus on
Place and Manner Features. Ph.D. Thesis. University of Lisbon.
Participant Name
Date of Birth
Age range
Number of
Sessions
Sex
Clara
01-Jun-2005
0;11 - 1;10
12
F
Inês
19-Nov-1992
0;11 - 4;2
30
F
24
Joana
22-Jun-1995
0;11 - 4;10
33
F
João
02-May-2004
1;0-2;0
22
M
Luma
24-Oct-2003
0;11-2;6
55
F
This corpus contains data from 5 European Portuguese monolingual children during 134
recording sessions. The children’s speech was recorded every other week in the case of João
and Luma, and monthly in the case of InInês, Clara and Joana. The data were recorded at the
children’s homes, during spontaneous speech. The mother and/or caretaker and the researcher
were present during the recording sessions.
The data from Clara, João and Luma were recorded in a Canon MV210 digital video camcorder a
nd videotaped in analogue format (Hi8 cassettes), between 2004 and 2007. Inês and Joana were
recorded on a Sony Handycam Video 8, AF Hi‐ Fi Stereo, during the 1990's. In all cases, the anal
ogue format was imported to digital format using iMovie© software, and compressed for 320x240
(.mov) format.
The data were manually entered into the Phon application. Orthographic and phonetic
transcriptions were made of the target and children's actual forms. The sessions were transcribed
and exchanged between two transcribers for revisions. Each transcriber was responsible for
transcribing 50% of the children's speech and revising the other 50%. In dubious cases, a third
researcher carried out a blind transcription and, in the case of persistent indecision, that utterance
was marked with an asterisk (*) as unintelligible.
All of the specific research questions formulated within this project were related to the
development of segmental properties (place and manner features) in Portuguese children. The
analytic tools available through the Phon application enabled the researcher to extract information
on the general order of acquisition of the EP consonantal inventory and the segmental shape of
early words. The analytic features used in the study were: (i) searching and analyzing
consonantal place and manner features in onset position; (ii) discriminating between homorganic
and non-homorganic feature combinations at the word-level; (iii) searching and analyzing
consonantal place and manner features, discriminating between stressed and unstressed
syllables; (iv) searching/analyzing consonant harmony and metathesis in children's speech.
The research questions underlying this project deal with the identification of and order of
acquisition for primary word stress and the study of the early shape of prosodic words in
Portuguese children. The analytic features used in this study were: (i) searching for and analyzing
the number of words uttered by the children; (ii) searching for and analyzing the number of
syllables within a word; (iii) searching for and analyzing the syllable structure in different word
positions; (iv) searching for and analyzing stress patterns.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
25
Portuguese-Freitas
Project Name:
Portuguese-Freitas
Investigator(s):
Maria João Freitas
Contact:
[email protected]
Location:
Lisbon area, PT
Number of Participants:
7
Nature of study:
naturalistic, longitudinal & cross-sectional samples
Media Type (if any):
Video with Hi-Fi stereo
Citation Information:
Freitas, M. J. (1997). Aquisição da Estrutura Silábica do Português Europeu. Ma, Universidade
de Lisboa.
Participant Name
Date of Birth
Age range
Number of
Sessions
Sex
Inês
João
Laura
Luís
Marta
Pedro
Raquel
1992-11-19
1992-03-29
1991-03-20
1991-12-23
1992-11-03
1991-03-27
1991-06-01
0;11.13 – 1;10.29
0;10.02 – 2;08.27
0;02.29 – 3;03.10
1;09.29 – 2;11.02
1;02.00 – 2;02.17
2;07.00 – 3;07.24
1;10.02 – 2;10.08
12
23
12
12
12
12
10
F
M
F
M
F
M
F
The recordings for this project were made in the child’s home, normally in the child’s bedroom.
Naturalistic data were collected both longitudinally and cross-sectionally. Each child was
videotaped for a period of one year longitudinally (João was videotaped for a period of 2 years).
The age of recordings for all sessions were from 0;10 onward and 3;7 onward.
Each recording session lasted from 30 to 60 minutes. The recordings were made using a Sony
Handycam video 8, AF Hi-FI Stereo. Digital recordings will be made available through Centro de
Linguística da Universidade de Lisboa by December, 2010.
The data collection was supported by Fundação para Ciência e a Tecnologia, research project
PCSH/C/LIN/524/93. The data transcription was first compiled within the CHILDPHON database,
developed at the Max-Planch Institut for Psycholinguistics in 1990 and first used by Fikkert (1994)
and Levelt (1994).
All sessions from João, Inês and Marta were transcribed in full. All sessions from Luís, Raquel,
Laura and Pedro were partially transcribed from the 10 minute to the 25 minute mark in each
session. Transcriptions were performed by a native speaker of European Portuguese. All
problematic transcriptions were noted and listened to by at least two other judges. In case of
remaining doubts on transcription, the utterance was not included in the database.
Due to the goals of the original research project, the corpus represents only isolated words or
sequences containing external sandhi phenomenon.
26
See Costa 2010 and Correia 2010 for full access to the Inês corpus, and for documentation of
early stress patterns. Transcription of stress in reduplicated forms was not revised in this corpus.
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
Romanian-Kern
Project Name:
Romanian-Kern
Investigator(s):
Sophie Kern
Contact:
[email protected]
Location:
Bucharest, RO
Number of Participants:
4
Nature of study:
naturalistic, longitudinal, monolingual
Media Type (if any):
digital video (.mov, .mp4)
Citation Information:
Kern, S, B.L. Davis, & I. Zink. (2009). From babbling to first words in four languages: Common
trends, cross language and individual differences. In Francesco d’Errico & Jean-Marie
Hombert (eds) Becoming eloquent: Advances in the Emergence of language, human
cognition and modern culture. John Benjamins' Publishing Company.
Kern, S. & B.L. Davis .(2009). Emergent complexity in early vocal acquisition: Cross-linguistic
comparisons of canonical babbling. In F. Pellegrino, E. Marsico, I. Chitoran & C. Coupé
(eds), Approaches to phonological complexity. Berlin, Mouton de Gruyter.
Kern, S. (2005). De l'universalité et des spécificités du développement langagier précoce. In
Hombert, J.M. (ed) Aux origines du langage et des langues. Fayard
Davis, B., S. Kern, A. Vilain & C. Lalevée (2008). Des babils à Babel: les premiers pas de la
parole. Revue Française de Linguistique Appliquée. 13(2), 81-91.
Participant Name
Date of Birth
Age range
Alice
2000-11-08
0;07.04 – 2;00.04
Number of
Sessions
31
Anna
2000-11-22
0;08.11 – 1;02.10
10
F
Matei
2000-11-09
0;08.10 – 2;01.11
34
M
Vlad
2000-09-08
0;08.28 - 2;00.27
28
M
Sex
F
Five types of data were collected. First, one hour of spontaneous vocalization data was audio and
video recorded every two weeks from 8 months of age through 25 months of age. Recording took
place in children’s homes. The parents were told to follow their normal types of activities with their
child. Second, minimally 1,000 dictionary entries from the ambient language employed by the
parents of each child participant were analyzed for comparison with the child data for that
language. Parental reports were administered using adaptations of the MacArthur Development
Inventories (Fenson et al., 1993) respectively elaborated for Dutch for French participants.
Mothers filled out the questionnaire once in a month. For the remaining languages there is no
27
adaptation yet, but one could imagine using the spontaneous data to elaborate the same
instrument. An object manipulation categorization task was administered every two months. This
task was conceived to evaluate the children’s spontaneous nonverbal categorization abilities.
Several toys, which were consistent across the language groups served as stimuli. Each task
involved a contrast of objects from two different categories (animal, means of locomotion,
furniture).
Children were developing normally by community standards as well as reports from parents and
physicians regarding developmental milestones were observed in the normal daily environment.
These children were becoming monolingual speakers of French, Romanian and Tunisian Arabic.
Languages were chosen to include different language families as well as for the contrasts they
show in phonetic and phonological features of interest for early language development. These
characteristics include word length, syllabic types as well as phonemic inventory diversity
One hour of spontaneous vocalization data was audio and video recorded every two weeks from
8 months of age through 25 months of age. Recording took place in children’s homes. The
parents were told to follow their normal types of activities with their child. After collection, data
was phonetically transcribed using the International Phonetic Alphabet (IPA). Broad phonetic
transcriptions were used, supplemented by some diacritics (mainly for palatalized,
pharyngealized, nasalized sounds and duration of sounds). Tokens considered as single
utterance strings were bounded by one second of silence, noise or adult speech.
This work was supported by the EUROCORE Program “The Origin of Man, Language and
Languages” (OMLL) and the French CNRS program “Origine de l’Homme, du Langage et des
Langues” (OHLL).
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
Spanish-PAIDUS
Project Name:
Spanish-PAIDUS
Investigator(s):
Antonio Maldonado
Conxita Lleó
Contact:
Universidad Autónoma de Madrid
[email protected]
Location:
Madrid, ES
Number of Participants:
5
Nature of study:
naturalistic, longitudinal/cross-sectional, babbling, first words, informal
testing
Media Type (if any):
Wave files
Citation Information:
Lleó, Conxita, Christliebe El Mogharbel & Michael Prinz (1994). Babbling und FrühwortProduktion im Deutschen und Spanischen. Linguistische Berichte 151, 191-217.
Lleó, Conxita & Michael Prinz (1996). Consonant clusters in child phonology and the directionality
of syllable structure assignment. Journal of Child Language 23, 31-56.
Lleó, Conxita, Michael Prinz, Christliebe El Mogharbel & Antonio Maldonado (1996). Early
phonological acquisition of German and Spanish: A reinterpretation of the continuity
28
issue within the Principles and Parameters model. In C.E. Johnson & J.H.V.
Gilbert (eds.), Children's Language, Volume 9, 11-31. Hillsdale: Lawrence Erlbaum
Associates.
Lleó, Conxita (2001). The Transition From Prenominal Fillers to Articles in Spanish and German.
Proceedings from the 8th International Congress of the International Association for the
Study of Child Language, 713-737. Donostia (Spain), Juli 1999.
Kehoe, Margaret & Conxita Lleó (2003a). The acquisition of syllable types in monolingual
and bilingual German and Spanish children. In B. Beachley, A. Brown, & F. Conlin (eds.),
BUCLD 27 Proceedings, 402-413. Somerville, Ma: Cascadilla Press.
Kehoe, Margaret & Conxita Lleó (2003b). The acquisition of nuclei: a longitudinal analysis
of phonological vowel length in three German-speaking children. Journal of Child
Language 30 (3), 527-556.
Kehoe, Margaret & Conxita Lleó (2003c). A Phonological Analysis of Schwa in German
First Language Acquisition. Canadian Journal of Linguistics 48 (3/4), 289-327.
Lleó, Conxita, Imme Kuchenbrandt, Margaret Kehoe & Cristina Trujillo (2003). Syllable
final consonants in Spanish and German monolingual and bilingual acquisition. In
N. Müller (ed.), (In)vulnerable Domains in Multilingualism, 191-220. Amsterdam,
Philadelphia: John Benjamins.
Kehoe, Margaret, Conxita Lleó & Margaret Rakow (2004). Voice Onset Time in Bilingual
German-Spanish Children. Bilingualism: Language and Cognition 7, 71-88.
Lleó, Conxita, Martin Rakow & Margaret Kehoe (2004). Acquisition of language-specific pitch
accent by Spanish and German monolingual and bilingual children. In T. Face
(ed.), Laboratory Approaches to Spanish Phonology, 3-27. Berlin, New York: Mouton.
Lleó, Conxita (2006). The Acquisition of Prosodic Word Structures in Spanish by Monolingual and
Spanish-German Bilingual Children. Language and Speech 49 (2), 207-231.
Lleó, Conxita & Javier Arias (2006). Foot, word and phrase constraints in first language
acquisition of Spanish stress. In Fernando Martínez-Gil & Sonia Colina (eds.), OptimalityTheoretic Studies in Spanish Phonology, 470-496. John Benjamins.
Lleó, Conxita & Martin Rakow (2006). The prosody of early two-word utterances by German and
Spanish monolingual and bilingual children. In C. Lleó (ed.), Interfaces in
Multilingualism: Acquisition and representation, 1-26. Amsterdam/Philadelphia: John
Benjamins.
Lleó, Conxita, Martin Rakow & Margaret Kehoe (2007). Acquiring rhythmically different
languages in a bilingual context. In Proceedings of the 16th International Congress of
Phonetic Sciences (ICPhS XVI), 1545-1548.
Lleó, C. & J. Arias (2009). The role of Weight-by-Position in the prosodic development of Spanish
and German. In Janet Grijzenhout & Baris Kabak (eds.), Phonological Domains:
Universals and Deviations, 221-247. Berlin, New York: Mouton de Gruyter.
Participant Name
Date of Birth
Age range
Ines
1991-02-23
0;09.27 – 1;06.12
Number of
Sessions
7
Sex
F
Jose
1991-06-17
0;09.16 – 2;06.12
21
Juan
1990-12-30
0;09.17 – 2;01.13
Maria
1990-10-17
0;09.15 – 2;07.17
11-12 (Bruder) M
16
F
Miguel
1990-08-13
0;10.04 – 3;01.28
22
M
M
The project BIDS/PAIDUS was carried out at the University of Hamburg (Department of Romance
Languages) with the support of the Deutsche Forschungsgemeinschaft under the direction of
Conxita Lleó. The project began in 1990 first as BIDS, with the goal of comparing the phonetic
characteristics of babbling in German and comparing it with babbling in Spanish. With this
purpose, a collaboration with Antonio Maldonado of the Universidad Autónoma de Madrid (Dept.
of Psychology) was started and we applied for a joint grant from “Acciones Integradas”, so that
29
the two teams could visit each other. The German data were collected in Hamburg, Germany,
and the Spanish data in Madrid, Spain.
All five German children had been born in Hamburg and lived there. Their parents were all
German and spoke the Northern variety of German. The children were visited from very early on,
before they began to produce their first words, as we wanted to study their babbling. Some
informal tests, in order to figure their (passive) vocabularies were conducted. After the children
began saying words, we continued recording them within the project PAIDUS, which had as its
goal to finding out how phonological parameters were possibly fixed for German as opposed to
Spanish. The German children were visited every 2 weeks at first, and later on once a month.
Recordings were made at the children’s homes by two research assistants of the project. At first,
the mother used to be present, but later on, we made most of the recordings with the child alone
interacting with one research assistant. We tried to record the child under spontaneous
conditions, playing, chatting, etc. We also brought some objects (little animals or toys) in order to
elicit certain words from all children, making their phonetic and phonological characteristics
comparable. The mother tried to keep a record of possible new words that were new in the child’s
vocabularies, and we tried to elicit them, as well. We also used books with pictures, which
belonged to the child and others that we brought. The recordings were audio recordings, and in
order not to miss the ambient setting and context, one of the researchers was writing down all
possible details, including what the child was doing, while the other assistant was interacting with
the child.
Both projects, BIDS and PAIDUS, were financed by the DFG with grant to Conxita Lleó, BIDS
from 1989 until 1991 and PAIDUS from 1991 until 1994. Research Assistants in the project BIDS
were Christliebe El Mogharbel and Kerstin Feuge. Research Assistants in the project PAIDUS
were Christliebe El Mogharbel and Michael Prinz.
Reserachers using the BIDS/PAIDUS data should send a copy of their publication based on the
data to the contributor of the data. In accordance with CHILDES rules, any use of data from this
corpus must be accompanied by at least one of the above references.
Swedish-Lacerda
Cornelia was born 23-SEP-2004. Ebba was born 02-SEP-2004. Teresa was born 17AUG-2004.
Session
c06
c08
c09
c10
c12
c13
e05
e06
e07
e08
Date
2005-OCT-10
2005-DEC-05
2006-JAN-13
2006-FEB-17
2006-DEC-18
2007-MAY-07
2005-OCT-10
2005-NOV-08
2006-JAN-16
2006-FEB-06
Age
1;0.17
1;3,02
1;3.21
1;4.25
2;2.25
2;7.14
1;1.08
1;2.06
1;4.14
1;5.04
30
e09
e10
t02
t03
t05
t06
t09
t10
2006-NOV-14
2007-MAY-29
2005-SEP-12
2005-OCT-04
2005-DEC-06
2006-JAN-09
2006-NOV-16
2007-MAY-09
2;2.12
2;8,27
1;0.26
1;1.17
1;3.18
1;4.23
2;2.30
2;8.22
Sound files are named for session. There were separate WaveSurfer transcript files
for the child, caregiver, and the investigator and these were merged to create the
CHILDES transcripts.
Swedish-Stanford
Incomplete
Project Name:
Swedish-Stanford
Investigator(s):
Marilyn Vihman
Contact:
[email protected]
Location:
Stockholm, SE
Number of Participants:
4 (or 5)
Nature of study:
naturalistic, longitudinal
Media Type (if any):
audio cassettes and/or reel-to-reel/VHS video
Citation Information:
Vihman, M. M., Macken, M.A., Miller, R., Simmons, H. & Miller, J. (1985). From babbling to
speech: a reassessment of the continuity issue. Language, 61, 395-443.
Vihman, M. M., C.A. Ferguson & M. Elbert. (1986). Phonological development from babbling to
speech: Common tenden-cies and individual differences. Applied Psycholinguistics, 7, 340.
Boysson-Bardies, B. de & Vihman, M. M. (1991). Adaptation to language: Evidence from
babbling and first words in four languages. Language, 67, 297-319.
Vihman, M. M., Kay, E., Boysson-Bardies, B. de, Durand, C. & Sundberg, U. (1994). External
sources of individual differences? A cross-linguistic analysis of the phonetics of mothers'
speech to one-year-old children. Developmental Psychology, 30, 652-663.
Vihman, M. M., DePaolis, R.A. & Davis, B.L. (1998). Is there a “trochaic bias” in early word
learning? Evidence from English and French. Child Development, 69, 933-947.
Participant Name
Date of Birth
Age range
cur
TBA
0;09.03 – 1;04.05
Number of
Sessions
TBA
Sex
M
31
did
TBA
0;09.16 – 1;04.18
3
M
F
han
TBA
0;09.15 – 1;05.10
3
lin
TBA
0;09.07 – 1;04.12
3
F
0.09.25 – 1;06.29
3
M
sti
TBA
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
Tunisian-Kern
Project Name:
Tunisian-Kern
Investigator(s):
Sophie Kern
Contact Information: Laboratoire Dynamique du Langage
[email protected]
Location:
Tunis, TN
Number of Participants:
4
Nature of study:
naturalistic, longitudinal, monolingual
Media Type (if any):
digital video (.mov, .mp4)
Citation Information:
Kern, S, B.L. Davis, & I. Zink. (2009). From babbling to first words in four languages: Common
trends, cross language and individual differences. In Francesco d’Errico & Jean-Marie
Hombert (eds) Becoming eloquent: Advances in the Emergence of language, human
cognition and modern culture. John Benjamins' Publishing Company.
Kern, S. & B.L. Davis .(2009). Emergent complexity in early vocal acquisition: Cross-linguistic
comparisons of canonical babbling. In F. Pellegrino, E. Marsico, I. Chitoran & C. Coupé
(eds), Approaches to phonological complexity. Berlin, Mouton de Gruyter.
Kern, S. (2005). De l'universalité et des spécificités du développement langagier précoce. In
Hombert, J.M. (ed) Aux origines du langage et des langues. Fayard
Davis, B., S. Kern, A. Vilain & C. Lalevée (2008). Des babils à Babel: les premiers pas de la
parole. Revue Française de Linguistique Appliquée. 13(2), 81-91.
Participant Name
Date of Birth
Age range
Fereyel
2002-04-09
0;08.09 – 2;00.21
Number of
Sessions
26
Iyed
Malek
Zaidaan
2002-01-30
2002-02-24
2002-01-24
0;08.22 - 2;00.00
0;10.11 – 2;00.04
0;04.15 – 2;00.10
30
24
31
Sex
F
F
M
F
Five types of data were collected. First, one hour of spontaneous vocalization data was audio and
video recorded every two weeks from 8 months of age through 25 months of age. Recording took
place in children’s homes. The parents were told to follow their normal types of activities with their
32
child. Second, minimally 1,000 dictionary entries from the ambient language employed by the
parents of each child participant were analyzed for comparison with the child data for that
language. Parental reports were administered using adaptations of the MacArthur Development
Inventories (Fenson et al., 1993) respectively elaborated for Dutch for French participants.
Mothers filled out the questionnaire once in a month. For the remaining languages there is no
adaptation yet, but one could imagine using the spontaneous data to elaborate the same
instrument. An object manipulation categorization task was administered every two months. This
task was conceived to evaluate the children’s spontaneous nonverbal categorization abilities.
Several toys, which were consistent across the language groups served as stimuli. Each task
involved a contrast of objects from two different categories (animal, means of locomotion,
furniture).
Children were developing normally by community standards as well as reports from parents and
physicians regarding developmental milestones were observed in the normal daily environment.
These children were becoming monolingual speakers of French, Romanian and Tunisian Arabic.
Languages were chosen to include different language families as well as for the contrasts they
show in phonetic and phonological features of interest for early language development. These
characteristics include word length, syllabic types as well as phonemic inventory diversity
One hour of spontaneous vocalization data was audio and video recorded every two weeks from
8 months of age through 25 months of age. Recording took place in children’s homes. The
parents were told to follow their normal types of activities with their child. After collection, data
was phonetically transcribed using the International Phonetic Alphabet (IPA). Broad phonetic
transcriptions were used, supplemented by some diacritics (mainly for palatalized,
pharyngealized, nasalized sounds and duration of sounds). Tokens considered as single
utterance strings were bounded by one second of silence, noise or adult speech.
This work was supported by the EUROCORE Program “The Origin of Man, Language and
Languages” (OMLL) and the French CNRS program “Origine de l’Homme, du Langage et des
Langues” (OHLL).
In accordance with CHILDES rules, any use of data from this corpus must be accompanied by at
least one of the above references.
_____________________________________________________________________________
33