CHILDES English CAPIL Corpus


Melissa Redford
Department of Linguistics
University of Oregon

Kristopher Kyle
Department of Linguistics
University of Oregon

Participants: 30, 30
Type of Study: narrative
Location: United States
Media type: audio
DOI: doi:10.21415/wsvw-0f88

Browsable transcripts

Download transcripts

Media folder

Citation information

Project Description

CAPIL (Child and Adult Phrasing, Inhalation, and Language) is an audio-linked corpus of nearly 18 hours of transcribed spontaneous, pause-delimited speech. It is unique in pairing utterance-by-utterance transcriptions of language with inhalation counts during the silent pauses between utterances. The corpus was created as part of the NSF-funded project ‘Pausing, Inhalation, and Language Structure’ (PILS; Award BCS-2315232), under PI Melissa (Lisa) Redford and co-PI Kristopher Kyle at the University of Oregon. Data collection was carried out under University of Oregon IRB protocol STUDY00000676. No restrictions beyond TalkBank's general Ground Rules apply to use of this corpus, except that we ask researchers using it to send us copies of any resulting publications.

Participants

The corpus includes data from 60 participants: 30 five-year-olds (16 girls, 14 boys, ages 5;1– 6;10) and 30 young adults (15 women, 12 men, 1 transgender man, 2 non-binary, ages 18;0–24;5), all speakers of a West Coast variety of American English. Roughly three- quarters of both groups identified as white only (23 of 30 children, 18 of 30 adults), with the remainder identifying as Asian, Hispanic or Latino, Black or African American, Native Hawaiian or Other Pacific Islander, multiracial, or some combination. Everyone passed a pure-tone hearing screening and scored in the typical range on the Expressive Vocabulary Test (EVT-3; Williams, 2019). Although no participants were in speech-language therapy at the time of the study, three children and four adults were reported to have experienced a speech delay, stutter, or past speech-language therapy; also, three children and four adults reported an ADHD diagnosis. Although four adults also self-reported asthma, spirometry confirmed typical pulmonary function for age in both groups of participants. Written consent was obtained from all adults and from each child's parent or guardian, with written assent from the children themselves. Per-participant age, sex, height, weight, and vocabulary score are given in the table at the end of this document.

Elicitation and Recording

Recording took place in the Spoken Language Research Laboratories at the University of Oregon in a sound-dampened experimental room. Participants were seated at a height- adjusted table facing a computer screen, with an experimenter to the side. Participants were asked to keep their hands placed on marked outlines to discourage movement during recording of speech breathing kinematics. Each participant narrated two short wordless Claymation episodes from the series Pingu, created by Otmar Gutmann and Erika Brueggemann, while watching a silenced replay after having watched them once through with sound ('Artist' = 'Pingu the Painter,' S3 E6; 'School' = 'Pingu and Pinga at the Kindergarten,' S2 E23, replaced by 'Skating' = 'Pingu's First Kiss,' S2 E5, for the first two participants), about 5 minutes of narration each. Each participant also produced expository speech in response to a descriptive prompt and an explanatory prompt (3–5 minutes each). Audio, video, and breathing kinematics were captured simultaneously: respiratory inductance plethysmography (RIP; Inductotrace, Ambulatory Monitoring Inc.) for the breathing signal, a directional microphone for kinematic-synchronized audio, a lavalier microphone feeding a separate high-quality recording (Zoom F1 Field Recorder, .WAV), and a side-mounted camera (.mp4) to monitor for movement that could distort the kinematic signal.

Transcript and Coding

Recordings were first auto-transcribed with timestamps, with text files then used to produce a preliminary pause-delimited TextGrid aligned to the audio signal in Praat (Boersma & Weenink, 2023–2026). Transcriptions and pause boundaries were then painstakingly corrected and additionally coded by hand, with Communication units (C-units; Loban, 1976) marked on a separate tier, following SALT (Heilmann & Miller, 2023) conventions. Breathing kinematics (recorded via RIP, not part of the deposited CHAT files) were used to identify inhalations during each pause; the deposited transcripts carry only the resulting count of separate inhalations detected during each pause (0 for none, 1, 2, and so on), not the underlying kinematic depth/amplitude data itself. In 96% of pauses this count is 0 or 1; a small remainder (about 4%) have 2 or more, up to 8 in longer pauses. Every TextGrid was second-checked by another member of the research team; a further pass specifically re- checked the kinematic (breathing) segmentation.

Project-Specific Codes

The original transcripts were done in an older version of CHAT and needed to be reformatted to pass Chatter. The transcripts is organized in breath groups with the speech on the *SPK line and the pauses and inhalations noted in the %cod line. Each *SPK line has a following %cod line which gives the duration of the pause in milliseconds. The number on that line indicates the number of inhalations during that period. This "0" indicates a pause with no inhalation and "1" indicates a pause with one inhalation.

Participant Table

Per-participant values for the 60 deposited participants (IDs match the deposited CHA filenames), scoped to the fields the consent forms authorize for public sharing. Height and weight are given in consistent units (feet/inches; pounds) rather than mixed units per cell. Height was measured directly; weights are self-/parent-reported estimates. A dash (–) indicates data that was not collected or could not be located, rather than a true zero or an inapplicable field.
Participant ID Group Age Sex Height (ft, in) Weight (lbs) Vocab (EVT Standard Score)
01-A_19FAdult19;0F5'6"145115
01-C_69MChild5;9M3'7"43110
02-A_22MAdult22;7M5'11"150104
02-C_66FChild5;6F4'0"50108
03-A_22MAdult22;7M6'2"180117-118
04-A_22FAdult22;8F5'0"135107
04-C_66FChild5;6F3'6"3694
05-A_19TMAdult19;5TM5'6"145117
06-A_21MAdult21;11M5'10"135110
06-C_67FChild5;7F3'9"4595
07-A_19FAdult19;6F5'5"150109
07-C_69MChild5;9M3'8"4095
08-C_69FChild5;9F99
09-A_20FAdult20;8F5'8"120104
09-C_65MChild5;5M4'0"55160
10-A_19FAdult19;3F5'7"180101
10-C_71FChild5;11F4'0"6598
11-C_72FChild6;0F96
12-A_18NBAdult18;0NB5'2"95121
13-A_20FAdult20;5F5'5"143108
13-C_66MChild5;6M3'4"44114
14-A_19FAdult19;10F5'4"140120
14-C_72FChild6;0F3'11"103
15-A_19FAdult19;6F5'8"150105
15-C_68FChild5;8F116
16-A_19MAdult19;9M5'8"155107
16-C_72FChild6;0F3'11"5599
17-A_24MAdult24;5M5'9"15879
17-C_69MChild5;9M3'2"58107
18-A_18MAdult18;4M5'9"135107
18-C_67MChild5;7M3'6"45120
19-C_61FChild5;1F3'6"3685
20-C_74FChild6;2F4'0"4095
21-A_22MAdult22;11M5'8"140119
21-C_66FChild5;6F4'0"5097
23-A_20FAdult20;9F5'3"120122
23-C_67MChild5;7M3'9"4293
24-A_21FAdult21;8F5'6"175123
24-C_64MChild5;4M4'1"48121
25-A_20FAdult20;10F5'8"160119
25-C_64MChild5;4M4'0"4099
26-A_22FAdult22;6F5'7"220105
26-C_80MChild6;8M4'6"6097
27-C_82FChild6;10F3'11"4692
28-A_21MAdult21;4M5'10"240119
29-A_18MNBAdult18;6M, NB5'11"130134
30-C_65MChild5;5M3'11"51108
31-C_65FChild5;5F3'11"4893
32-A_19FAdult19;8F5'3"148102
32-C_65FChild5;5F3'10"50118
33-A_20MAdult20;3M6'1"16696
33-C_78MChild6;6M4'4"65104
34-A_20MAdult20;8M6'0"135121
34-C_78FChild6;6F4'0"4894
35-A_20MAdult20;7M6'0"153101
35-C_65MChild5;5M3'6"36136
36-A_19FAdult19;8F5'4"150106
36-C_71MChild5;11M4'0"45118
37-A_19FAdult19;10F5'4"145100
38-A_19MAdult19;8M6'3"200106

Acknowledgments

We are grateful to the many undergraduate and graduate research assistants who helped build this corpus. Bess Frerichs deserves special thanks: she contributed to every aspect of the work, from writing scripts and transcribing/segmenting/coding to training new research assistants, organizing the work of the team, and leading authorship of the lab's process manual. We also thank, in alphabetical order by last name: Siri Chotechuang, Maya Darmawi-Hicks, Sofia James, Lukas Klotz, Ava Lindon, Ella MacIsaac, Estelle Roering, Sarah Shellow, and James Taylor. This list reflects those who contributed most substantially to the project; it is not exhaustive.

References

Boersma, P., & Weenink, D. (2023–2026). Praat: Doing phonetics by computer [Computer software]. https://www.praat.org/

Gutmann, O., & Brueggemann, E. (Creators). (1990–2000). Pingu [TV series]. Pingu Filmstudio; HIT Entertainment.

Heilmann, J., & Miller, J. F. (2023). Systematic analysis of language transcripts solutions: A tutorial. Perspectives of the ASHA Special Interest Groups, 8(1), 1–18.

Loban, W. (1976). Language development: Kindergarten through grade twelve (NCTE Committee on Research Report No. 18). National Council of Teachers of English.

Williams, K. T. (2019). Expressive Vocabulary Test (3rd ed.). NCS Pearson.