CHILDES English CAPIL Corpus
|
Melissa Redford
Department of Linguistics
University of Oregon
redford@uoregon.edu
|
|
Kristopher Kyle
Department of Linguistics
University of Oregon
kkyle2@uoregon.edu
|
| Participants: | 30, 30 |
| Type of Study: | narrative |
| Location: | United States |
| Media type: | audio |
| DOI: | doi:10.21415/wsvw-0f88 |
Browsable transcripts
Download transcripts
Media folder
Citation information
Redford, M., Frerichs, B., & Kyle, K. (2026). On the Relationship between Breathing and
Prosody in the CAPIL Corpus. In Proc. Speech Prosody 2026 (pp. 98–102).
https://doi.org/10.21437/SpeechProsody.2026-20 pdf here.
In accordance with TalkBank rules, any use of data from this corpus must be accompanied
by the above reference.
Project Description
CAPIL (Child and Adult Phrasing, Inhalation, and Language) is an audio-linked corpus of
nearly 18 hours of transcribed spontaneous, pause-delimited speech. It is unique in pairing
utterance-by-utterance transcriptions of language with inhalation counts during the silent
pauses between utterances. The corpus was created as part of the NSF-funded project
‘Pausing, Inhalation, and Language Structure’ (PILS; Award BCS-2315232), under PI
Melissa (Lisa) Redford and co-PI Kristopher Kyle at the University of Oregon. Data collection
was carried out under University of Oregon IRB protocol STUDY00000676. No restrictions
beyond TalkBank's general Ground Rules apply to use of this corpus, except that we ask
researchers using it to send us copies of any resulting publications.
Participants
The corpus includes data from 60 participants: 30 five-year-olds (16 girls, 14 boys, ages 5;1–
6;10) and 30 young adults (15 women, 12 men, 1 transgender man, 2 non-binary, ages
18;0–24;5), all speakers of a West Coast variety of American English. Roughly three-
quarters of both groups identified as white only (23 of 30 children, 18 of 30 adults), with the
remainder identifying as Asian, Hispanic or Latino, Black or African American, Native
Hawaiian or Other Pacific Islander, multiracial, or some combination. Everyone passed a
pure-tone hearing screening and scored in the typical range on the Expressive Vocabulary
Test (EVT-3; Williams, 2019). Although no participants were in speech-language therapy at
the time of the study, three children and four adults were reported to have experienced a
speech delay, stutter, or past speech-language therapy; also, three children and four adults
reported an ADHD diagnosis. Although four adults also self-reported asthma, spirometry
confirmed typical pulmonary function for age in both groups of participants. Written consent
was obtained from all adults and from each child's parent or guardian, with written assent
from the children themselves. Per-participant age, sex, height, weight, and vocabulary score
are given in the table at the end of this document.
Elicitation and Recording
Recording took place in the Spoken Language Research Laboratories at the University of
Oregon in a sound-dampened experimental room. Participants were seated at a height-
adjusted table facing a computer screen, with an experimenter to the side. Participants were
asked to keep their hands placed on marked outlines to discourage movement during
recording of speech breathing kinematics. Each participant narrated two short wordless
Claymation episodes from the series Pingu, created by Otmar Gutmann and Erika
Brueggemann, while watching a silenced replay after having watched them once through
with sound ('Artist' = 'Pingu the Painter,' S3 E6; 'School' = 'Pingu and Pinga at the
Kindergarten,' S2 E23, replaced by 'Skating' = 'Pingu's First Kiss,' S2 E5, for the first two
participants), about 5 minutes of narration each. Each participant also produced expository
speech in response to a descriptive prompt and an explanatory prompt (3–5 minutes each).
Audio, video, and breathing kinematics were captured simultaneously: respiratory
inductance plethysmography (RIP; Inductotrace, Ambulatory Monitoring Inc.) for the
breathing signal, a directional microphone for kinematic-synchronized audio, a lavalier
microphone feeding a separate high-quality recording (Zoom F1 Field Recorder, .WAV), and
a side-mounted camera (.mp4) to monitor for movement that could distort the kinematic
signal.
Transcript and Coding
Recordings were first auto-transcribed with timestamps, with text files then used to produce
a preliminary pause-delimited TextGrid aligned to the audio signal in Praat (Boersma &
Weenink, 2023–2026). Transcriptions and pause boundaries were then painstakingly
corrected and additionally coded by hand, with Communication units (C-units; Loban, 1976)
marked on a separate tier, following SALT (Heilmann & Miller, 2023) conventions. Breathing
kinematics (recorded via RIP, not part of the deposited CHAT files) were used to identify
inhalations during each pause; the deposited transcripts carry only the resulting count of
separate inhalations detected during each pause (0 for none, 1, 2, and so on), not the
underlying kinematic depth/amplitude data itself. In 96% of pauses this count is 0 or 1; a
small remainder (about 4%) have 2 or more, up to 8 in longer pauses. Every TextGrid was
second-checked by another member of the research team; a further pass specifically re-
checked the kinematic (breathing) segmentation.
Project-Specific Codes
The original transcripts were done in an older version of CHAT and needed to be reformatted
to pass Chatter. The transcripts is organized in breath groups with the speech on the *SPK
line and the pauses and inhalations noted in the %cod line. Each *SPK line has a following
%cod line which gives the duration of the pause in milliseconds. The number on that line indicates
the number of inhalations during that period. This "0" indicates a pause with no inhalation
and "1" indicates a pause with one inhalation.
- C-units are marked with &-cu
- [% cont] marks that an utterance continues the same C-unit/sentence across an
intervening pause
- within-word pause breaks are marked by showing material at the break on the first line
at the end of the word and then at the beginning of the word on the next line.
Participant Table
Per-participant values for the 60 deposited participants (IDs match the deposited CHA
filenames), scoped to the fields the consent forms authorize for public sharing. Height and
weight are given in consistent units (feet/inches; pounds) rather than mixed units per cell.
Height was measured directly; weights are self-/parent-reported estimates. A dash (–)
indicates data that was not collected or could not be located, rather than a true zero or an
inapplicable field.
| Participant ID |
Group |
Age |
Sex |
Height (ft, in) |
Weight (lbs) |
Vocab (EVT Standard Score) |
| 01-A_19F | Adult | 19;0 | F | 5'6" | 145 | 115 |
| 01-C_69M | Child | 5;9 | M | 3'7" | 43 | 110 |
| 02-A_22M | Adult | 22;7 | M | 5'11" | 150 | 104 |
| 02-C_66F | Child | 5;6 | F | 4'0" | 50 | 108 |
| 03-A_22M | Adult | 22;7 | M | 6'2" | 180 | 117-118 |
| 04-A_22F | Adult | 22;8 | F | 5'0" | 135 | 107 |
| 04-C_66F | Child | 5;6 | F | 3'6" | 36 | 94 |
| 05-A_19TM | Adult | 19;5 | TM | 5'6" | 145 | 117 |
| 06-A_21M | Adult | 21;11 | M | 5'10" | 135 | 110 |
| 06-C_67F | Child | 5;7 | F | 3'9" | 45 | 95 |
| 07-A_19F | Adult | 19;6 | F | 5'5" | 150 | 109 |
| 07-C_69M | Child | 5;9 | M | 3'8" | 40 | 95 |
| 08-C_69F | Child | 5;9 | F | – | – | 99 |
| 09-A_20F | Adult | 20;8 | F | 5'8" | 120 | 104 |
| 09-C_65M | Child | 5;5 | M | 4'0" | 55 | 160 |
| 10-A_19F | Adult | 19;3 | F | 5'7" | 180 | 101 |
| 10-C_71F | Child | 5;11 | F | 4'0" | 65 | 98 |
| 11-C_72F | Child | 6;0 | F | – | – | 96 |
| 12-A_18NB | Adult | 18;0 | NB | 5'2" | 95 | 121 |
| 13-A_20F | Adult | 20;5 | F | 5'5" | 143 | 108 |
| 13-C_66M | Child | 5;6 | M | 3'4" | 44 | 114 |
| 14-A_19F | Adult | 19;10 | F | 5'4" | 140 | 120 |
| 14-C_72F | Child | 6;0 | F | 3'11" | – | 103 |
| 15-A_19F | Adult | 19;6 | F | 5'8" | 150 | 105 |
| 15-C_68F | Child | 5;8 | F | – | – | 116 |
| 16-A_19M | Adult | 19;9 | M | 5'8" | 155 | 107 |
| 16-C_72F | Child | 6;0 | F | 3'11" | 55 | 99 |
| 17-A_24M | Adult | 24;5 | M | 5'9" | 158 | 79 |
| 17-C_69M | Child | 5;9 | M | 3'2" | 58 | 107 |
| 18-A_18M | Adult | 18;4 | M | 5'9" | 135 | 107 |
| 18-C_67M | Child | 5;7 | M | 3'6" | 45 | 120 |
| 19-C_61F | Child | 5;1 | F | 3'6" | 36 | 85 |
| 20-C_74F | Child | 6;2 | F | 4'0" | 40 | 95 |
| 21-A_22M | Adult | 22;11 | M | 5'8" | 140 | 119 |
| 21-C_66F | Child | 5;6 | F | 4'0" | 50 | 97 |
| 23-A_20F | Adult | 20;9 | F | 5'3" | 120 | 122 |
| 23-C_67M | Child | 5;7 | M | 3'9" | 42 | 93 |
| 24-A_21F | Adult | 21;8 | F | 5'6" | 175 | 123 |
| 24-C_64M | Child | 5;4 | M | 4'1" | 48 | 121 |
| 25-A_20F | Adult | 20;10 | F | 5'8" | 160 | 119 |
| 25-C_64M | Child | 5;4 | M | 4'0" | 40 | 99 |
| 26-A_22F | Adult | 22;6 | F | 5'7" | 220 | 105 |
| 26-C_80M | Child | 6;8 | M | 4'6" | 60 | 97 |
| 27-C_82F | Child | 6;10 | F | 3'11" | 46 | 92 |
| 28-A_21M | Adult | 21;4 | M | 5'10" | 240 | 119 |
| 29-A_18MNB | Adult | 18;6 | M, NB | 5'11" | 130 | 134 |
| 30-C_65M | Child | 5;5 | M | 3'11" | 51 | 108 |
| 31-C_65F | Child | 5;5 | F | 3'11" | 48 | 93 |
| 32-A_19F | Adult | 19;8 | F | 5'3" | 148 | 102 |
| 32-C_65F | Child | 5;5 | F | 3'10" | 50 | 118 |
| 33-A_20M | Adult | 20;3 | M | 6'1" | 166 | 96 |
| 33-C_78M | Child | 6;6 | M | 4'4" | 65 | 104 |
| 34-A_20M | Adult | 20;8 | M | 6'0" | 135 | 121 |
| 34-C_78F | Child | 6;6 | F | 4'0" | 48 | 94 |
| 35-A_20M | Adult | 20;7 | M | 6'0" | 153 | 101 |
| 35-C_65M | Child | 5;5 | M | 3'6" | 36 | 136 |
| 36-A_19F | Adult | 19;8 | F | 5'4" | 150 | 106 |
| 36-C_71M | Child | 5;11 | M | 4'0" | 45 | 118 |
| 37-A_19F | Adult | 19;10 | F | 5'4" | 145 | 100 |
| 38-A_19M | Adult | 19;8 | M | 6'3" | 200 | 106 |
Acknowledgments
We are grateful to the many undergraduate and graduate research assistants who helped
build this corpus. Bess Frerichs deserves special thanks: she contributed to every aspect of
the work, from writing scripts and transcribing/segmenting/coding to training new research
assistants, organizing the work of the team, and leading authorship of the lab's process
manual. We also thank, in alphabetical order by last name: Siri Chotechuang, Maya
Darmawi-Hicks, Sofia James, Lukas Klotz, Ava Lindon, Ella MacIsaac, Estelle Roering,
Sarah Shellow, and James Taylor. This list reflects those who contributed most substantially
to the project; it is not exhaustive.
References
Boersma, P., & Weenink, D. (2023–2026). Praat: Doing phonetics by computer [Computer
software]. https://www.praat.org/
Gutmann, O., & Brueggemann, E. (Creators). (1990–2000). Pingu [TV series]. Pingu
Filmstudio; HIT Entertainment.
Heilmann, J., & Miller, J. F. (2023). Systematic analysis of language transcripts solutions: A
tutorial. Perspectives of the ASHA Special Interest Groups, 8(1), 1–18.
Loban, W. (1976). Language development: Kindergarten through grade twelve (NCTE
Committee on Research Report No. 18). National Council of Teachers of English.
Williams, K. T. (2019). Expressive Vocabulary Test (3rd ed.). NCS Pearson.