|
Haolong Zheng University of Illinois |
|
Mark Hasegawa-Johnson University of Illinois |
Zheng, H., Hu, Y., Liang, X., Sunder, V., Liu, D., Xiong, J., Thomas, S., Kingsbury, B., Wu, Z., & Hasegawa-Johnson, M. A. (2026). CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling. arXiv preprint arXiv:2607.03670. https://arxiv.org/abs/2607.03670 with PDF available here.
When using this database, please cite the above paper.
CHILDES-Aligned is a corpus of **501,880 transcript-paired English child-speech clips (413.3 hours)** with corrected utterance-level timestamps, derived from CHILDES. Utterance boundaries were recovered with BEACON (Boundary Estimation via Alignment CONsensus): word-level timestamps from four ASR systems (WhisperX, Parakeet, Canary, Qwen3) are aligned to the original human CHAT transcripts and fused by consensus voting, without altering the transcripts. Adjacent target-child utterances separated by at most 1 s are merged into single continuous clips.
All clips are target-child (CHI) speech from 49 English CHILDES corpora (child ages 0;6–14;3), cut as 16 kHz mono WAV.
There are two versions of the database.
Metadata per clip (CSV and JSONL provided) provide: relative audio path, transcript (text; the ASR version also carries text_original raw CHAT), the recovered span on the original recording (start_sec, end_sec), duration, recording_id, source corpus, and child_age (CHILDES years;months.days notation, e.g. 2;00.29 = 2 years, 0 months, 29 days), child_sex, child_group. corpora.tsv lists all 49 source corpora with PIDs and clip/hour counts.
Downloads
Copyright is CC BY-NC-SA 4.0 (attribution, non-commercial, share-alike), inherited from CHILDES, under the TalkBank ground rules.