We expect that researchers will contribute corpora constructed with TalkBank programs and tools. We strongly recommend
that projects collecting new data make use of the Batchalign pipeline and the CLAN editor. Batchalign is available from
https://github.com/talkbank. When applied to newly collected audio or video, it uses ASR to create transcripts in the format required
for inclusion in TalkBank. If you are unable to use Batchalign, you can send us media for processing
and we will send the results back to you for cleanup.
Cleanup of any mistakes can then be done using the CLAN editor, as describe in
this sheet.
It is the obligation of TalkBank users to assure that contributions are properly acknowledged.
It is the responsibility of TalkBank to make sure that data are correctly registered and properly accessible.
To contribute a new data set:
First, please send an email message to me (Brian MacWhinney) at macw@cmu.edu describing your contribution.
For FluencyBank contributions, please write to both Brian MacWhinney and Nan Bernstein Ratner (nratner@umd.edu).
Additional instructions that are specific to HomeBank are here.
Additional instructions that are specific to PsychosisBank are here.
If you were not able to obtain informed consent for sharing of media, we still need to receive a copy of the media for
processing and offline archiving.
If you have created transcripts in current CHAT format. You can send us your transcripts and media
together, as described below for WeTransfer.
If you need help in creating transcripts, you can send us your media for us to process.
TalkBank uses a system which requires that each transcript align with
only one media file and that the names of the transcript file and the media file be the same (ignoring the extensions).
For example, the file 020456.cha must have a matching 020456.mp4 (or .wav or .mp3) media file. In addition, the @Media
line in the *.cha file should use the name of the media which matches the name of the transcript.
Please use names for your files that are as short as possible with nothing more than the code name for the participant,
session number, and for children, the age in the format YYMMDD. Information already provided in folder names and the @ID lines
does not need to be duplicated in file names.
Please send us the information needed to create a corpus web page, such as
this one .
Just send this in MS-Word or other text format, rather than in HTML.
Your corpus page should provide corpus and project documentation. The guidelines for information to include
are given in chapter 6 of the CHAT manual, which is copied
here.
Also, please complete this contribution form, scan it, and
send us the scan and documentation either through WeTransfer or as an email attachment to macw@cmu.edu.
Once you have organized your contribution materials, send an email to Brian MacWhinney at macw@cmu.edu
and he will send you a link for transferring files using WeTransfer.com
After everything is in the database, we will create a webpage for your corpus along with a DOI number,
and we will eventually announce the addition of the new corpus to the Info-CHILDES list or the AphasiaBank list.
If you need to cite your corpus, you can use the format in this example, as suggested by the APA manual:
Bernstein Ratner, Nan. (1988). Bernstein Ratner Corpus. (data file) Retrieved from https://childes.talkbank.org doi:10.21415/T5CC7X.
We are very thankful for the kindness and collegiality you are showing in contributing your hard-won data.