ClassBank English MultiEd Corpus


Samantha Sie
Theoretical and Applied Linguistics
University of Cambridge

Ianthi Maria Tsimpli
Theoretical and Applied Linguistics
University of Cambridge

Ayesha Kidwai
Centre for Linguistics
Jawaharlal Nehru University

Participants: 7
Type of Study: teacher talk
Location: India
Media type: audio

Browsable transcripts

Download transcripts

Media folder

Citation information

In accordance with TalkBank rules, any use of data from this corpus must be accompanied by at least one of the above references.

Project Description

The Multilingualism in Education (MultiEd) project (2023-2026) was conducted with the aim to investigate, document, and optimise lesson delivery in English medium instruction (EMI) by means of pedagogical translanguaging in India. In this context, translanguaging involves the use of the home and regional languages of students and teachers to scaffold learning and facilitate teaching.

The strong societal demand for EMI schools in India reflects the value and prestige English carries locally. English is perceived not only as a language that grants access to knowledge but also as a language that enables upward social mobility. In spite of the Indian government’s effort to promote mother-tongue education, particularly in the early years of school, many parents across social classes aspire to enrol their children in EMI schools. Subsequently, India has over the years witnessed an increasing number of EMI schools established in response to this societal demand.

As a result, children attending EMI government schools mostly come from socioeconomically underprivileged families. Without adequate exposure to English and literacy support at home, these children are faced with a language barrier that impedes their overall learning experience and outcomes. To mitigate this, teachers engage in translingual practices, utilising their students’ linguistic resources to scaffold learning in the target language that is English.

Accordingly, the Teacher Talk Corpus (TTC) study – one of the research strands of the MultiEd project – presents an evidence-based documentation of lessons carried out by teachers in real classrooms. Importantly, it showcases teachers’ use of language, including naturalistic translanguaging.

Corpus description

The TTC is a small corpus comprising nine transcripts, which amount to 26,186 words and 5,587 utterances from approximately 6 hours 47 minutes of recorded speech. The audio recordings were carried out in Grades 4 and 5 English Language classrooms across four government-aided primary schools in New Delhi, India.

Seven English Language teachers (sex: 5 female, 2 male; age range: 31 – 45 years) took part in the TTC study. They were all multilingual speakers of at least Tamil, Hindi, and English. Most of them (n=6) acquired Tamil as their home or first language (L1), apart from one teacher whose L1 was Kannada. Another teacher acquired four languages including Malayalam in addition to the languages mentioned above. For all of them, Hindi and English were acquired as subsequent languages.

All seven teachers were university graduates: four of them were bachelor’s degree holders whereas the remaining three attained a master’s degree. All of them also undertook formal teacher training: most of them (n=6) had a Bachelor of Education, apart from one, who obtained a Diploma in Education.

The teachers had an average of 10 years (range: 7 – 15 years) of teaching experience at the time of data collection (2024).

In compliance with the UK’s Data Protection Act 2018, unique IDs were assigned to the teacher participants to protect their anonymity. Any callouts of personal names (e.g. names of students, teachers, research assistants) in the audio recordings were bleeped out; these names have also been replaced with fictitious names in the transcripts to ensure the anonymity of individuals.

Acknowledgements

The MultiEd project was funded by the British Council and awarded to the University of Cambridge (Cambridge Partnership for Education at Cambridge University Press & Assessment). The project was led by Professor Ianthi M. Tsimpli from the Department of Theoretical and Applied Linguistics as the Principal Investigator. We would like to thank Akshit Singh, Sankrithi Loganathan, and Sarannaya Bose for assisting with the transcriptions, as well as to Professor Brian MacWhinney and his team for their support with the morphological tagging and formatting of the corpus.