Capturing children’s voices beyond words: Quantitative methods for coding vocalisation and gesture in early communication
Background: Understanding child communication requires approaches that reflect its multimodal nature. Young children communicate through vocalisations, gestures, facial expressions, and emerging words before speech is fully conventionalised. However, research and clinical frameworks rely heavily on speech-based measures, which often under-represent multimodal behaviour. This limits our ability to detect developmental variability and may obscure communicative competence in children who are not yet relying on speech. More inclusive methods are needed to capture communication across modalities.
Aim: This presentation introduces a quantitative multimodal coding framework for analysing vocalisation and gesture in naturalistic settings. The goal of the framework is to support analysis of communicative variability across children with diverse developmental profiles, including those who rely on non-verbal or emerging communicative forms. Ethical approval was obtained through institutional review board procedures for observational video and audio recording with caregiver consent.
Method/approach: The method uses video-based behavioural coding to analyse longitudinal child interaction data. Vocal and gestural forms are coded in parallel. Communicative acts are classified along four dimensions: modality (vocal or gestural), form (modality-specific features), communicative function (intended meaning or purpose), and social directivity (whether the act is socially oriented). This allows direct comparison across developmental time. The first cohort (n=10) was observed at 4, 7, and 11 months during parent-infant interactions. The second cohort (n=12) was observed at 13, 16, and 20 months during caregiver-child play sessions.
Results/findings: Across more than 9,000 communicative events, the framework captured systematic, quantifiable patterns of multimodal communication. In the infant cohort, prelinguistic vocalisations occurred substantially more frequently than gestures (3903 vs. 752 events) and were significantly more likely to be socially directed (34.6% vs. 17.7%). In the toddler cohort (4729 events), vocalisations outnumber gestures, though their distribution varied across communicative functions. Vocalisations were more often socially directed than gestures. The framework captured how relations between modality, function, and social directivity change across developmental time.
Conclusion: Overall, this framework enables direct quantitative comparison of vocal and gestural communication across early development. It provides a structured and transparent approach for studying multimodal communication and supports more inclusive and developmentally sensitive models of child communication.
Implications for children and families: Children communicate in many ways beyond speech, and this approach helps better recognise and understand how your child expresses meaning.
Implications for practitioners: This framework helps you systematically capture and interpret vocal and gestural communication, supporting more complete assessment of early communicative variability beyond spoken language.
Keywords: multimodal communication, quantitative coding methods, early language development, gesture and vocalisation analysis
This presentation relates to the following United Nations Sustainable Development Goals: