The Speech Health Lab Department of Psychiatry · AIIMS New Delhi
Phase II · Clinical validation
Acoustic biomarkers in mental health
Depression often hides in plain sight. At the Lab we are developing tools that listen for the subtle
acoustic changes that accompany it, with the aim of making screening timely, objective
and accessible.
Every utterance carries two layers: what is said, and how it is said. The second is where
we look.
Our philosophy
Bridging care and technology
Technology should extend clinical judgement, not stand in for it. Speech is one of the
few clinical signals a person produces continuously, at no cost, without instrumentation
which makes it worth understanding properly.
Linguistic content what is said
Word choice, semantic coherence, narrative structure. Well studied, but bound to
language and culture, and dependent on transcription.
Acoustic form how it is said
Rhythm, prosody, pitch variability, pause structure, voice quality. Largely
language-independent and measurable directly from the waveform. This is the layer the
lab investigates.
Two directions, one question
Whether an acoustic signal can support depression assessment depends on holding it up
against both clinical populations and the wider community it would eventually serve.
Clinical focus
Patient outreach
Depression assessment and assistive diagnostic technology sit at the core of the lab.
Working under the guidance of investigators in the Department of Psychiatry, and with
funding support from Coal India Limited, we develop and refine assessment tools through
direct interaction with people living with depression.
Expanding horizons
Community outreach
Mental health research has to reach past the laboratory. Any diagnostic model needs
robust, unbiased baseline data from people who are not patients, across age groups and
generations. Working with communities on the foundations of stress and anxiety is how
that baseline gets built, and how methods find their way onto web and mobile platforms
where people can actually use them.
Method
From recording to validated measure
Each stage is designed to be auditable, so a result can be traced back to the audio it came from.
Acquisition
Consented, ethically approved collection of speech samples under controlled conditions.
Pre-processing
Noise reduction, segmentation and normalisation, so recordings made in different rooms remain comparable.
Extraction
Paralinguistic feature extraction prosody, timing, pause structure and voice quality.
Modelling
Statistical and machine learning models relating acoustic features to clinical measures.
Validation
Cross-reference against structured clinical assessment, with held-out data and clinician review.
Laboratory data pipeline.
The lab
Who this work runs with
Dr Himanshu Singh works full time at the Speech Health Lab as Scientist-C. The work
described on this page is carried out under Professor Dr Nand Kumar, Department of
Psychiatry, All India Institute of Medical Sciences, New Delhi, and is actively funded
by Coal India Limited.
The lab maintains its own site at
speechhealth.org.
The work is presented here to document the projects Dr Singh contributes to; the lab
and its outputs belong to the Department.
Research in progress. Methods described here are under active
development and clinical validation. Nothing on this page is a diagnostic tool, and
none of it should be used in place of assessment by a qualified clinician.
Work with the lab
For collaboration, participation or questions about the methods, get in touch directly.