Speech Pathways vs “Language Module”
Jarvis argues there is no separate brain “language module.” Instead, speech production circuits and auditory perception circuits each contain the computations needed for speaking and understanding.
In this Huberman Lab Essentials episode, my guest is Dr. Erich Jarvis, PhD, a professor and Head of the Laboratory of Neurogenetics of Language at Rockefeller University and an investigator at the Howard Hughes Medical Institute (HHMI). We discuss the brain circuits and genes underlying spoken language and why the ability to learn and produce vocalizations is extraordinarily rare in the animal kingdom. We also explore why song likely evolved before language, how gesture and movement share deep neural roots with speech, the neurobiology of stuttering, why childhood is the optimal window for language acquisition, and how physical movement — including dance — may help preserve speech and cognitive function across a lifetime. Read the show notes at hubermanlab.com. Thank you to our sponsors AG1: https://drinkag1.com/huberman Function: https://functionhealth.com/huberman Eight Sleep: https://eightsleep.com/huberman
Jarvis argues there is no separate brain “language module.” Instead, speech production circuits and auditory perception circuits each contain the computations needed for speaking and understanding.
Spoken language depends on specialized motor control of the larynx, jaw, and related muscles. Jarvis emphasizes this vocal production specialization is rare across species compared with sound perception.
Many animals can learn to understand human spoken commands without speaking. Jarvis uses dogs and great apes to illustrate that comprehension can be strong even when vocal production is limited.
Hand gesture control regions sit adjacent to speech production regions. Jarvis suggests speech circuits evolved from broader body-movement control circuits, which may explain why people gesture while talking.
Most vertebrates vocalize innately, like crying or barking. Learned vocal imitation is rare and is presented as the key ingredient that makes spoken language possible.
Primitive vocal sounds rely heavily on brainstem and basic emotional circuits. Vocal learners are described as having forebrain circuits that can exert control over brainstem vocal machinery to enable imitation.
Jarvis speculates spoken language is older than often assumed and may have existed in Neanderthals. He cites shared gene sequences in speech-related genes across ancient hominins, suggesting 500,000 to 1,000,000 years.
Human language learning is easier before puberty, and birds show analogous windows for tutor-song learning. Loss of hearing can degrade learned vocal output in vocal learners, unlike many non-learners.
Distinct bird nuclei like Area X and related pathways map onto analogous functions in human speech circuitry. Jarvis emphasizes convergence despite different anatomical names and distant evolutionary histories.
Specialized gene expression patterns in vocal learning circuits are reported as similar in humans and vocal-learning birds. Mutations tied to human speech deficits, including FOXP2-related effects, can produce comparable impairments in birds.
Some hummingbirds produce percussive wing sounds synchronized with their vocal song. Jarvis describes this as coordinated multimodal signaling that can mimic song syllables.
Young birds prefer learning their own species’ song but can learn other species’ songs if social exposure changes. Jarvis relates this to an inborn bias that constrains what is learned rather than replacing learning.
Children exposed to multiple languages during the critical period can blend phonemes and words into hybrid forms. Jarvis frames this as cultural evolution tracking genetic constraints, with shared sound units persisting most.
Jarvis highlights axon-guidance genes that can be turned off to allow new connections to form. He also emphasizes neuroprotection and calcium-buffering genes for high firing rates, plus plasticity genes that support complex learning.
Jarvis suggests critical periods reflect broad brain development, not just speech circuits. He argues the brain balances rapid learning with limits on memory and the need to stabilize skills for long-term use.
A proposed advantage is retaining a wider set of producible phonemes rather than permanently higher plasticity. Having more sound building blocks available can make acquiring additional languages faster in adulthood.
Jarvis distinguishes emotional “affective” signaling from meaning-focused “semantic” signaling. He argues the same vocal circuits can support both, but are deployed differently depending on context and goals.
He describes left dominance for speech production and more right involvement for singing and musical processing, while noting both hemispheres contribute. This is linked to common ideas about lateralization of language and artistry.
Jarvis notes that most vocal learners use learned sounds mainly for emotional communication. He presents a hypothesis that human spoken language may have built on earlier selection for song-like courtship and display.
Non-human primates have strong cortical control over facial musculature, enabling rich expression without vocal imitation. Jarvis argues humans layered learned voice on top of existing facial signaling to reduce ambiguity in communication.
Reading is described as visual input flowing to speech-production circuits, creating silent inner speech and an auditory-like “hearing” of it. Writing is framed as translation from speech and auditory representations into hand-motor output.
Jarvis reports stuttering-like phenomena in songbirds after basal ganglia damage during recovery. He connects this to human neurogenic stuttering and broader evidence implicating basal ganglia disruption in developmental stuttering.
He suggests many stuttering interventions leverage sensory-motor integration. More deliberate control of what is heard relative to what is produced can reduce disfluency.
Jarvis argues heavy texting increases rapid communication and shifts practice to different modalities. He frames it as “use it or lose it,” where some circuits strengthen with use while others may receive less practice.
Jarvis links consistent whole-body movement, including dance, to keeping brain circuits engaged and cognition resilient. He recommends ongoing movement and vocal practice like speaking, singing, or oratory to keep related circuits tuned.