Summary
A Nature Neuroscience study used a single high-density brain implant to decode attempted speech and upper-body gestures in parallel from people with paralysis. The outputs drove a virtual avatar and displayed speech as text, while context-inclusive training improved decoding across isolated and simultaneous tasks.
Researchers have shown that a single brain implant can decode attempted speech and upper-body gestures in parallel in people with severe paralysis. In a proof-of-concept study published in Nature Neuroscience on 14 September 2026, the neural signals drove a personalised full-body virtual avatar: decoded gestures were animated by the avatar, while decoded speech outputs appeared as text.
The study involved three participants in the BRAVO clinical trial, although real-time avatar demonstrations and simultaneous speech-and-gesture results involved two of them. The work addresses a practical challenge for brain–computer interfaces (BCIs): natural communication often combines speech with facial or hand movements rather than using each function separately.
Contents
- One cortical grid covered multiple motor representations
- Concurrent behaviour required context-inclusive training
- The decoded outputs controlled a virtual avatar
- The study points towards larger communication systems
One cortical grid covered multiple motor representations
The researchers recorded electrical activity using a high-density electrocorticography (ECoG) array. The implant has 253 electrodes arranged with 3-millimetre centre-to-centre spacing and was placed beneath the skull on the surface of the left hemisphere, over areas involved in speech production and movement. A percutaneous connector carried signals from the array to external recording equipment.
The three participants had different conditions and movement capabilities. Bravo-1r and Bravo-3 had severe paralysis and loss of intelligible speech after brainstem strokes. Bravo-6 had amyotrophic lateral sclerosis (ALS), severe vocal-tract paralysis and moderate limb impairment. Their attempted movements also differed: some speech attempts were silent, one participant vocalised unintelligibly, and some gestures were minimally attempted or imagined.
Neural activity was recorded from the sensorimotor cortex while participants attempted movements involving the hands, arms, head, eyes and speech-related facial structures. Activity in the high-gamma range, 70–150 hertz, showed a broad organisation in which hand-related signals were generally more dorsal and speech- and orofacial signals more ventral. The areas were not completely separated. Some electrodes responded during both speech and gesture attempts, while others were more selective for one type of behaviour.
That partial overlap mattered because a decoder trained on one behaviour in isolation could learn a neural pattern that changed when another behaviour was attempted at the same time.
Concurrent behaviour required context-inclusive training
For the speech-and-gesture experiments, Bravo-1r used five speech phrases and four gestures, producing 12 selected speech–gesture combinations. Bravo-6 used 10 phrases, 10 gestures and all 100 possible pairings between them. The researchers compared three conditions: speech alone, gesture alone and simultaneous speech plus gesture.
Models trained only on isolated speech or gesture attempts did not fully generalise to simultaneous trials. Models trained only on simultaneous trials had the opposite problem, performing less well on isolated behaviour. The strongest performance across both contexts came from hybrid models trained on both isolated and simultaneous examples.
The researchers also trained each decoder to recognise the opposite modality as a rest condition. Without this cross-modality training, the median false-positive rate during opposite-modality trials was 30.6% for Bravo-6 and 76.0% for Bravo-1r. After opposite-modality examples were added to training, the median rate fell to 0.0% for both participants while remaining low during true rest.
Using the complete available datasets, mean offline simultaneous accuracy for Bravo-6 was 68.8% for gestures and 77.5% for speech. For Bravo-1r, the corresponding figures were 88.3% and 84.0%. These were classification results over small, participant-specific sets of phrases and gestures, rather than general speech-transcription performance.
An additional leave-one-out analysis in Bravo-6 found that the models could decode some simultaneous phrase–gesture pairings that had not appeared together during training. Median gesture accuracy was 66.7% for unseen pairings and 65.8% for seen pairings; median speech accuracy was 80.0% for unseen pairings and 79.3% for seen pairings.
The decoded outputs controlled a virtual avatar
In real-time copy tasks, Bravo-6’s gesture decoder reached 66.0% accuracy and its speech decoder 70.0% during simultaneous attempts involving 10 gestures and 10 phrases. In a limited conversational task across five blocks, average accuracy was 85.0% for gestures and 75.0% for speech. The chance level for each decoder in these 10-class tasks was 9.1% when rest was included.
Bravo-1r’s simultaneous speech-and-gesture performance was evaluated offline rather than in the real-time simultaneous task. In a limited conversational demonstration across three blocks, the median accuracy was 100.0% for both decoders, with chance levels of 16.7% for speech and 20.0% for gestures because the participant used different-sized vocabularies.
The speech channel displayed decoded phrases as text, while the gesture channel animated the corresponding movement on a personalised avatar. The system therefore restored a communication interface based on attempted speech and movement; it did not restore intelligible natural vocal speech.
The study points towards larger communication systems
The result is a technical feasibility demonstration rather than a clinical effectiveness study. The implant was used under a phase I, single-centre early-feasibility trial whose primary outcomes concern device safety and the feasibility of ECoG-controlled neuroprostheses. The present experiments included only three participants, and the avatar work involved two. Differences in paralysis cause, residual movement and attempt strategy may also affect the neural patterns available to a decoder.
The vocabulary was deliberately restricted so that the researchers could test every speech–gesture combination for Bravo-6. Practical systems would need larger repertoires, incremental or continuous decoding to reduce latency, and testing across more people with different forms and degrees of paralysis. The study’s central engineering finding is that such systems need training data representing both isolated and concurrent behaviour, rather than assuming that a decoder trained on one context will automatically transfer to the other.