In everyday life, humans are constantly immersed in acoustically complex environments, yet they can effortlessly focus on a single attended speech stream while suppressing other competing unattended streams. This remarkable ability is known as the cocktail-party effect (Ahmed et al., 2023). Understanding how the central nervous system selectively extracts relevant auditory information while suppressing competing inputs, and identifying the neural mechanisms underlying this process, have long been central topics in auditory and cognitive neuroscience (Bednar and Lalor, 2020).
Recent studies have revealed that the cocktail-party effect is not solely a property of individual brains, but also involves inter-brain neural synchronization between communicating individuals (Holtze et al., 2022). Compared with traditional brain-sound neural resonance approaches (i.e., brain-to-stimulus coupling), inter-brain synchronization directly examines neural alignment between interacting brains (Rosenkranz et al., 2021), bypassing the acoustic signal as an intermediate representation. This perspective offers a more direct window into the neural mechanisms underlying selective listening in natural communication. Clarifying how attended and unattended streams differentially shape inter-brain coupling is therefore crucial not only for understanding the neural basis of the cocktail-party effect, but also for informing the development of brain-brain interfaces under realistic multi-talker conditions.
Limitations of existing approachesTo date, most research on the neural mechanisms of the cocktail-party effect has focused on neural tracking of speech signals, using EEG, MEG, or fNIRS (Ahmed et al., 2023; Keshavarzi and Varano, 2021; Mesgarani and Chang, 2012; O'Sullivan et al., 2019). These studies have provided compelling evidence that attended speech is preferentially represented at higher levels of the auditory hierarchy, while unattended speech is largely confined to early auditory cortex.
However, speech signals themselves are the result of multiple stages of neural processing and complex articulatory transformations. As such, brain-sound coupling inevitably reflects a mixture of sensory encoding, motor production, and environmental distortion. In contrast, directly measuring neural signals from interacting brains allows researchers to examine the neural dynamics of communication at their source. Inter-brain synchronization therefore offers a more direct and theoretically grounded approach for probing the neural basis of selective listening in cocktail-party scenarios.
Despite its promise, inter-brain research on the cocktail-party effect remains underdeveloped. Most existing studies rely on EEG or fNIRS (Dai et al., 2018; Holtze et al., 2022; Kuhlen et al., 2012; Li J. et al., 2023; Li et al., 2021; Li Z. et al., 2023; Rosenkranz et al., 2021), which, although sensitive to temporal dynamics, suffer from limited spatial resolution and are largely insensitive to deep-brain (subcortical) activity. Consequently, these studies are unable to precisely localize where inter-brain synchronization emerges, particularly within deeper cortical and subcortical regions, nor can they adequately distinguish the neural mechanisms supporting attended vs. unattended streams.
Advantages of fMRI-based hyperscanningfMRI-based hyperscanning has emerged in recent years as a powerful tool for studying neural coupling during social communication (Hausfeld et al., 2024; Liu et al., 2020; Speer et al., 2024; Stephens et al., 2010; Xie et al., 2020). Compared with EEG, MEG, and fNIRS, fMRI offers superior spatial resolution and whole-brain coverage that includes both cortical and subcortical structures, while avoiding the invasiveness and limited applicability of electrocorticography (ECoG).
The high spatial precision of fMRI hyperscanning is particularly advantageous for research on the neural mechanisms of the cocktail-party effect. It enables accurate localization of neural activity and inter-brain synchronization across distributed cortical and subcortical systems involved in speech processing and attentional control. This opens the door to disentangling the neural substrates associated with attended vs. unattended speech at multiple hierarchical levels, a distinction that remains difficult to achieve with existing noninvasive techniques. Importantly, compared with EEG- or fNIRS-based hyperscanning, which offers only coarse and spatially limited measures of inter-brain coupling, fMRI hyperscanning provides anatomically precise mapping of cross-brain synchrony. This allows researchers to identify the specific networks involved and to advance from descriptive observations toward mechanistic explanations of how selective listening shapes communication-related circuitry.
Methodological challenges and practical solutionsApplying fMRI hyperscanning to cocktail-party effect paradigms is not without challenges. First, MRI scanners generate substantial acoustic noise, imposing stringent requirements on real-time speech denoising. Second, the scanner environment precludes face-to-face interaction, potentially reducing ecological validity. Third, true hyperscanning beyond dyads would require the simultaneous operation of three or more MRI scanners, which poses substantial logistical, financial, and technical challenges. Fourth, compared with indirect neuroimaging techniques such as EEG and fNIRS, fMRI has relatively limited temporal resolution, which may constrain the investigation of rapid, dynamic neural processes underlying spoken language communication.
These challenges, however, are not insurmountable. A staged approach that incrementally balances technical feasibility with scientific ambition may offer a pragmatic pathway. In the first stage, pseudo-hyperscanning paradigms can be employed to lower initial technical barriers and achieve partial objectives such as establishing robust inter-brain coupling metrics, validating speech denoising pipelines, and refining stimulus design under controlled conditions. In such designs, neural signals are first recorded from multiple speakers during speech production, and their denoised speech signals are then combined offline to construct cocktail-party stimuli subsequently presented to listeners during fMRI scanning. This approach preserves key advantages of hyperscanning—namely inter-brain coupling analysis—while dramatically reducing immediate technical demands and scanner requirements.
In the second stage, fully real hyperscanning research using simultaneous fMRI of multiple participants should be undertaken to address real-time cocktail-party language communication. In this phase, pairs or triplets of interacting participants are scanned simultaneously in different MR scanners with synchronized acquisition. Each speaker's speech signal should undergo real-time denoising prior to mixing, independent of the experimental manipulation. The two denoised streams can then be adaptively combined to construct a cocktail-party acoustic scene, with mixing parameters experimentally controlled (e.g., via spatial or speaker-identity cues) before real-time presentation to the listener inside the scanner. To enhance ecological validity, a video interface displaying the communication partners can be incorporated so that participants engage in more face-to-face–like interaction despite the physical constraints of the scanner environment.
In parallel with experimental implementation, the analytical framework for inter-brain data should be explicitly specified. Several established and complementary approaches can be employed for fMRI hyperscanning data analysis. First, inter-brain functional coupling can be quantified using Pearson correlation, which has been extensively validated in functional connectivity research and provides a straightforward measure of temporal synchrony between homologous or functionally defined regions across participants (Biswal et al., 1995). Despite its simplicity, correlation-based inter-subject coupling has proven robust in naturalistic paradigms, including narrative communication (Stephens et al., 2010). Second, multivariate linear regression models offer important advantages by enabling the explicit modeling and removal of confounding factors, such as head motion parameters, physiological noise regressors, scanner drift, and task-related covariates within a general linear model (GLM) framework (Satterthwaite et al., 2013). Such approaches improve the specificity of inter-brain coupling estimates by reducing shared artifactual variance, which is particularly critical in hyperscanning contexts where motion and acoustic artifacts may be correlated across participants. Third, wavelet coherence analysis provides a powerful time–frequency framework for assessing inter-brain coupling across multiple temporal scales. Unlike static correlation measures, wavelet coherence can characterize non-stationary and frequency-specific synchronization patterns, making it especially suitable for investigating dynamic social interaction and speech-related rhythms (Chang and Glover, 2010). Given that conversational speech contains hierarchical temporal structure (e.g., syllabic and phrasal rhythms), time–frequency approaches may reveal frequency-dependent inter-brain alignment that is not captured by stationary metrics.
To further validate inter-brain coupling measures, stimulus-driven benchmarks can be incorporated. Specifically, the speech amplitude envelope of each speaker can be convolved with a canonical hemodynamic response function (HRF) and correlated with the listener's BOLD signal. Such speech–brain coupling analyses have been widely used to quantify neural entrainment to continuous speech (Lerner et al., 2011). Convergence between envelope-based stimulus–brain coupling and inter-brain coupling metrics would provide an important cross-validation of the analytical framework, ensuring that observed interpersonal neural alignment reflects meaningful speech-driven processes rather than shared noise or scanner-related artifacts.
A key limitation of conventional fMRI is its relatively low temporal resolution compared with techniques like EEG or MEG, which may constrain the characterization of rapid dynamic neural processes. However, recent advances in multiband (simultaneous multi-slice, SMS) have substantially improved temporal sampling. For example, multiband SMS protocols have enabled whole-brain coverage with repetition times (TRs) reduced from the typical 2–3 s to sub-second ranges such as ~0.72 s in large population projects like the Human Connectome Project (Wall, 2023), thereby increasing the effective sampling rate of BOLD signals. More sophisticated hybrid methods combining multiband encoding with advanced readout strategies (e.g., echo-volumar encoding) have been developed that sustainably support even shorter TRs (e.g., 118–650 ms) without sacrificing whole-brain coverage and with maintained temporal SNR, enabling the sensitive mapping of neural processes at higher frequency bands (e.g., above 0.3 Hz) (Posse et al., 2025). At the extreme, proof-of-principle work combining multiband acceleration with innovative reshuffling strategies has demonstrated effective BOLD sampling rates of ~75 ms, which would, according to Nyquist's sampling theorem, support neural signal components up to ~6–7 Hz in principle (Schmidt et al., 2023). Importantly, these accelerated sampling regimes also make it feasible—at least from a sampling-adequacy perspective—to probe faster neural responses aligned with speech temporal structure, including syllabic (~1.5 Hz) and phrasal (~3 Hz) rhythmic components (Meng et al., 2021). Such improvements in temporal resolution bring fMRI closer to the timescales of many aspects of speech and language exchange, making real-time hyperscanning more feasible.
Collectively, these methodological advances establish a scalable framework for investigating the neural mechanisms underlying the cocktail-party effect through fMRI hyperscanning, bridging controlled experimental models and real-time multi-speaker social interaction.
Beyond a unitary cocktail-party effectFinally, it is worth noting that the cocktail-party effect is not a single, homogeneous phenomenon. Selective listening can rely on multiple cues, including voice identity, timbre, semantics, and spatial location (Liu et al., 2024). Yet many studies treat these diverse mechanisms as manifestations of a single effect. The spatial resolution of fMRI hyperscanning makes it possible to systematically dissociate the neural mechanisms supporting different cue-based forms of selective listening, thereby refining theoretical understanding of the cocktail-party effect.
ConclusionfMRI-based hyperscanning offers a powerful and timely approach for advancing research on the neural mechanisms underlying the cocktail-party effect. By leveraging its high spatial resolution and whole-brain coverage, this method enables systematic dissociation of the neural processes supporting different cue-based forms of selective listening, thereby overcoming key conceptual and methodological limitations of prior approaches. Integrating fMRI hyperscanning into research on the cocktail-party effect thus provides a more precise mechanistic framework for understanding how brains dynamically coordinate during complex communicative environments.
StatementsAuthor contributionsQL: Writing – original draft, Writing – review & editing.
FundingThe author(s) declared that financial support was received for this work and/or its publication. This work was supported by the National Natural Science Foundation of China (Grant No. 32460203) and the Guizhou Provincial Basic Research Program (General Project) (Grant No. QKHJCMS [2026] 485).
AcknowledgmentsThe author gratefully acknowledges Prof. Xiaochu Zhang from the University of Science and Technology of China for insightful discussions and valuable suggestions during the development of this work.
Conflict of interestThe author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statementThe author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s noteAll claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
ReferencesAhmedF.NidifferA. R.O'SullivanA. E.ZukN. J.LalorE. C. (2023). The integration of continuous audio and visual speech in a cocktail-party environment depends on attention. Neuroimage274:120143. doi: 10.1016/j.neuroimage.2023.120143
BednarA.LalorE. C. (2020). Where is the cocktail party? Decoding locations of attended and unattended moving sound sources using EEG. Neuroimage205:116283. doi: 10.1016/j.neuroimage.2019.116283
BiswalB.Zerrin YetkinF.HaughtonV. M.HydeJ. S. (1995). Functional connectivity in the motor cortex of resting human brain using echo-planar mri. Magn. Reson. Med. 34, 537–541. doi: 10.1002/mrm.1910340409
ChangC.GloverG. H. (2010). Time–frequency dynamics of resting-state brain connectivity measured with fMRI. Neuroimage50, 81–98. doi: 10.1016/j.neuroimage.2009.12.011
DaiB.ChenC.LongY.ZhengL.ZhaoH.BaiX.et al. (2018). Neural mechanisms for selectively tuning in to the target speaker in a naturalistic noisy situation. Nat. Commun. 9:2405. doi: 10.1038/s41467-018-04819-z
HausfeldL.HamersI. M. H.FormisanoE. (2024). FMRI speech tracking in primary and non-primary auditory cortex while listening to noisy scenes. Commun. Biol. 7:1217. doi: 10.1038/s42003-024-06913-z
HoltzeB.RosenkranzM.JaegerM.DebenerS.MirkovicB. (2022). Ear-EEG measures of auditory attention to continuous speech. Front. Neurosci. 16:869426. doi: 10.3389/fnins.2022.869426
KeshavarziM.VaranoE. (2021). Cortical tracking of a background speaker modulates the comprehension of a foreground speech. Signal41, 5093–5101. doi: 10.1523/JNEUROSCI.3200-20.2021
KuhlenA. K.AllefeldC.HaynesJ.-D. (2012). Content-specific coordination of listeners' to speakers' EEG during communication. Front. Hum. Neurosci.6:266. doi: 10.3389/fnhum.2012.00266
LernerY.HoneyC. J.SilbertL. J.HassonU. (2011). Topographic mapping of a hierarchy of temporal receptive windows using a narrated story. J. Neurosci. 31, 2906−2915. doi: 10.1523/JNEUROSCI.3684-10.2011
LiJ.HongB.NolteG.EngelA. K.ZhangD. (2023). EEG-based speaker–listener neural coupling reflects speech-selective attentional mechanisms beyond the speech stimulus. Cereb. Cortex33, 11080–11091. doi: 10.1093/cercor/bhad347
LiZ.HongB.WangD.NolteG.EngelA. K.ZhangD. (2023). Speaker–listener neural coupling reveals a right-lateralized mechanism for non-native speech-in-noise comprehension. Cereb. Cortex33, 3701–3714. doi: 10.1093/cercor/bhac302
LiZ.LiJ.HongB.NolteG.EngelA. K.ZhangD. (2021). Speaker–listener neural coupling reveals an adaptive mechanism for speech comprehension in a noisy environment. Cereb. Cortex31, 4719–4729. doi: 10.1093/cercor/bhab118
LiuH.BaiY.ZhengQ.LiuJ.ZhuJ.NiG. (2024). Electrophysiological correlation of auditory selective spatial attention in the “cocktail party” situation. Hum. Brain Mapp. 45, 1–14. doi: 10.1002/hbm.26793
LiuL.ZhangY.ZhouQ.GarrettD. D.LuC.ChenA.et al. (2020). Auditory–articulatory neural alignment between listener and speaker during verbal communication. Cereb. Cortex30, 942–951. doi: 10.1093/cercor/bhz138
MengQ.HegnerY. L.GiblinI.McMahonC.JohnsonB. W. (2021). Lateralized cerebral processing of abstract linguistic structure in clear and degraded speech. Cereb. Cortex31, 591–602. doi: 10.1093/cercor/bhaa245
MesgaraniN.ChangE. F. (2012). Selective cortical representation of attended speaker in multi-talker speech perception. Nature485, 233–236. doi: 10.1038/nature11020
O'SullivanJ.HerreroJ.SmithE.ShethS. A.MehtaA. D.MesgaraniN.et al. (2019). Hierarchical encoding of attended auditory objects in multi-talker speech perception. Neuron104, 1195–1209.e3. doi: 10.1016/j.neuron.2019.09.007
PosseS.RamannaS.MoellerS.VakamudiK.OtazoR.Sa de La Rocque GuimaraesB.et al. (2025). Real-time fMRI using multi-band echo-volumar imaging with millimeter spatial resolution and sub-second temporal resolution at 3 tesla. Front. Neurosci. 19:1543206. doi: 10.3389/fnins.2025.1543206
RosenkranzM.HoltzeB.JaegerM.DebenerS. (2021). EEG-based intersubject correlations reflect selective attention in a competing speaker scenario. Front. Neurosci. 15:685774. doi: 10.3389/fnins.2021.685774
SatterthwaiteT. D.ElliottM. A.GerratyR. T.RuparelK.LougheadJ.CalkinsM. E.et al. (2013). An improved framework for confound regression and filtering for control of motion artifact in the preprocessing of resting-state functional connectivity data. Neuroimage64, 240–256. doi: 10.1016/j.neuroimage.2012.08.052
SchmidtT.VannesjoS. J.SommerS.NagyZ. (2023). fMRI with whole-brain coverage, 75-ms temporal resolution and high SNR by combining HiHi reshuffling and multiband imaging. Magn. Reson. Imaging103, 48–53. doi: 10.1016/j.mri.2023.06.015
SpeerS. P. H.Mwilambwe-TshiloboL.TsoiL.BurnsS. M.FalkE. B.TamirD. I. (2024). Hyperscanning shows friends explore and strangers converge in conversation. Nat. Commun.15:7781. doi: 10.1038/s41467-024-51990-7
StephensG. J.SilbertL. J.HassonU.
Comments (0)