Work in progress — pages and figures are still being written.

Alicia Zeng

Research

Cortical flatmap of one subject from the NeurIPS 2025 poster, with voxels coloured by their selectivity for number, time and space.

Encoding models using word embeddings or artificial neural network (ANN) features reliably predict brain responses to naturalistic stimuli, yet interpreting these models remains challenging. A central limitation is superposition: distinct semantic features become entangled along correlated directions in dense embeddings when latent features outnumber embedding dimensions. This entanglement renders regression weights non-identifiable—different combinations of semantic directions can produce identical predictions, precluding principled interpretation of voxel selectivity. To address this, we introduce the Sparse Concept Encoding Model, which transforms dense embeddings into a higher-dimensional, sparse, non-negative space of learned concept atoms. This transformation yields an axis-aligned semantic basis where each dimension corresponds to an interpretable concept, enabling direct readout of conceptual selectivity from voxel weights. When applied to fMRI data collected during story listening, our model matches the prediction performance of conventional dense models while substantially enhancing interpretability. It enables novel neuroscientific analyses such as disentangling overlapping cortical representations of time, space, and number, and revealing structured similarity among distributed conceptual maps. This framework offers a scalable and interpretable bridge between ANN-derived features and human conceptual representations in the brain.

From Figure 3.1 of the thesis: neural spectrograms from three recording sites over one day at home, under over-, under- and preferred-stimulation.

Understanding brain function requires modeling a high-dimensional, nonlinear dynamical system continuously driven by sensory input, shaped by internal state and context, and expressed through behavior. Machine learning can be used to extract predictive structure from large-scale neural recordings, but its scientific value hinges on whether fitted parameters admit interpretation as testable hypotheses about neural computation. This thesis develops and applies interpretable predictive models—primarily linearized encoding models that map stimulus or behavioral features onto neural responses—to investigate (i) how human cortex represents abstract concepts during naturalistic language comprehension and (ii) how deep brain stimulation (DBS) modulates movement-related neural dynamics in Parkinson’s disease. The first study addresses a central limitation of encoding models built on dense word embeddings or neural network activations: superposition. When latent semantic factors outnumber embedding dimensions, distinct concepts become entangled along shared directions, rendering voxelwise regression weights uninterpretable and formally non-identifiable. I introduce the Sparse Concept Encoding Model, which applies sparse dictionary learning to project dense embeddings into an overcomplete basis of interpretable concept atoms. Because each atom isolates a coherent semantic direction, voxel tuning can be read directly from regression coefficients. Applied to whole-brain fMRI acquired during naturalistic story listening, this approach matches the prediction accuracy of conventional dense models while substantially improving interpretability. The resulting cortical maps reveal convergent selectivity for time, space, and number in the intraparietal sulcus and the inferior frontal gyrus—consistent with the shared-magnitude hypothesis—and uncover systematic similarity structure among distributed semantic representations. The second study moves to a clinical domain where characterizing neural dynamics carries immediate therapeutic implications. Parkinson’s disease is increasingly understood as a disorder of network dynamics wherein motor symptoms arise from aberrant oscillatory activity across basal ganglia and sensorimotor cortex. Yet, DBS remains an open-loop therapy with parameters tuned heuristically. Using a home-deployed multimodal data collection platform integrating bidirectional neurostimulators, bilateral wrist accelerometry, and video-based pose estimation, we recorded synchronized neural and behavioral data from two participants performing standardized motor tasks across multiple stimulation amplitudes. By building encoding models that predict neural spectral power data from behavioral features, we demonstrate how temporally rich predictors such as full acceleration spectrograms and task labels outperform coarse movement summaries. Thus, this study highlights the importance of fine-grained behavioral structure for identifying movement-related variance in sensed neural signals. In the participant with stronger baseline model performance, higher stimulation amplitudes enhanced prediction performance across sensorimotor cortex and subthalamic nucleus. We show that these gains arise from amplification of a movement-related spectral component characterized by beta suppression alongside alpha/delta and gamma enhancement, including a prominent peak at the stimulation-entrained subharmonic. This interpretation is consistent with “information lesion” accounts of DBS, wherein pathological synchronization impedes movement-related signaling and high-frequency stimulation restores information flow through basal ganglia–cortical circuits. Together, these studies illustrate how high-volume data collected under naturalistic or semi-naturalistic conditions can be analyzed with interpretable predictive models to transform accurate prediction into mechanistic neuroscientific hypotheses.

Principled Neuroscientific Discovery with Machine Learning PhD thesis, UC Berkeley,

Alicia Zeng

From Fig. 1 of the paper: movement threshold under movement-responsive DBS next to a participant wearing the implanted stimulator, two wrist watches and being filmed at home.

Deep brain stimulation (DBS) has garnered widespread use as an effective treatment for advanced Parkinson’s disease. Conventional DBS (cDBS) provides electrical stimulation to the basal ganglia at fixed amplitude and frequency, yet patients’ therapeutic needs are often dynamic with residual symptom fluctuations or side effects. Adaptive DBS (aDBS) is an emerging technology that modulates stimulation with respect to real-time clinical, physiological or behavioural states, enabling therapy to dynamically align with patient-specific symptoms. Here we report an aDBS algorithm intended to mitigate movement slowness by delivering targeted stimulation increases during movement using decoded motor signals from the brain. Our approach demonstrated improvements in dominant hand movement speeds and study participant-reported therapeutic efficacy compared with an inverted control, as well as increased typing speed and reduced dyskinesia compared with cDBS. Furthermore, we demonstrate proof of principle of a machine learning pipeline capable of remotely optimizing aDBS parameters in a home setting. This work illustrates the potential of movement-responsive aDBS as a promising therapeutic approach and highlights how machine learning-assisted programming can simplify complex optimization to facilitate translational scalability.

Movement-responsive deep brain stimulation for Parkinson’s disease using a remotely optimized neural decoder Nature Biomedical Engineering,

Tanner C. Dixon, Gabrielle Strandquist, Alicia Zeng, Tomasz Frączek, Raphael Bechtold, Daryl Lawrence, Shravanan Ravi, Philip A. Starr, Jack L. Gallant, Jeffrey A. Herron, Simon J. Little

From Figure 3.6 of the thesis: the precentral watch-acceleration weight matrix and its first principal component, negative in beta and positive in gamma.

Parkinson’s disease (PD) is a dynamical disease: motor symptoms arise not from static lesions but from aberrant neural activity across basal ganglia and sensorimotor cortex. Deep brain stimulation (DBS) is effective but remains open-loop, with parameters tuned heuristically. Here we use a home-deployed multimodal platform to quantify how DBS amplitude modulates movement-related neural activity. Two participants implanted with bidirectional neurostimulators (Medtronic Summit RC+S) completed six recording days at home, performing a UPDRS-derived motor battery at three stimulation amplitudes (under-, preferred-, and over-stimulation). We recorded local field potentials from subthalamic nucleus and electrocorticography from precentral and postcentral gyri, alongside bilateral wrist accelerometry and video data for pose estimation. Encoding models predicted neural spectral power from behavioral features: mean acceleration, full acceleration spectrogram, task labels, pose kinematics, and a combined set. Across participants, hemispheres, and recording sites, behavioral features robustly predicted neural activity; temporally rich representations—full acceleration spectrogram and task labels—substantially outperformed coarse summaries, indicating that fine-grained movement structure is necessary for capturing behaviorally driven neural variance. In the participant with stronger model performance, higher stimulation amplitudes enhanced prediction throughout the sensorimotor network by amplifying a movement-locked spectral component characterized by beta suppression and gamma enhancement. This pattern aligns with “information lesion” accounts of DBS in which stimulation restores movement-related signaling through pathologically synchronized circuits, and provides an empirical basis for adaptive neuromodulation informed by behavioral state.

From Figure 2.6 of the thesis: the word cloud of the number concept atom — fifteen, eleven, eighteen, seven, sixteen, nine, sixty, twenty, thirty, nineteen, ten, eight, seventeen.

Prior studies have shown that perceptual representations of number, space, and time concepts are represented in several different regions of the human cerebral cortex. These quantities can also be represented as lexical- semantic concepts, but little is known about how lexical-semantic representations of quantity are organized across the cerebral cortex? We used voxelwise encoding models and a sparse lexical-semantic feature space, to map quantity representations in 19 participants who listened to narrative stories in the fMRI scanner. We find that quantity is represented in a distributed temporal, parietal, and frontal network. Inspection of voxel tuning profiles showed that tuning for space, time or number falls along a continuum. Voxels in some regions are highly selective for a single quantity, while those in other regions respond to all three domains. These results suggest that the brain represents space, time and number concepts in a distributed network of regions with varying degrees of cross-domain integration. This distributed network may support the flexible use of quantity concepts in everyday tasks.

A Distributed Cortical Network Integrates Semantic Representations of Number, Space, and Time CCN , poster

Evi Hendrikx, Alicia Zeng, Yashaswini, Matteo Visconti di Oleggio Castello, Jack L. Gallant

[Still not complete — any remaining papers, preprints and talks go here.]