
I am involved in exciting research projects. I am looking for talented and committed Ph.D. students and/or Post-Doctoral trainees with strong background in statistical signal processing and machine learning. Experience in audio signal processing and high programming skills will be considered as an advantage. If interested, please contact me.
Please see our sponsorship.

Timo Gerkmann, University of Hamburg, Germany and Sharon GannotISF-DFG, 2026-2028

-Hearables are wearable devices such as headphones or earbuds that typically incorporate various sensors, including health monitoring sensors and multiple microphones, and offer advanced wireless connectivity, usually via Bluetooth. The market share of wireless devices in overall headphone sales is steadily increasing. Given the widespread availability of hearables, there is a strong demand for sophisticated algorithms tailored to their functionalities. This project specifically focuses on enhancing the assistive listening capabilities of hearables. An essential requirement for hearables is the preservation of binaural cues, which provide the necessary information for accurate spatial perception. Consequently, this project’s objective goes beyond creating state-of-the-art speaker separation and noise reduction algorithms; it also aims to maintain the auditory spatial impression for the hearing device user. To address these research topics, we will first explore data-driven methods for source localization and tracking. Tracking the direction of arrival of sources in an acoustic environment complicates when applied to hearables, primarily due to the rapid head movements of the device wearer, which can lead to fast changes in the entire scene. Consequently, we aim to develop a fast and efficient tracking mechanism that leverages contemporary techniques from the time-series analysis domain. Secondly, accurate binaural reproduction, along with informed speaker extraction and separation, necessitates access to individualized head-related transfer functions (HRTFs). Traditional procedures for measuring HRTFs are labor-intensive and require extensive setups with specialized equipment in an anechoic chamber. Our objective is to investigate methods that allow for the measurement of HRTFs in regular room environments, minimizing the number of measurements while ensuring high quality. Ultimately, we aim to facilitate the integration of real acoustically measured individualized HRTFs into future hearables through a user-friendly procedure. Next, we will investigate two paradigms for acoustics-aware binaural speech extraction and reproduction: predictive models and generative models. While predictive methods are widely utilized for speech extraction, generative methods are gaining significant attention, especially in single-microphone scenarios. By leveraging acoustic models, we aim to adapt these data-driven predictive and generative methods for the binaural context, with an emphasis on cue preservation and enhancing spatial interpretability within the network design. We will assess the advantages and limitations of both approaches in these tasks and derive valuable insights. In addition to scientific advancements, we will build a demo platform that integrates all relevant modules to work collaboratively. All algorithms and their integration will undergo comprehensive evaluation using both public databases and data collected specifically for this project.
Sharon Gannot
ISF, 2027-2030


Many State-of-the-Art (SOTA) multi-microphone Acoustic Signal Processing (ASP) methods—such as noise reduction, speaker separation, and acoustic source tracking-rely on accurate knowledge of the acoustic propagation from the source to the microphones, or on related descriptors of the acoustic environment.Typical acoustic features in this context are Acoustic Transfer Functions (ATFs) and Relative Transfer Functions (RTFs). However, estimating these high-dimensional features under adverse conditions remains a longstanding and unresolved challenge. Recent studies show that, despite their complexity, ATFs and RTFs lie on a low-dimensional manifold. Learning compact representations (embeddings) offers clear benefits: improved robustness by discouraging off-manifold solutions, inference of acoustic quantities such as room parameters and source positions, and controlled manipulation of acoustic responses. To fully exploit these benefits, the embeddings must remain physically interpretable and tightly linked to the underlying acoustic scene.
In this project, we propose Manifold Learning (ML) schemes that leverage modern deep-learning architectures. Specifically, we will investigate Variational Autoencoder (VAE) and Graph Neural Network (GNN) approaches for mapping high-dimensional acoustic quantities to compact low-dimensional embeddings. A central requirement is the preservation of the relative structure of the original data. By linking embeddings to Physically Interpretable Quantities (PIQs), such as source locations, changes in the ATF domain can be interpreted via their effects in the PIQ domain. This enables high-dimensional acoustic features to be analyzed, enhanced, and manipulated. In particular, we are interested in a specific implementation of GNNs, namely Graph Convolutional Networks (GCNs), which extend classical manifold learning concepts, traditionally based on geometric harmonics, toward more expressive and structured latent representations. While these ideas are gaining traction in computer vision and three-dimensional graphics, they remain largely unexplored in ASP.
The project will proceed in two phases. Phase A develops a deep ML framework for ASP. Phase B integrates classical statistical signal-processing paradigms with the ML models developed in Phase A to form hybrid algorithms.


Develop a novel paradigm and novel concept of socially-aware robots, and to conceive innovative methods and algorithms for computer vision, audio processing, sensor-based control, and spoken dialog systems based on modern statistical- and deep-learning to ground the required social robot skills.
Create and launch a brand new generation of robots that are flexible enough to adapt to the needs of the users, and not the other way around.
Validate the technology based on HRI experiments in a gerontology hospital, and to assess its acceptability by patients and medical staff.
Despite the many recent achievements in developing and deploying social robotics, there are still many underexplored environments and applications for which systematic evaluation of such systems by end-users is necessary. While several robotic platforms have been used in gerontological healthcare, the question of whether or not a social interactive robot with multi-modal conversational capabilities will be useful and accepted in real-life facilities is yet to be answered.
In the H2020 SPRING project we developed a software architecture, giving a full-sized humanoid robot social and conversational interaction capabilities. We conducted experiments with patients and companions in a day-care gerontological facility in Paris. Overall, the users were receptive to this technology, especially when the robot perception and action skills are robust to environmental clutter and flexible to handle a plethora of different interactions.
The software architecture is depicted in the following figure:

And the audio pipeline, developed by my group:

With the following modules:With the following modules:

Modern day environments are laden with rich stimuli all competing for our attention, a reality that poses substantial challenges for the perceptual system. Focusing attention exclusively on one important speaker and avoiding distraction is a major feat, in particular for individuals with hearing or attentional impairment.
In this project, we propose a unique combination of methodologies from signal processing, machine learning and brain research disciplines that can jointly develop novel algorithms and evaluation procedures capable of extracting and enhancing a desired speaker in adverse acoustic scenarios. Specifically, we harness the power of deep neural networks (DNNs) for audio-processing classification tasks, to develop approaches for training a DNN with multi-microphone speech recordings, and to overcome the complex nature of dealing with natural speech data with its inherent dynamic nature. Moreover, determining which speaker should be selectively enhanced will be guided automatically by the user’s momentary internal preferences using a real-time EEG-based neural interface. The novel and unique neuro-engineering approach is critical for developing stable technological solutions meant for human use and that need to adhere to real-life behavioural and environmental constraints, and can be applied more broadly in the design of new-generation “attentive” hearing devices, “hearables” and teleconferencing systems.

This project deals with developing single- and multi-microphone algorithms for speech enhancement in adverse conditions. Special concern will be given to hardware constraints and to real-time requirements.

As rich datasets for visual and auditory modalities accumulate, the field of machine perception progresses towards the challenge of understanding rich scenes that involve multiple co-interacting entities. Going beyond low-level perception, semantic understanding of rich-scenes is viewed as the next frontier of machine perception.
To evaluate understanding of rich scenes, one often aims at high-level decisions and reasoning about the scene. Specifically, research on language grounding, aims to connect high-level
representations in plain language referring to perceived objects and situations. As one concrete important example, in the task of referring expressions, a model is trained to detect an object of interest, based on its visual properties. As an example, a deep network can find a “boy wearing a red hat” in an image of several boys.
We propose to advance the state of the art in this domain by generating perception signals that are more natural in two ways. First, they refer to dynamic scenes (videos) where objects can move around. Second, objects are also referred to by the sounds they emit, forming a unified audio-visual reasoning dataset.
Two main challenges are addressed in this project: