Biography
I am a Postdoctoral Researcher at Tsinghua University, working with Prof. Chao Zhang. My research focuses on multimodal foundation models for speech and physiological signals, with applications in sleep studies, speech disorders, and clinical health assessment. I received my B.Eng. in Software Engineering from Dalian University of Technology and my Ph.D. in Systems Engineering and Engineering Management from The Chinese University of Hong Kong, where I was supervised by Prof. Xunying Liu. My doctoral research examined adversarial and reinforcement learning methods for data augmentation in dysarthric and elderly speech recognition.
Selected Publications
Contribution to Open Source
JinZr/sleep2vec
Unified cross-modal alignment for heterogeneous nocturnal biosignals.
k2-fsa/icefall
icefall contains ASR recipes for various datasets using https://github.com/k2-fsa/k2.
lhotse-speech/lhotse
Lhotse is a Python library aiming to make speech and audio data preparation flexible and accessible to a wider community. Alongside k2, it is a part of the next generation Kaldi speech processing library.
Academic Service
My Polaroid Gallery
I have one Polaroid Spectra for shooting B&W film, one SX-70 Sonar, and one SLR680 for regular shooting. My Polaroid camera collection also includes an SLR680 Special Edition (Blue Button Version), an SX-70 Model 2, a 670-AF, and a 670-AF Special Edition (also known as the Blue Button Version).