本文经arXiv每日学术速递授权转载
【1】 Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
标题: 多个众包录音的音频增强:简单有效的基线
作者:Shiran Aziz,Yossi Adi,Shmuel Peleg
链接:点击下载PDF文件
【2】 Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
标题: 紧紧抓住我:语音增强的稳定编码器-解码器设计
作者:Daniel Haider,Felix Perfler,Vincent Lostanlen,Martin Ehler,Peter Balazs
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【3】 AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
标题: AASIST 3:KAN增强AASIST语音Deepfake检测使用SSL功能和额外正规化应对ASVspoof 2024挑战
作者:Kirill Borodin,Vasiliy Kudryavtsev,Dmitrii Korzh,Alexey Efimenko,Grach Mkrtchian,Mikhail Gorodnichev,Oleg Y. Rogov
备注:8 pages, 2 figures, 2 tables. Accepted paper at the ASVspoof 2024 (the 25th Interspeech Conference)
链接:点击下载PDF文件
【4】 Utilizing Speaker Profiles for Impersonation Audio Detection
标题: 利用扬声器配置文件进行模拟音频检测
作者:Hao Gu,JiangYan Yi,Chenglong Wang,Yong Ren,Jianhua Tao,Xinrui Yan,Yujie Chen,Xiaohui Zhang
备注:Accepted by ACM MM2024
链接:点击下载PDF文件
【5】 Point Neuron Learning: A New Physics-Informed Neural Network Architecture
标题: 点神经元学习:一种新的物理信息神经网络架构
作者:Hanwen Bi,Thushara D. Abhayapala
备注:under the review process of EURASIP Journal on Audio, Speech, and Music Processing
链接:点击下载PDF文件
【6】 Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
标题: 编解码器确实很重要:探索音频语言模型编解码器的语义缺陷
作者:Zhen Ye,Peiwen Sun,Jiahe Lei,Hongzhan Lin,Xu Tan,Zheqi Dai,Qiuqiang Kong,Jianyi Chen,Jiahao Pan,Qifeng Liu,Yike Guo,Wei Xue
链接:点击下载PDF文件
【7】 Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
标题: 从多说话人录音中提取说话人嵌入的循环注意池
作者:Shota Horiguchi,Atsushi Ando,Takafumi Moriya,Takanori Ashihara,Hiroshi Sato,Naohiro Tawara,Marc Delcroix
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【8】 User-Driven Voice Generation and Editing through Latent Space Navigation
标题: 通过潜在空间导航的用户驱动语音生成和编辑
作者:Yusheng Tian,Junbin Liu,Tan Lee
链接:点击下载PDF文件
标题: SelectTTC:通过基于离散单元的帧选择合成任何人的语音
作者:Ismail Rasim Ulgen,Shreeram Suresh Chandra,Junchen Lu,Berrak Sisman
备注:Submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
【2】 Advancing Multi-talker ASR Performance with Large Language Models
标题: 利用大型语言模型提高多说话者ASB性能
作者:Mohan Shi,Zengrui Jin,Yaoxun Xu,Yong Xu,Shi-Xiong Zhang,Kun Wei,Yiwen Shao,Chunlei Zhang,Dong Yu
备注:8 pages, accepted by IEEE SLT 2024
链接:点击下载PDF文件
【3】 Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
标题: 编解码器确实很重要:探索音频语言模型编解码器的语义缺陷
作者:Zhen Ye,Peiwen Sun,Jiahe Lei,Hongzhan Lin,Xu Tan,Zheqi Dai,Qiuqiang Kong,Jianyi Chen,Jiahao Pan,Qifeng Liu,Yike Guo,Wei Xue
链接:点击下载PDF文件
【4】 Learning Multi-Target TDOA Features for Sound Event Localization and Detection
标题: 学习多目标TDOE特征以进行声音事件定位和检测
作者:Axel Berg,Johanna Engman,Jens Gulin,Karl Åström,Magnus Oskarsson
备注:DCASE 2024
链接:点击下载PDF文件
【5】 Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
标题: 从多说话人录音中提取说话人嵌入的循环注意池
作者:Shota Horiguchi,Atsushi Ando,Takafumi Moriya,Takanori Ashihara,Hiroshi Sato,Naohiro Tawara,Marc Delcroix
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【6】 User-Driven Voice Generation and Editing through Latent Space Navigation
标题: 通过潜在空间导航的用户驱动语音生成和编辑
作者:Yusheng Tian,Junbin Liu,Tan Lee
链接:点击下载PDF文件
【7】 Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
标题: 多个众包录音的音频增强:简单有效的基线
作者:Shiran Aziz,Yossi Adi,Shmuel Peleg
链接:点击下载PDF文件
【8】 Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
标题: 紧紧抓住我:语音增强的稳定编码器-解码器设计
作者:Daniel Haider,Felix Perfler,Vincent Lostanlen,Martin Ehler,Peter Balazs
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【9】 AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
标题: AASIST 3:KAN增强AASIST语音Deepfake检测使用SSL功能和额外正规化应对ASVspoof 2024挑战
作者:Kirill Borodin,Vasiliy Kudryavtsev,Dmitrii Korzh,Alexey Efimenko,Grach Mkrtchian,Mikhail Gorodnichev,Oleg Y. Rogov
备注:8 pages, 2 figures, 2 tables. Accepted paper at the ASVspoof 2024 (the 25th Interspeech Conference)
链接:点击下载PDF文件
【10】 Utilizing Speaker Profiles for Impersonation Audio Detection
标题: 利用扬声器配置文件进行模拟音频检测
作者:Hao Gu,JiangYan Yi,Chenglong Wang,Yong Ren,Jianhua Tao,Xinrui Yan,Yujie Chen,Xiaohui Zhang
备注:Accepted by ACM MM2024
链接:点击下载PDF文件
【11】 Point Neuron Learning: A New Physics-Informed Neural Network Architecture
标题: 点神经元学习:一种新的物理信息神经网络架构
作者:Hanwen Bi,Thushara D. Abhayapala
备注:under the review process of EURASIP Journal on Audio, Speech, and Music Processing
链接:点击下载PDF文件
标题: 多个众包录音的音频增强:简单有效的基线
作者:Shiran Aziz,Yossi Adi,Shmuel Peleg
链接:点击下载PDF文件
摘要:随着手机的普及,事件通常由来自不同位置的多个设备记录并在社交媒体上共享。许多事件都有不同的记录。这样的记录通常是有噪声的,其中每个设备的噪声是本地的并且与其他设备无关。这种情况下,多个麦克风在未知的位置,捕捉本地,不相关的噪音,很少在文献中处理。在这项工作中,我们提出了一个简单而有效的众包音频增强方法,以消除每个输入音频信号的局部噪声。然后,平均所有清洁的源信号给出了事件的改进的音频。我们证明了我们的方法使用合成音频信号的有效性,以及现实世界的录音。这种简单的方法可以为众包音频增强建立一个新的基线,我们希望研究界能够开发出更复杂的方法。摘要:With the popularity of cellular phones, events are often recorded by multiple devices from different locations and shared on social media. Several different recordings could be found for many events. Such recordings are usually noisy, where noise for each device is local and unrelated to others. This case of multiple microphones at unknown locations, capturing local, uncorrelated noise, was rarely treated in the literature. In this work we propose a simple and effective crowdsourced audio enhancement method to remove local noises at each input audio signal. Then, averaging all cleaned source signals gives an improved audio of the event. We demonstrate the effectiveness of our method using synthetic audio signals, together with real-world recordings. This simple approach can set a new baseline for crowdsourced audio enhancement for more sophisticated methods which we hope will be developed by the research community.
【2】 Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
标题: 紧紧抓住我:语音增强的稳定编码器-解码器设计
作者:Daniel Haider,Felix Perfler,Vincent Lostanlen,Martin Ehler,Peter Balazs
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:具有1-D滤波器的卷积层通常用作前端来编码音频信号。与固定的时频表示不同,它们可以适应输入数据的局部特征。然而,原始音频上的1-D滤波器很难训练,并且经常遭受不稳定性。在本文中,我们解决这些问题的混合解决方案,即,结合理论驱动和数据驱动的方法。首先,我们通过听觉滤波器组对音频信号进行预处理,保证学习编码器的良好频率定位。其次,我们使用框架理论的结果来定义一个无监督的学习目标,鼓励节能和完美的重建。第三,我们适应混合压缩谱范数作为学习目标的编码器系数。在低复杂度的编码器-掩码-解码器模型中使用这些解决方案显著地改善了语音增强中的语音质量的感知评估(PESQ)。摘要:Convolutional layers with 1-D filters are often used as frontend to encode audio signals. Unlike fixed time-frequency representations, they can adapt to the local characteristics of input data. However, 1-D filters on raw audio are hard to train and often suffer from instabilities. In this paper, we address these problems with hybrid solutions, i.e., combining theory-driven and data-driven approaches. First, we preprocess the audio signals via a auditory filterbank, guaranteeing good frequency localization for the learned encoder. Second, we use results from frame theory to define an unsupervised learning objective that encourages energy conservation and perfect reconstruction. Third, we adapt mixed compressed spectral norms as learning objectives to the encoder coefficients. Using these solutions in a low-complexity encoder-mask-decoder model significantly improves the perceptual evaluation of speech quality (PESQ) in speech enhancement.
【3】 AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
标题: AASIST 3:KAN增强AASIST语音Deepfake检测使用SSL功能和额外正规化应对ASVspoof 2024挑战
作者:Kirill Borodin,Vasiliy Kudryavtsev,Dmitrii Korzh,Alexey Efimenko,Grach Mkrtchian,Mikhail Gorodnichev,Oleg Y. Rogov
备注:8 pages, 2 figures, 2 tables. Accepted paper at the ASVspoof 2024 (the 25th Interspeech Conference)
链接:点击下载PDF文件
摘要:自动说话者验证(ASV)系统根据说话者的语音特征识别说话者,具有众多应用,例如金融交易中的用户身份验证、智能设备中的独家访问控制和取证欺诈检测。然而,深度学习算法的进步使得通过文本到语音(TTS)和语音转换(VC)系统生成合成音频成为可能,从而使ASV系统面临潜在的漏洞。为了解决这个问题,我们提出了一个新的架构AASIST 3。通过使用Kolmogorov-Arnold网络、附加层、编码器和预加重技术来增强现有的AASIST框架,AASIST 3在性能上实现了两倍以上的改进。它展示了在关闭条件下为0.5357的minDCF结果和在打开条件下为0.1414的minDCF结果,显著增强了对合成语音的检测并提高了ASV安全性。摘要:Automatic Speaker Verification (ASV) systems, which identify speakers based on their voice characteristics, have numerous applications, such as user authentication in financial transactions, exclusive access control in smart devices, and forensic fraud detection. However, the advancement of deep learning algorithms has enabled the generation of synthetic audio through Text-to-Speech (TTS) and Voice Conversion (VC) systems, exposing ASV systems to potential vulnerabilities. To counteract this, we propose a novel architecture named AASIST3. By enhancing the existing AASIST framework with Kolmogorov-Arnold networks, additional layers, encoders, and pre-emphasis techniques, AASIST3 achieves a more than twofold improvement in performance. It demonstrates minDCF results of 0.5357 in the closed condition and 0.1414 in the open condition, significantly enhancing the detection of synthetic voices and improving ASV security.
【4】 Utilizing Speaker Profiles for Impersonation Audio Detection
标题: 利用扬声器配置文件进行模拟音频检测
作者:Hao Gu,JiangYan Yi,Chenglong Wang,Yong Ren,Jianhua Tao,Xinrui Yan,Yujie Chen,Xiaohui Zhang
备注:Accepted by ACM MM2024
链接:点击下载PDF文件
摘要:虚假音频检测是一个新兴的活跃话题。越来越多的文献致力于检测虚假话语,这些虚假话语大多是通过文本到语音(TTS)或语音转换(VC)产生的。然而,针对假冒的对策仍然是一个未充分探索的领域。模仿是指模仿者模仿目标说话人的特定特征和说话风格。与TTS和VC不同,它们通常会留下数字痕迹或信号伪影,模仿涉及真人产生完全自然的语音,从而使模仿音频的检测成为一项具有挑战性的任务。因此,我们提出了一种新的方法,将扬声器配置文件的过程中的模仿音频检测。说话人特征是指说话人的年龄、职业等固有特征,这些特征对模仿者的准确模仿具有挑战性。我们的目标是利用这些功能来提取识别信息,用于检测模仿音频。此外,目前还没有大规模的模仿语料库可用于模仿影响的定量研究。为了解决这个问题,我们进一步设计了第一个大规模的,不同的扬声器中文模仿数据集,命名为模仿音频检测(iPad),以推进社区对模仿音频检测的研究。我们在我们提出的数据集iPad上评估了几种现有的虚假音频检测方法,证明了其必要性和挑战。此外,我们的研究结果表明,将扬声器配置文件可以显着提高模型的性能,在检测模仿音频。摘要:Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice conversion (VC). However, countermeasures against impersonation remain an underexplored area. Impersonation is a fake type that involves an imitator replicating specific traits and speech style of a target speaker. Unlike TTS and VC, which often leave digital traces or signal artifacts, impersonation involves live human beings producing entirely natural speech, rendering the detection of impersonation audio a challenging task. Thus, we propose a novel method that integrates speaker profiles into the process of impersonation audio detection. Speaker profiles are inherent characteristics that are challenging for impersonators to mimic accurately, such as speaker's age, job. We aim to leverage these features to extract discriminative information for detecting impersonation audio. Moreover, there is no large impersonated speech corpora available for quantitative study of impersonation impacts. To address this gap, we further design the first large-scale, diverse-speaker Chinese impersonation dataset, named ImPersonation Audio Detection (IPAD), to advance the community's research on impersonation audio detection. We evaluate several existing fake audio detection methods on our proposed dataset IPAD, demonstrating its necessity and the challenges. Additionally, our findings reveal that incorporating speaker profiles can significantly enhance the model's performance in detecting impersonation audio.
【5】 Point Neuron Learning: A New Physics-Informed Neural Network Architecture
标题: 点神经元学习:一种新的物理信息神经网络架构
作者:Hanwen Bi,Thushara D. Abhayapala
备注:under the review process of EURASIP Journal on Audio, Speech, and Music Processing
链接:点击下载PDF文件
摘要:机器学习和神经网络已经推进了许多研究领域,但诸如大量训练数据需求和不一致的模型性能等挑战阻碍了它们在某些科学问题中的应用。为了克服这些挑战,研究人员研究了将物理原理集成到机器学习模型中,主要通过:(i)物理指导的损失函数,通常称为物理信息神经网络,以及(ii)物理指导的架构设计。虽然这两种方法都在多个科学学科中取得了成功,但它们都有局限性,包括被困在局部最小值,可解释性差和有限的普遍性。本文提出了一种新的物理信息神经网络(PINN)架构,通过将波动方程的基本解嵌入到网络架构中,结合了两种方法的优点,使学习模型严格满足波动方程。所提出的点神经元学习方法可以在没有任何数据集的情况下基于麦克风观测对任意声场进行建模。与其他PINN方法相比,我们的方法直接处理复数,并提供更好的可解释性和通用性。我们通过混响环境中的声场重建问题来评估所提出的架构的多功能性。结果表明,点神经元方法优于两个竞争的方法,可以有效地处理噪声环境与稀疏麦克风观察。摘要:Machine learning and neural networks have advanced numerous research domains, but challenges such as large training data requirements and inconsistent model performance hinder their application in certain scientific problems. To overcome these challenges, researchers have investigated integrating physics principles into machine learning models, mainly through: (i) physics-guided loss functions, generally termed as physics-informed neural networks, and (ii) physics-guided architectural design. While both approaches have demonstrated success across multiple scientific disciplines, they have limitations including being trapped to a local minimum, poor interpretability, and restricted generalizability. This paper proposes a new physics-informed neural network (PINN) architecture that combines the strengths of both approaches by embedding the fundamental solution of the wave equation into the network architecture, enabling the learned model to strictly satisfy the wave equation. The proposed point neuron learning method can model an arbitrary sound field based on microphone observations without any dataset. Compared to other PINN methods, our approach directly processes complex numbers and offers better interpretability and generalizability. We evaluate the versatility of the proposed architecture by a sound field reconstruction problem in a reverberant environment. Results indicate that the point neuron method outperforms two competing methods and can efficiently handle noisy environments with sparse microphone observations.
【6】 Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
标题: 编解码器确实很重要:探索音频语言模型编解码器的语义缺陷
作者:Zhen Ye,Peiwen Sun,Jiahe Lei,Hongzhan Lin,Xu Tan,Zheqi Dai,Qiuqiang Kong,Jianyi Chen,Jiahao Pan,Qifeng Liu,Yike Guo,Wei Xue
链接:点击下载PDF文件
摘要:音频生成的最新进展受到大型语言模型(LLM)功能的极大推动。关于音频LLM的现有研究主要集中在增强音频语言模型的架构和规模,以及利用更大的数据集,并且通常,声学编解码器(例如EnCodec)用于音频标记化。然而,这些编解码器最初被设计用于音频压缩,这可能导致在音频LLM的上下文中的次优性能。我们的研究旨在解决当前音频LLM编解码器的缺点,特别是在生成的音频中保持语义完整性方面的挑战。例如,现有的方法,如VALL-E,其条件声令牌生成文本transmittance,往往遭受内容不准确和提高字错误率(WER)由于语义误解的声令牌,导致单词跳过和错误。为了克服这些问题,我们提出了一种简单而有效的方法,称为X-Codec。X-Codec在残差矢量量化(RVQ)阶段之前合并来自预训练的语义编码器的语义特征,并在RVQ之后引入语义重建损失。通过增强编解码器的语义能力,X-Codec显著降低了语音合成任务中的WER,并将这些优势扩展到非语音应用,包括音乐和声音生成。我们在文本到语音,音乐延续和文本到声音任务的实验表明,整合语义信息大大提高了音频生成中的语言模型的整体性能。我们的代码和demo都是可用的(Demo:https: x-codec-audio.github.io Code:https: github.com zhenye234 xcodec)摘要:Recent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and scale of audio language models, as well as leveraging larger datasets, and generally, acoustic codecs, such as EnCodec, are used for audio tokenization. However, these codecs were originally designed for audio compression, which may lead to suboptimal performance in the context of audio LLM. Our research aims to address the shortcomings of current audio LLM codecs, particularly their challenges in maintaining semantic integrity in generated audio. For instance, existing methods like VALL-E, which condition acoustic token generation on text transcriptions, often suffer from content inaccuracies and elevated word error rates (WER) due to semantic misinterpretations of acoustic tokens, resulting in word skipping and errors. To overcome these issues, we propose a straightforward yet effective approach called X-Codec. X-Codec incorporates semantic features from a pre-trained semantic encoder before the Residual Vector Quantization (RVQ) stage and introduces a semantic reconstruction loss after RVQ. By enhancing the semantic ability of the codec, X-Codec significantly reduces WER in speech synthesis tasks and extends these benefits to non-speech applications, including music and sound generation. Our experiments in text-to-speech, music continuation, and text-to-sound tasks demonstrate that integrating semantic information substantially improves the overall performance of language models in audio generation. Our code and demo are available (Demo: https: x-codec-audio.github.io Code: https: github.com zhenye234 xcodec)
【7】 Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
标题: 从多说话人录音中提取说话人嵌入的循环注意池
作者:Shota Horiguchi,Atsushi Ando,Takafumi Moriya,Takanori Ashihara,Hiroshi Sato,Naohiro Tawara,Marc Delcroix
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:本文提出了一种从包含多个说话人的变长录音中提取每个说话人的说话人嵌入的方法。说话人嵌入不仅在说话人识别中有着重要的应用,而且在多说话人语音处理中也有着重要的应用。尽管在多说话人场景中获得单个说话人的语音而不进行预注册的挑战,但大多数关于说话人嵌入提取的研究都集中在仅从单个说话人录音中提取嵌入。已经提出了一些方法用于直接从多说话者录音中提取说话者嵌入,但是它们通常需要为每个可能数量的说话者准备模型或者涉及复杂的训练过程。所提出的方法通过关注从输入多扬声器音频中提取的逐帧嵌入的不同部分来计算多扬声器的嵌入。这是通过递归计算注意力权重来池化逐帧嵌入来实现的。此外,我们建议使用计算出的注意力权重来估计录音中的扬声器数量,这允许将相同的模型应用于不同数量的扬声器。实验结果表明,该方法在说话人确认和日记化任务中的有效性。摘要:This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not only for speaker recognition but also for various multi-speaker speech applications such as speaker diarization and target-speaker speech processing. Despite the challenges of obtaining a single speaker's speech without pre-registration in multi-speaker scenarios, most studies on speaker embedding extraction focus on extracting embeddings only from single-speaker recordings. Some methods have been proposed for extracting speaker embeddings directly from multi-speaker recordings, but they typically require preparing a model for each possible number of speakers or involve complicated training procedures. The proposed method computes the embeddings of multiple speakers by focusing on different parts of the frame-wise embeddings extracted from the input multi-speaker audio. This is achieved by recursively computing attention weights for pooling the frame-wise embeddings. Additionally, we propose using the calculated attention weights to estimate the number of speakers in the recording, which allows the same model to be applied to various numbers of speakers. Experimental evaluations demonstrate the effectiveness of the proposed method in speaker verification and diarization tasks.
【8】 User-Driven Voice Generation and Editing through Latent Space Navigation
标题: 通过潜在空间导航的用户驱动语音生成和编辑
作者:Yusheng Tian,Junbin Liu,Tan Lee
链接:点击下载PDF文件
摘要:本文提出了一种基于用户反馈的用户驱动的方法来合成高度特定的目标声音,这对于希望重建丢失的声音但缺乏先前记录的言语障碍者特别有益。具体来说,我们利用神经分析和合成框架来构建一个低维的,但足够表达的潜在说话人嵌入空间。在这个潜在的空间内,我们实现了一个搜索算法,通过完成一系列简单的比较任务,引导用户找到他们想要的声音。仿真实验和真实用户实验表明,该方法能够有效地逼近目标语音。此外,通过分析梅尔频谱生成器的雅可比,我们确定了一组有意义的语音编辑方向内的潜在空间。这些指示使用户能够进一步微调所生成的语音的特定属性,包括音高水平、音高范围、音量、声音张力、鼻音和音色。音频样本可在https: myspeechprojects.github.io voicedesign 上获得。摘要:This paper presents a user-driven approach for synthesizing highly specific target voices based on user feedback, which is particularly beneficial for speech-impaired individuals who wish to recreate their lost voices but lack prior recordings. Specifically, we leverage the neural analysis and synthesis framework to construct a low-dimensional, yet sufficiently expressive latent speaker embedding space. Within this latent space, we implement a search algorithm that guides users to their desired voice through completing a sequence of straightforward comparison tasks. Both synthetic simulations and real-world user studies demonstrate that the proposed approach can effectively approximate target voices. Moreover, by analyzing the mel-spectrogram generator's Jacobians, we identify a set of meaningful voice editing directions within the latent space. These directions enable users to further fine-tune specific attributes of the generated voice, including the pitch level, pitch range, volume, vocal tension, nasality, and tone color. Audio samples are available at https: myspeechprojects.github.io voicedesign .
eess.AS音频处理
【1】 SelectTTS: Synthesizing Anyone's Voice via Discrete Unit-Based Frame Selection标题: SelectTTC:通过基于离散单元的帧选择合成任何人的语音
作者:Ismail Rasim Ulgen,Shreeram Suresh Chandra,Junchen Lu,Berrak Sisman
备注:Submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:合成看不见的说话者的声音是多说话者文本到语音(TTS)中的一个持续挑战。大多数多说话人TTS模型依赖于在训练期间通过说话人条件反射来建模说话人特征。通过这种方法对看不见的说话者属性进行建模需要增加模型复杂性,这使得重现结果并对其进行改进具有挑战性。我们设计了一个简单的替代方案。我们提出了SelectTTS,一种新的方法来选择合适的帧从目标扬声器和解码使用帧级自监督学习(SSL)功能。我们表明,这种方法可以有效地捕捉扬声器的特征看不见的扬声器,并取得了可比的结果,其他多扬声器TTS框架在客观和主观指标。与SelectTTS,我们表明,从目标扬声器的语音帧选择是一种直接的方式来实现推广,在看不见的扬声器具有低模型复杂度。我们实现了比SOTA基线XTTS-v2和VALL-E更好的说话人相似性性能,模型参数减少了8倍以上,训练数据减少了270倍摘要:Synthesizing the voices of unseen speakers is a persisting challenge in multi-speaker text-to-speech (TTS). Most multi-speaker TTS models rely on modeling speaker characteristics through speaker conditioning during training. Modeling unseen speaker attributes through this approach has necessitated an increase in model complexity, which makes it challenging to reproduce results and improve upon them. We design a simple alternative to this. We propose SelectTTS, a novel method to select the appropriate frames from the target speaker and decode using frame-level self-supervised learning (SSL) features. We show that this approach can effectively capture speaker characteristics for unseen speakers, and achieves comparable results to other multi-speaker TTS frameworks in both objective and subjective metrics. With SelectTTS, we show that frame selection from the target speaker's speech is a direct way to achieve generalization in unseen speakers with low model complexity. We achieve better speaker similarity performance than SOTA baselines XTTS-v2 and VALL-E with over an 8x reduction in model parameters and a 270x reduction in training data
【2】 Advancing Multi-talker ASR Performance with Large Language Models
标题: 利用大型语言模型提高多说话者ASB性能
作者:Mohan Shi,Zengrui Jin,Yaoxun Xu,Yong Xu,Shi-Xiong Zhang,Kun Wei,Yiwen Shao,Chunlei Zhang,Dong Yu
备注:8 pages, accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:在会话场景中识别多个说话人的重叠语音是自动语音识别(ASR)中最具挑战性的问题之一。串行输出训练(SOT)是解决多说话者ASR的经典方法,其思想是根据多个说话者的语音的发射时间将来自多个说话者的传输串接起来进行训练。然而,SOT风格的transmittance,来自连接多个相关的话语在一个会话中,显着依赖于建模长的上下文。因此,与主要强调基于注意力的编码器-解码器(AED)架构中的编码器性能的传统方法相比,利用利用预训练的解码器的能力的大型语言模型(LLM)的新颖方法可以更好地适合于这种复杂且具有挑战性的场景。在本文中,我们提出了一种基于LLM的SOT方法用于多人ASR,利用预训练的语音编码器和LLM,使用适当的策略在多人数据集上对其进行微调。实验结果表明,我们的方法在模拟数据集LibriMix上超越了传统的基于AED的方法,并在真实数据集AMI的评估集上实现了最先进的性能,优于在以前的作品中使用1000倍以上的监督数据训练的AED模型。摘要:Recognizing overlapping speech from multiple speakers in conversational scenarios is one of the most challenging problem for automatic speech recognition (ASR). Serialized output training (SOT) is a classic method to address multi-talker ASR, with the idea of concatenating transcriptions from multiple speakers according to the emission times of their speech for training. However, SOT-style transcriptions, derived from concatenating multiple related utterances in a conversation, depend significantly on modeling long contexts. Therefore, compared to traditional methods that primarily emphasize encoder performance in attention-based encoder-decoder (AED) architectures, a novel approach utilizing large language models (LLMs) that leverages the capabilities of pre-trained decoders may be better suited for such complex and challenging scenarios. In this paper, we propose an LLM-based SOT approach for multi-talker ASR, leveraging pre-trained speech encoder and LLM, fine-tuning them on multi-talker dataset using appropriate strategies. Experimental results demonstrate that our approach surpasses traditional AED-based methods on the simulated dataset LibriMix and achieves state-of-the-art performance on the evaluation set of the real-world dataset AMI, outperforming the AED model trained with 1000 times more supervised data in previous works.
【3】 Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
标题: 编解码器确实很重要:探索音频语言模型编解码器的语义缺陷
作者:Zhen Ye,Peiwen Sun,Jiahe Lei,Hongzhan Lin,Xu Tan,Zheqi Dai,Qiuqiang Kong,Jianyi Chen,Jiahao Pan,Qifeng Liu,Yike Guo,Wei Xue
链接:点击下载PDF文件
摘要:音频生成的最新进展受到大型语言模型(LLM)功能的极大推动。关于音频LLM的现有研究主要集中在增强音频语言模型的架构和规模,以及利用更大的数据集,并且通常,声学编解码器(例如EnCodec)用于音频标记化。然而,这些编解码器最初被设计用于音频压缩,这可能导致在音频LLM的上下文中的次优性能。我们的研究旨在解决当前音频LLM编解码器的缺点,特别是在生成的音频中保持语义完整性方面的挑战。例如,现有的方法,如VALL-E,其条件声令牌生成文本transmittance,往往遭受内容不准确和提高字错误率(WER)由于语义误解的声令牌,导致单词跳过和错误。为了克服这些问题,我们提出了一种简单而有效的方法,称为X-Codec。X-Codec在残差矢量量化(RVQ)阶段之前合并来自预训练的语义编码器的语义特征,并在RVQ之后引入语义重建损失。通过增强编解码器的语义能力,X-Codec显著降低了语音合成任务中的WER,并将这些优势扩展到非语音应用,包括音乐和声音生成。我们在文本到语音,音乐延续和文本到声音任务的实验表明,整合语义信息大大提高了音频生成中的语言模型的整体性能。我们的代码和demo都是可用的(Demo:https: x-codec-audio.github.io Code:https: github.com zhenye234 xcodec)摘要:Recent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and scale of audio language models, as well as leveraging larger datasets, and generally, acoustic codecs, such as EnCodec, are used for audio tokenization. However, these codecs were originally designed for audio compression, which may lead to suboptimal performance in the context of audio LLM. Our research aims to address the shortcomings of current audio LLM codecs, particularly their challenges in maintaining semantic integrity in generated audio. For instance, existing methods like VALL-E, which condition acoustic token generation on text transcriptions, often suffer from content inaccuracies and elevated word error rates (WER) due to semantic misinterpretations of acoustic tokens, resulting in word skipping and errors. To overcome these issues, we propose a straightforward yet effective approach called X-Codec. X-Codec incorporates semantic features from a pre-trained semantic encoder before the Residual Vector Quantization (RVQ) stage and introduces a semantic reconstruction loss after RVQ. By enhancing the semantic ability of the codec, X-Codec significantly reduces WER in speech synthesis tasks and extends these benefits to non-speech applications, including music and sound generation. Our experiments in text-to-speech, music continuation, and text-to-sound tasks demonstrate that integrating semantic information substantially improves the overall performance of language models in audio generation. Our code and demo are available (Demo: https: x-codec-audio.github.io Code: https: github.com zhenye234 xcodec)
【4】 Learning Multi-Target TDOA Features for Sound Event Localization and Detection
标题: 学习多目标TDOE特征以进行声音事件定位和检测
作者:Axel Berg,Johanna Engman,Jens Gulin,Karl Åström,Magnus Oskarsson
备注:DCASE 2024
链接:点击下载PDF文件
摘要:使用来自麦克风阵列的音频记录的声音事件定位和检测(SELD)系统依赖于用于确定声音事件的位置的空间线索。因此,这种系统的定位性能在很大程度上由用作系统输入的音频特征的质量确定。我们提出了一个新的功能,基于神经广义互相关与相位变换(NGCC-PHAT),学习适合本地化的音频表示。使用排列不变训练的到达时间差(TDOA)估计问题,使NGCC-PHAT学习多个重叠的声音事件的TDOA特征。这些特性可用作SELD网络GCC-PHAT输入的直接替代。我们在STARSS 23数据集上测试了我们的方法,并证明了与使用标准GCC-PHAT或SALSA-Lite输入功能相比,改进的本地化性能。摘要:Sound event localization and detection (SELD) systems using audio recordings from a microphone array rely on spatial cues for determining the location of sound events. As a consequence, the localization performance of such systems is to a large extent determined by the quality of the audio features that are used as inputs to the system. We propose a new feature, based on neural generalized cross-correlations with phase-transform (NGCC-PHAT), that learns audio representations suitable for localization. Using permutation invariant training for the time-difference of arrival (TDOA) estimation problem enables NGCC-PHAT to learn TDOA features for multiple overlapping sound events. These features can be used as a drop-in replacement for GCC-PHAT inputs to a SELD-network. We test our method on the STARSS23 dataset and demonstrate improved localization performance compared to using standard GCC-PHAT or SALSA-Lite input features.
【5】 Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
标题: 从多说话人录音中提取说话人嵌入的循环注意池
作者:Shota Horiguchi,Atsushi Ando,Takafumi Moriya,Takanori Ashihara,Hiroshi Sato,Naohiro Tawara,Marc Delcroix
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:本文提出了一种从包含多个说话人的变长录音中提取每个说话人的说话人嵌入的方法。说话人嵌入不仅在说话人识别中有着重要的应用,而且在多说话人语音处理中也有着重要的应用。尽管在多说话人场景中获得单个说话人的语音而不进行预注册的挑战,但大多数关于说话人嵌入提取的研究都集中在仅从单个说话人录音中提取嵌入。已经提出了一些方法用于直接从多说话者录音中提取说话者嵌入,但是它们通常需要为每个可能数量的说话者准备模型或者涉及复杂的训练过程。所提出的方法通过关注从输入多扬声器音频中提取的逐帧嵌入的不同部分来计算多扬声器的嵌入。这是通过递归计算注意力权重来池化逐帧嵌入来实现的。此外,我们建议使用计算出的注意力权重来估计录音中的扬声器数量,这允许将相同的模型应用于不同数量的扬声器。实验结果表明,该方法在说话人确认和日记化任务中的有效性。摘要:This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not only for speaker recognition but also for various multi-speaker speech applications such as speaker diarization and target-speaker speech processing. Despite the challenges of obtaining a single speaker's speech without pre-registration in multi-speaker scenarios, most studies on speaker embedding extraction focus on extracting embeddings only from single-speaker recordings. Some methods have been proposed for extracting speaker embeddings directly from multi-speaker recordings, but they typically require preparing a model for each possible number of speakers or involve complicated training procedures. The proposed method computes the embeddings of multiple speakers by focusing on different parts of the frame-wise embeddings extracted from the input multi-speaker audio. This is achieved by recursively computing attention weights for pooling the frame-wise embeddings. Additionally, we propose using the calculated attention weights to estimate the number of speakers in the recording, which allows the same model to be applied to various numbers of speakers. Experimental evaluations demonstrate the effectiveness of the proposed method in speaker verification and diarization tasks.
【6】 User-Driven Voice Generation and Editing through Latent Space Navigation
标题: 通过潜在空间导航的用户驱动语音生成和编辑
作者:Yusheng Tian,Junbin Liu,Tan Lee
链接:点击下载PDF文件
摘要:本文提出了一种基于用户反馈的用户驱动的方法来合成高度特定的目标声音,这对于希望重建丢失的声音但缺乏先前记录的言语障碍者特别有益。具体来说,我们利用神经分析和合成框架来构建一个低维的,但足够表达的潜在说话人嵌入空间。在这个潜在的空间内,我们实现了一个搜索算法,通过完成一系列简单的比较任务,引导用户找到他们想要的声音。仿真实验和真实用户实验表明,该方法能够有效地逼近目标语音。此外,通过分析梅尔频谱生成器的雅可比,我们确定了一组有意义的语音编辑方向内的潜在空间。这些指示使用户能够进一步微调所生成的语音的特定属性,包括音高水平、音高范围、音量、声音张力、鼻音和音色。音频样本可在https: myspeechprojects.github.io voicedesign 上获得。摘要:This paper presents a user-driven approach for synthesizing highly specific target voices based on user feedback, which is particularly beneficial for speech-impaired individuals who wish to recreate their lost voices but lack prior recordings. Specifically, we leverage the neural analysis and synthesis framework to construct a low-dimensional, yet sufficiently expressive latent speaker embedding space. Within this latent space, we implement a search algorithm that guides users to their desired voice through completing a sequence of straightforward comparison tasks. Both synthetic simulations and real-world user studies demonstrate that the proposed approach can effectively approximate target voices. Moreover, by analyzing the mel-spectrogram generator's Jacobians, we identify a set of meaningful voice editing directions within the latent space. These directions enable users to further fine-tune specific attributes of the generated voice, including the pitch level, pitch range, volume, vocal tension, nasality, and tone color. Audio samples are available at https: myspeechprojects.github.io voicedesign .
【7】 Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
标题: 多个众包录音的音频增强:简单有效的基线
作者:Shiran Aziz,Yossi Adi,Shmuel Peleg
链接:点击下载PDF文件
摘要:随着手机的普及,事件通常由来自不同位置的多个设备记录并在社交媒体上共享。许多事件都有不同的记录。这样的记录通常是有噪声的,其中每个设备的噪声是本地的并且与其他设备无关。这种情况下,多个麦克风在未知的位置,捕捉本地,不相关的噪音,很少在文献中处理。在这项工作中,我们提出了一个简单而有效的众包音频增强方法,以消除每个输入音频信号的局部噪声。然后,平均所有清洁的源信号给出了事件的改进的音频。我们证明了我们的方法使用合成音频信号的有效性,以及现实世界的录音。这种简单的方法可以为众包音频增强建立一个新的基线,我们希望研究界能够开发出更复杂的方法。摘要:With the popularity of cellular phones, events are often recorded by multiple devices from different locations and shared on social media. Several different recordings could be found for many events. Such recordings are usually noisy, where noise for each device is local and unrelated to others. This case of multiple microphones at unknown locations, capturing local, uncorrelated noise, was rarely treated in the literature. In this work we propose a simple and effective crowdsourced audio enhancement method to remove local noises at each input audio signal. Then, averaging all cleaned source signals gives an improved audio of the event. We demonstrate the effectiveness of our method using synthetic audio signals, together with real-world recordings. This simple approach can set a new baseline for crowdsourced audio enhancement for more sophisticated methods which we hope will be developed by the research community.
【8】 Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
标题: 紧紧抓住我:语音增强的稳定编码器-解码器设计
作者:Daniel Haider,Felix Perfler,Vincent Lostanlen,Martin Ehler,Peter Balazs
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:具有1-D滤波器的卷积层通常用作前端来编码音频信号。与固定的时频表示不同,它们可以适应输入数据的局部特征。然而,原始音频上的1-D滤波器很难训练,并且经常遭受不稳定性。在本文中,我们解决这些问题的混合解决方案,即,结合理论驱动和数据驱动的方法。首先,我们通过听觉滤波器组对音频信号进行预处理,保证学习编码器的良好频率定位。其次,我们使用框架理论的结果来定义一个无监督的学习目标,鼓励节能和完美的重建。第三,我们适应混合压缩谱范数作为学习目标的编码器系数。在低复杂度的编码器-掩码-解码器模型中使用这些解决方案显著地改善了语音增强中的语音质量的感知评估(PESQ)。摘要:Convolutional layers with 1-D filters are often used as frontend to encode audio signals. Unlike fixed time-frequency representations, they can adapt to the local characteristics of input data. However, 1-D filters on raw audio are hard to train and often suffer from instabilities. In this paper, we address these problems with hybrid solutions, i.e., combining theory-driven and data-driven approaches. First, we preprocess the audio signals via a auditory filterbank, guaranteeing good frequency localization for the learned encoder. Second, we use results from frame theory to define an unsupervised learning objective that encourages energy conservation and perfect reconstruction. Third, we adapt mixed compressed spectral norms as learning objectives to the encoder coefficients. Using these solutions in a low-complexity encoder-mask-decoder model significantly improves the perceptual evaluation of speech quality (PESQ) in speech enhancement.
【9】 AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
标题: AASIST 3:KAN增强AASIST语音Deepfake检测使用SSL功能和额外正规化应对ASVspoof 2024挑战
作者:Kirill Borodin,Vasiliy Kudryavtsev,Dmitrii Korzh,Alexey Efimenko,Grach Mkrtchian,Mikhail Gorodnichev,Oleg Y. Rogov
备注:8 pages, 2 figures, 2 tables. Accepted paper at the ASVspoof 2024 (the 25th Interspeech Conference)
链接:点击下载PDF文件
摘要:自动说话人验证(ASV)系统基于说话人的语音特征来识别说话人,具有许多应用,例如金融交易中的用户身份验证、智能设备中的独占访问控制和取证欺诈检测。然而,深度学习算法的进步使得通过文本到语音(TTS)和语音转换(VC)系统生成合成音频成为可能,从而使ASV系统面临潜在的漏洞。为了解决这个问题,我们提出了一个新的架构AASIST 3。通过使用Kolmogorov-Arnold网络、附加层、编码器和预加重技术来增强现有的AASIST框架,AASIST 3在性能上实现了两倍以上的改进。它展示了在关闭条件下为0.5357的minDCF结果和在打开条件下为0.1414的minDCF结果,显著增强了对合成语音的检测并提高了ASV安全性。摘要:Automatic Speaker Verification (ASV) systems, which identify speakers based on their voice characteristics, have numerous applications, such as user authentication in financial transactions, exclusive access control in smart devices, and forensic fraud detection. However, the advancement of deep learning algorithms has enabled the generation of synthetic audio through Text-to-Speech (TTS) and Voice Conversion (VC) systems, exposing ASV systems to potential vulnerabilities. To counteract this, we propose a novel architecture named AASIST3. By enhancing the existing AASIST framework with Kolmogorov-Arnold networks, additional layers, encoders, and pre-emphasis techniques, AASIST3 achieves a more than twofold improvement in performance. It demonstrates minDCF results of 0.5357 in the closed condition and 0.1414 in the open condition, significantly enhancing the detection of synthetic voices and improving ASV security.
【10】 Utilizing Speaker Profiles for Impersonation Audio Detection
标题: 利用扬声器配置文件进行模拟音频检测
作者:Hao Gu,JiangYan Yi,Chenglong Wang,Yong Ren,Jianhua Tao,Xinrui Yan,Yujie Chen,Xiaohui Zhang
备注:Accepted by ACM MM2024
链接:点击下载PDF文件
摘要:虚假音频检测是一个新兴的活跃话题。越来越多的文献致力于检测虚假话语,这些虚假话语大多是通过文本到语音(TTS)或语音转换(VC)产生的。然而,针对假冒的对策仍然是一个未充分探索的领域。模仿是指模仿者模仿目标说话人的特定特征和说话风格。与TTS和VC不同,它们通常会留下数字痕迹或信号伪影,模仿涉及真人产生完全自然的语音,从而使模仿音频的检测成为一项具有挑战性的任务。因此,我们提出了一种新的方法,将扬声器配置文件的过程中的模仿音频检测。说话人特征是指说话人的年龄、职业等固有特征,这些特征对模仿者的准确模仿具有挑战性。我们的目标是利用这些功能来提取识别信息,用于检测模仿音频。此外,目前还没有大规模的模仿语料库可用于模仿影响的定量研究。为了解决这一问题,我们进一步设计了第一个大规模的,不同说话人的中文模仿数据集,命名为模仿音频检测(iPad),以推进社区对模仿音频检测的研究。我们在我们提出的数据集iPad上评估了几种现有的虚假音频检测方法,证明了其必要性和挑战。此外,我们的研究结果表明,将扬声器配置文件可以显着提高模型的性能,在检测模仿音频。摘要:Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice conversion (VC). However, countermeasures against impersonation remain an underexplored area. Impersonation is a fake type that involves an imitator replicating specific traits and speech style of a target speaker. Unlike TTS and VC, which often leave digital traces or signal artifacts, impersonation involves live human beings producing entirely natural speech, rendering the detection of impersonation audio a challenging task. Thus, we propose a novel method that integrates speaker profiles into the process of impersonation audio detection. Speaker profiles are inherent characteristics that are challenging for impersonators to mimic accurately, such as speaker's age, job. We aim to leverage these features to extract discriminative information for detecting impersonation audio. Moreover, there is no large impersonated speech corpora available for quantitative study of impersonation impacts. To address this gap, we further design the first large-scale, diverse-speaker Chinese impersonation dataset, named ImPersonation Audio Detection (IPAD), to advance the community's research on impersonation audio detection. We evaluate several existing fake audio detection methods on our proposed dataset IPAD, demonstrating its necessity and the challenges. Additionally, our findings reveal that incorporating speaker profiles can significantly enhance the model's performance in detecting impersonation audio.
【11】 Point Neuron Learning: A New Physics-Informed Neural Network Architecture
标题: 点神经元学习:一种新的物理信息神经网络架构
作者:Hanwen Bi,Thushara D. Abhayapala
备注:under the review process of EURASIP Journal on Audio, Speech, and Music Processing
链接:点击下载PDF文件
摘要:机器学习和神经网络已经推进了许多研究领域,但诸如大量训练数据需求和不一致的模型性能等挑战阻碍了它们在某些科学问题中的应用。为了克服这些挑战,研究人员研究了将物理原理集成到机器学习模型中,主要通过:(i)物理指导的损失函数,通常称为物理信息神经网络,以及(ii)物理指导的架构设计。虽然这两种方法都在多个科学学科中取得了成功,但它们都有局限性,包括被困在局部最小值,可解释性差和有限的普遍性。本文提出了一种新的物理信息神经网络(PINN)架构,通过将波动方程的基本解嵌入到网络架构中,结合了两种方法的优点,使学习的模型能够严格满足波动方程。所提出的点神经元学习方法可以在没有任何数据集的情况下基于麦克风观测对任意声场进行建模。与其他PINN方法相比,我们的方法直接处理复数,并提供更好的可解释性和通用性。我们通过混响环境中的声场重建问题来评估所提出的架构的多功能性。结果表明,点神经元方法优于两个竞争的方法,可以有效地处理噪声环境与稀疏麦克风观察。摘要:Machine learning and neural networks have advanced numerous research domains, but challenges such as large training data requirements and inconsistent model performance hinder their application in certain scientific problems. To overcome these challenges, researchers have investigated integrating physics principles into machine learning models, mainly through: (i) physics-guided loss functions, generally termed as physics-informed neural networks, and (ii) physics-guided architectural design. While both approaches have demonstrated success across multiple scientific disciplines, they have limitations including being trapped to a local minimum, poor interpretability, and restricted generalizability. This paper proposes a new physics-informed neural network (PINN) architecture that combines the strengths of both approaches by embedding the fundamental solution of the wave equation into the network architecture, enabling the learned model to strictly satisfy the wave equation. The proposed point neuron learning method can model an arbitrary sound field based on microphone observations without any dataset. Compared to other PINN methods, our approach directly processes complex numbers and offers better interpretability and generalizability. We evaluate the versatility of the proposed architecture by a sound field reconstruction problem in a reverberant environment. Results indicate that the point neuron method outperforms two competing methods and can efficiently handle noisy environments with sparse microphone observations.
机器翻译,仅供参考
