TorchDIVA

TorchDIVA implements computational simulations of speech motor control by reconstructing the DIVA (Directions Into Velocities of Articulators) model in Python using PyTorch tensors to enable neurobiological modeling and integration with generative vocoders such as DiffWave.


Key Features:

  • Matlab-to-Python Reimplementation: The original DIVA Matlab/Simulink model was translated to Python and reimplemented using PyTorch tensors while preserving original functionalities.
  • Block-by-Block Validation: Systematic block-by-block validation confirmed that TorchDIVA outputs closely match those of the original DIVA model with negligible differences.
  • PyTorch-based Computation: Core computations and simulations are performed with PyTorch tensors to support tensor-based operations and integration with PyTorch models.
  • Integration with DiffWave: A PyTorch generative vocoder, DiffWave, was incorporated with a modified mel-spectrum upsampler trained on human speech waveforms and conditioned on TorchDIVA speech-production outputs.
  • Demonstrated Improvement in Synthesis Quality: Conditioning DiffWave on TorchDIVA outputs produced improved speech quality metrics relative to the baseline implementation.
  • Neurobiological and Vocal Tract Modeling: The implementation simulates brain regions responsible for speech production alongside a model of the human vocal tract from the DIVA framework.

Scientific Applications:

  • Speech Motor Control Modeling: Simulation and analysis of neural mechanisms and motor control involved in speech production using the DIVA framework.
  • Speech Synthesis Research: Development and evaluation of speech synthesis methods by conditioning generative vocoders on model-derived spectrograms.
  • Model Replication and Validation: Replication of DIVA model behavior in Python for comparative studies against the original Matlab/Simulink implementation.
  • Integration of Neurobiology and Machine Learning: Experimental platform for combining neurobiological speech models with PyTorch-based machine learning and generative models.

Methodology:

The original DIVA Matlab/Simulink code was translated to Python and reimplemented with PyTorch tensors; systematic block-by-block validation compared outputs to the original model; a modified DiffWave mel-spectrum upsampler was trained on human speech waveforms and conditioned on outputs from TorchDIVA’s speech production module; the implementation simulates brain regions responsible for speech production and a human vocal tract model as conceptualized in the DIVA framework.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
MATLAB, Python
Added:
3/17/2023
Last Updated:
11/24/2024

Operations

Publications

Kinahan SP, Liss JM, Berisha V. TorchDIVA: An extensible computational model of speech production built on an open-source machine learning library. PLOS ONE. 2023;18(2):e0281306. doi:10.1371/journal.pone.0281306. PMID:36800358. PMCID:PMC9937462.

PMID: 36800358
PMCID: PMC9937462
Funding: - NIH-NIDCD: R01DC006859