I have been testing the model with Japanese audio inputs and have noticed that the reconstruction quality is often unstable, with a consistently low snr.
Would it be possible for you to disclose any information about the language distribution in the training data?
I have been testing the model with Japanese audio inputs and have noticed that the reconstruction quality is often unstable, with a consistently low snr.
Would it be possible for you to disclose any information about the language distribution in the training data?