How does audio language model hold speaker identity throughout an utterance.
the first hypothesis that comes to mind is that TTS models may have attention sink at speaker token or may have high attention score there and the model keeps looking at speaker token. but looking ...