NVIDIA has released a free, 100 million-parameter AI model capable of identifying and separating up to eight speakers from audio data in real-time. This model enables efficient execution in applications that require advanced audio processing.
The newly released model specializes in accurately determining "who is speaking and when" in multi-speaker environments. With a size of 100 million parameters, it is designed to operate even in environments with limited computational resources, achieving a balance between high processing capability and lightweight efficiency.
Traditional speaker diarization technologies often involve heavy processing loads, making real-time application difficult in many cases. Based on NVIDIA's optimization technologies, this model adopts an algorithm that can accurately separate up to eight simultaneous speakers even with limited computational resources. As a result, it is expected to be utilized in use cases requiring immediacy, such as automated meeting minutes generation and real-time translation tools.
The model is now publicly available, allowing developers to integrate it into their own applications. By providing this technology to the speech recognition community, NVIDIA aims to drive the evolution of next-generation voice dialogue systems and AI assistants.