NVIDIA Wants Everyone to Own a Digital Avatar with Open-Source Audio2Face Animation Model

NVIDIA is today open-sourcing its Audio2Face 3D animation model, allowing everyone to become the owner of their own digital avatar. The company has trained its Audio2Face technology to produce characters that are true to life, utilizing generative video models for facial expressions, text-to-speech models for human output, and large language models for actual conversations, all combined in a single pipeline to create a digital avatar. All of these stages are fed by analyzing acoustic input, which features characteristics such as intonation and phonemes, and are then translated into a stream of animation data. That data is later mapped to facial expressions, where the mood of the avatar can be interpreted simply by looking at it. The entire data pipeline can be either an offline scripted content or you can feed the model using live data for real-time lip-synching and facial expressions.

Additionally, NVIDIA is open-sourcing the entire Audio2Face stack, which includes the model itself and the SDK, so everyone, from game developers to home enthusiasts, can use it in any possible way. NVIDIA is now open-sourcing the entire training framework, allowing anyone to create their own avatar version with customized features. The company even provided Autodesk Maya and Unreal Engine plugins, allowing 3D workloads to integrate this generative model immediately.
Read full story