
The new model, called VSSFlow, leverages a creative architecture to generate sounds and speech with a single unified system, with state-of-the-art results. Watch (and hear) some demos below.
Currently, most video-to-sound models (that is, models that are trained to generate sounds from silent videos) aren’t that great at generating speech. Likewise, most text-to-speech models fail at generating non-speech sounds, since they’re designed for a different purpose



