User avatar
🇨🇦Samuel Proulx🇨🇦 @fastfinge@interfree.ca
1d
What happens if you take a sample of human speech and force an autoencoder to continue it? For those of you who are unfamiliar, an autoencoder is a type of neural network usually used to encode and reproduce speech data. In general, it would be guided as part of a text to speech model like omnivoice or kokoro. However, in this case, I just used the mimi 1.5b autoencoder on its own, gave it audio, and forced it to keep generating all by itself. It really really didn't want to do that, so I had to write a script to detect silent frames and force it to retry them. I also had to keep the original audio at the start of its context Window at all times, to force it to stop wandering off into shrieking and demonic growling. Interestingly, this can run on realtime on my GPU, so it would be possible to just set up an endless stream of this. The sample I used to prompt the autoencoder was the Star Trek computer saying "Enter when ready." What I thought would happen is it would just produce more of something that sounds like the Star Trek computer speaking in gibberish. That is...not what I got. Good grief, now I have to write alt text for this.
8
14
12
0
User avatar
Andre Louis @FreakyFwoof@universeodon.com
1d
@fastfinge I really really really want to try this.
2
0
0
0
User avatar
Derek Roberts @tcikoritys@failbox.xyz
1d
@FreakyFwoof @fastfinge Me as well. I wonder how much of that is broken training data or similar.
1
0
1
0
User avatar
🇨🇦Samuel Proulx🇨🇦 @fastfinge@interfree.ca
19h
@tcikoritys Here you go. It is really an MVP. Documentation only sort of exists. It's only tested to run on my one Windows machine. If it breaks, you get to keep all the pieces: github.com/fastfinge/audai
0
0
1
0