completely blind computer geek, lover of science fiction and fantasy (especially LitRPG). I work in accessibility, but my opinions are my own, not that of my employer. Fandoms: Harry Potter, Discworld, My Little Pony: Friendship is Magic, Buffy, Dead Like Me, Glee, and I'll read fanfic of pretty much anything that crosses over with one of those.
keyoxide: aspe:keyoxide.org:PFAQDLXSBNO7MZRNPUMWWKQ7TQ
xmpp
fastfinge@im.interfree.ca
keyoxide
aspe:keyoxide.org:PFAQDLXSBNO7MZRNPUMWWKQ7TQ
What happens if you take a sample of human speech and force an autoencoder to continue it? For those of you who are unfamiliar, an autoencoder is a type of neural network usually used to encode and reproduce speech data. In general, it would be guided as part of a text to speech model like omnivoice or kokoro. However, in this case, I just used the mimi 1.5b autoencoder on its own, gave it audio, and forced it to keep generating all by itself. It really really didn't want to do that, so I had to write a script to detect silent frames and force it to retry them. I also had to keep the original audio at the start of its context Window at all times, to force it to stop wandering off into shrieking and demonic growling. Interestingly, this can run on realtime on my GPU, so it would be possible to just set up an endless stream of this. The sample I used to prompt the autoencoder was the Star Trek computer saying "Enter when ready." What I thought would happen is it would just produce more of something that sounds like the Star Trek computer speaking in gibberish. That is...not what I got. Good grief, now I have to write alt text for this.