This is not a SFW instance. NSFW communities will be tagged, and efforts will be made to keep SFW communities clean, but proceed at your own comfort level.
I think there is a very narrow, specific implementation of LLMs for NPCs that could be used for games that I personally would deem as both ethical and acceptable:
The model would need to be trained exclusively on a deep write-up of source materials authored by human beings (something like what Patrick Stewart was given for his role in Oblivion); and human voice actors hired for all characters, who explicitly approve their use towards generative characterisation solely for this one role/game.
This is just my personal line in the sand, and I know not everyone would agree - and that’s fair enough.
Every time I read something like this I realize just how little people understand how these things work.
The tiniest LLMs you can run on your machine are 1B, those take around 10TB of unrepeatable and well varied text to train. Let me put this into perspective, if you downloaded the whole Wikipedia, you would get 1% of the amount of data needed to train the tiniest LLM you can run. Do you think you can ethically source the remaining 99%?
I think there is a very narrow, specific implementation of LLMs for NPCs that could be used for games that I personally would deem as both ethical and acceptable:
The model would need to be trained exclusively on a deep write-up of source materials authored by human beings (something like what Patrick Stewart was given for his role in Oblivion); and human voice actors hired for all characters, who explicitly approve their use towards generative characterisation solely for this one role/game.
This is just my personal line in the sand, and I know not everyone would agree - and that’s fair enough.
Every time I read something like this I realize just how little people understand how these things work.
The tiniest LLMs you can run on your machine are 1B, those take around 10TB of unrepeatable and well varied text to train. Let me put this into perspective, if you downloaded the whole Wikipedia, you would get 1% of the amount of data needed to train the tiniest LLM you can run. Do you think you can ethically source the remaining 99%?