I think the "it's just SEO" or "propaganda isn't new" comments are missing the bigger picture. Given how LLMs are effectively being treated as core processing components by so many automations, generating information to skew LLM training could be considered similar to adding a back door to an actual CPU.
For example, if someone writes a generic prompt to flag offensive language, the LLM would flag certain political views whether or not the language was offensive. Causing friction and effectively discrimination against groups of people even though it was not the writer's intention.
An even more serious example would be someone in the military screening targets and the LLM marking a school or hospital as a valid target without the user intending to.
To be clear, I'm not stating that either of these doesn't happen intentionally, but that any manipulation of LLM training data can cause these things to happen when the user did not intend it, and hence it's similar to adding a back door or corrupting a CPU.
For example, if someone writes a generic prompt to flag offensive language, the LLM would flag certain political views whether or not the language was offensive. Causing friction and effectively discrimination against groups of people even though it was not the writer's intention.
An even more serious example would be someone in the military screening targets and the LLM marking a school or hospital as a valid target without the user intending to.
To be clear, I'm not stating that either of these doesn't happen intentionally, but that any manipulation of LLM training data can cause these things to happen when the user did not intend it, and hence it's similar to adding a back door or corrupting a CPU.