· 杂谈

AI垃圾内容污染大模型检索,形成低质信息反馈循环。

中文翻译

只是一个随机想法,也许有人能创办一家初创公司来解决这个问题: 我观察到一个有趣的趋势,即人类可识别的内容农场垃圾网站,如“Ad Hoc News”。 它们每天针对特定公司发布大量旧的/重复的误导性AI新闻以获取流量。 但这些内容被AI检索视为高权重且新鲜。因此,Gemini等大语言模型(LLM)在引用这些数据时受到污染。 接着,像大学生这样没有技术背景的人使用AI撰写他们并不理解的报告。 然后LLM在训练/检索中也被这些内容污染,导致人们混淆错误的技术细节,例如将连续波分布反馈激光器(cw dfb laser)阵列的单通道与单高功率激光器的输出进行比较。 因此,低阶AI反馈进入前沿AI,形成循环。模型检索受到污染。

英文原文

Just a random thought, maybe someone can do a startup to solve this: Interesting trend I’m seeing is human-recognized content farm spam sites like “Ad Hoc News”. Flooding tens of old/repeated misleading AI news every day on specific companies for views. But it’s treated highly + fresh to AI retrieval. So LLMs like Gemini get polluted citing that data. Then you have people without technical backgrounds like college students making writeups on things they don’t understand, using AI. Then LLMs get polluted in training/retrieval with that too, so you see people conflating the wrong technical nuance like comparing single channels of a cw dfb laser array vs. single high power lasers outputs. So you have a feedback loop of low tier AI feeding into frontier AI. With model retrieval getting contaminated.

在 X 上查看原推 ↗