<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: blogus</title><link>https://news.ycombinator.com/user?id=blogus</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 23 Sep 2026 07:02:29 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=blogus" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by blogus in "OpenAI may not use lyrics without license, German court rules"]]></title><description><![CDATA[
<p>I would think as a matter of practice AI companies would attempt to detect long strings that appeared frequently in their corpus and dedup them out. There isn’t any value in training over and over again on the same data, and the copyright danger of being able to exactly reproduce your training set is obvious. Perhaps they did it intentionally, using the ability to reproduce copyrighted material as a way to get customers early on, knowing they would have to pay a paltry fee for it later.</p>
]]></description><pubDate>Tue, 11 Nov 2025 18:19:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=45890870</link><dc:creator>blogus</dc:creator><comments>https://news.ycombinator.com/item?id=45890870</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45890870</guid></item></channel></rss>