<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: omneity</title><link>https://news.ycombinator.com/user?id=omneity</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 06 Oct 2026 04:13:38 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=omneity" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by omneity in "LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents"]]></title><description><![CDATA[
<p>I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.<p><a href="https://huggingface.co/posts/omarkamali/593639295164067" rel="nofollow">https://huggingface.co/posts/omarkamali/593639295164067</a><p><a href="https://huggingface.co/blog/omarkamali/tokenization" rel="nofollow">https://huggingface.co/blog/omarkamali/tokenization</a></p>
]]></description><pubDate>Sun, 04 Oct 2026 14:55:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49954444</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=49954444</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49954444</guid></item><item><title><![CDATA[New comment by omneity in "Launch HN: Speko (YC S26) – OpenRouter for Voice AI"]]></title><description><![CDATA[
<p>Good old Whisper allows you to enter a prompt with domain specific terms and it will use them for transcription.</p>
]]></description><pubDate>Mon, 17 Aug 2026 18:11:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49335225</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=49335225</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49335225</guid></item><item><title><![CDATA[New comment by omneity in "Compression is prediction"]]></title><description><![CDATA[
<p>I'd approach this distinction differently. Prediction from compression is valid within the distribution of the compressed data. Which brings it much closer to LLMs in this case (can an LLM talk about a topic it has never seen in training? unlikely if it cannot be derived from other training data)</p>
]]></description><pubDate>Tue, 11 Aug 2026 23:24:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49265864</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=49265864</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49265864</guid></item><item><title><![CDATA[New comment by omneity in "Show HN: iPhone app takes simultaneous images from 2 lenses, fuses into 1 photo"]]></title><description><![CDATA[
<p>I’d argue the opposite. The geometry involved in reconciling photos from two separate lenses with different focal lengths is quite challenging. If you pay close attention you will see some differences in the before/after samples in the site (for example the ear in that kitchen selfie).<p>The math is surely more complex than aligning same-lens photography even when you consider the change in angle and perspective between two shots from the same lens.</p>
]]></description><pubDate>Tue, 11 Aug 2026 22:52:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49265590</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=49265590</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49265590</guid></item><item><title><![CDATA[New comment by omneity in "30papers.com – Ilya's 30 essential ML papers, in a beginner friendly format"]]></title><description><![CDATA[
<p>I thought the actual 30 papers have never been disclosed. Do you have a source tying the recommendations back to Ilya, or did you come up with this list?</p>
]]></description><pubDate>Tue, 07 Jul 2026 17:46:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48821117</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48821117</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48821117</guid></item><item><title><![CDATA[New comment by omneity in "How do you keep Web MIDI from crashing a 1983 synthesizer?"]]></title><description><![CDATA[
<p>Glad it helped! A little credit on the post would go a long way :)</p>
]]></description><pubDate>Mon, 29 Jun 2026 21:28:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48725500</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48725500</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48725500</guid></item><item><title><![CDATA[New comment by omneity in "How do you keep Web MIDI from crashing a 1983 synthesizer?"]]></title><description><![CDATA[
<p>The Web MIDI API[0] used by the author has a built-in precise scheduler, that has higher precision and works better than the unreliable setTimeout approach used by OP when coupled with the Performance API[1].<p>Pass a timestamp as the second argument to midiOutput.send(data, timestamp), calculated with performance.now. Something like midiOutput.send(data, performance.now() + offset)<p>0: <a href="https://developer.mozilla.org/en-US/docs/Web/API/MIDIOutput/send" rel="nofollow">https://developer.mozilla.org/en-US/docs/Web/API/MIDIOutput/...</a><p>1: <a href="https://developer.mozilla.org/en-US/docs/Web/API/Performance/now" rel="nofollow">https://developer.mozilla.org/en-US/docs/Web/API/Performance...</a></p>
]]></description><pubDate>Sun, 28 Jun 2026 11:03:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48706289</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48706289</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48706289</guid></item><item><title><![CDATA[New comment by omneity in "GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2"]]></title><description><![CDATA[
<p>I do think it might improve but only marginally.<p>You are however likely to observe better results in smaller models since they're usually more strapped for "cognitive capacity", so two separate calls reduce the load in each request, and hallucination in my experience is a common side effect of overloading an LLM cognitively.</p>
]]></description><pubDate>Sat, 20 Jun 2026 21:42:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48613245</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48613245</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48613245</guid></item><item><title><![CDATA[New comment by omneity in "GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2"]]></title><description><![CDATA[
<p>It’s not as simple. I trained an LLM before on exactly this, to scratch the itch of this question.<p>The task was simple, using the MS-MARCO[0] dataset which contains queries, search results, answers, I made a training set that has:<p>1. Questions paired with real results supporting them (mixed with some irrelevant results), and a correct answer<p>2. Questions paired only with irrelevant results, with the answer “No answer present”<p>The dataset was huge (close to 1M samples), and I trained using different techniques, from SFT (just mimicking the dataset) to DPO (good answer contrasted with a bad answer for the same user query) to GRPO (verifier that checks my annotations whether an answer was present or not)<p>Lo and behold, this didn’t reduce hallucination, rather made it much worse. Now the model started claiming “No answer present” even when it is, or even when the question didn’t need search results in the first place (simple stuff like what is X+Y).<p>Now you could argue that my training was basic compared to what frontier labs could do. Yet I think it hints at a more profound limitation. LLMs are finicky and don’t have a neat understand of things from first principles (list of search results, check relevance of result to user query, if answers are below a certain threshold of relevance then don’t consider them to answer …).<p>tl;dr: not as simple as one might think, perhaps not attainable at all.<p>0: <a href="https://huggingface.co/datasets/microsoft/ms_marco" rel="nofollow">https://huggingface.co/datasets/microsoft/ms_marco</a></p>
]]></description><pubDate>Sat, 20 Jun 2026 11:53:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48608562</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48608562</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48608562</guid></item><item><title><![CDATA[New comment by omneity in "OpenRouter raises $113M Series B"]]></title><description><![CDATA[
<p>The Open in OpenRouter is the same as in OpenSea, as it's the same founder. Make of that what you will.</p>
]]></description><pubDate>Sat, 30 May 2026 19:43:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48339907</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48339907</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48339907</guid></item><item><title><![CDATA[New comment by omneity in "Qwen 3.7 Preview"]]></title><description><![CDATA[
<p>You can increase the context window beyond its max trained context using RoPE scaling[0] which will require more VRAM.<p>But you can increase your context window for the same VRAM by quantizing the KV cache with FP8 (double the context) or TurboQuant (more than double)[1].<p>0: <a href="https://medium.com/@leannetan/extending-context-length-with-hugging-faces-transformers-6b04db05b39a" rel="nofollow">https://medium.com/@leannetan/extending-context-length-with-...</a><p>1: <a href="https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/#example-quantize-llama-attention-kv-cache-to-fp8" rel="nofollow">https://docs.vllm.ai/en/latest/features/quantization/quantiz...</a></p>
]]></description><pubDate>Mon, 18 May 2026 19:30:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48184367</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48184367</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48184367</guid></item><item><title><![CDATA[New comment by omneity in "The US is winning the AI race where it matters most: commercialization"]]></title><description><![CDATA[
<p>Funny, I’ve been cracking[0] at this exact problem with a purpose-built model[1]:<p>0: <a href="https://huggingface.co/posts/omarkamali/593639295164067" rel="nofollow">https://huggingface.co/posts/omarkamali/593639295164067</a><p>1: <a href="https://omneitylabs.com/models/sawtone" rel="nofollow">https://omneitylabs.com/models/sawtone</a></p>
]]></description><pubDate>Wed, 13 May 2026 17:43:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48125045</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=48125045</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48125045</guid></item><item><title><![CDATA[New comment by omneity in "We gave an AI a 3 year retail lease and asked it to make a profit"]]></title><description><![CDATA[
<p>Strong vibes from the novel Manna.<p><a href="https://marshallbrain.com/manna1" rel="nofollow">https://marshallbrain.com/manna1</a></p>
]]></description><pubDate>Thu, 16 Apr 2026 15:45:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=47795039</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47795039</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47795039</guid></item><item><title><![CDATA[New comment by omneity in "Audio Reactive LED Strips Are Diabolically Hard"]]></title><description><![CDATA[
<p>I'm pretty sure it should be possible to distill HS-TasNet into a version approximate and fast enough for the purpose of animating LEDs.<p>At the end it's "just" chunking streamed audio into windows and predicting which LEDs a window should activate. One can build a complex non-realtime pipeline, generate high-quality training data with it, and then train a much smaller model (maybe even an MLP) with it to predict just this task.</p>
]]></description><pubDate>Wed, 08 Apr 2026 18:01:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=47693909</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47693909</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47693909</guid></item><item><title><![CDATA[New comment by omneity in "Meta's Omnilingual MT for 1,600 Languages"]]></title><description><![CDATA[
<p>Excellent, thank you mandeepj! Curious about the language coverage of your agent and if / how you plan to eval your agent, if you're willing to share more.</p>
]]></description><pubDate>Mon, 23 Mar 2026 01:21:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=47484331</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47484331</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47484331</guid></item><item><title><![CDATA[New comment by omneity in "Meta's Omnilingual MT for 1,600 Languages"]]></title><description><![CDATA[
<p>Hey, this is super cool! I’ve been working on a similar problem, focusing on low-resource and underserved languages including the Mayan family, and have published some research and open resources around that [0, 1].<p>On the data side, I’ve found that the biggest bottleneck isn’t collecting text (it’s out there!) but reliable language identification. It’s often difficult or ambiguous to separate languages cleanly in datasets like Common Crawl, Fineweb, or others. I worked on improving this a bit for Fineweb 2 for my native language, that might inspire you [3].<p>Many of the challenges you mention seem to recur across regions and language families, so I’d love to connect and compare notes sometime. Feel free to reach me at omar [at] the labs site below.<p>0: <a href="https://wikilangs.org" rel="nofollow">https://wikilangs.org</a><p>1: <a href="https://omneitylabs.com" rel="nofollow">https://omneitylabs.com</a><p>2: <a href="https://huggingface.co/blog/omarkamali/gherbal-multilingual-fineweb-moroccan-arabic" rel="nofollow">https://huggingface.co/blog/omarkamali/gherbal-multilingual-...</a></p>
]]></description><pubDate>Sat, 21 Mar 2026 22:36:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=47472244</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47472244</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47472244</guid></item><item><title><![CDATA[Tokenization Is Killing Our Multilingual LLM Dream]]></title><description><![CDATA[
<p>Article URL: <a href="https://huggingface.co/blog/omarkamali/tokenization">https://huggingface.co/blog/omarkamali/tokenization</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47388410">https://news.ycombinator.com/item?id=47388410</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 15 Mar 2026 15:34:38 +0000</pubDate><link>https://huggingface.co/blog/omarkamali/tokenization</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47388410</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47388410</guid></item><item><title><![CDATA[I stopped trusting the official Wikipedia dataset, and what I did about it]]></title><description><![CDATA[
<p>Article URL: <a href="https://omarkamali.com/blog/wikipedia-monthly-pipeline">https://omarkamali.com/blog/wikipedia-monthly-pipeline</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47292266">https://news.ycombinator.com/item?id=47292266</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 07 Mar 2026 22:52:41 +0000</pubDate><link>https://omarkamali.com/blog/wikipedia-monthly-pipeline</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47292266</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47292266</guid></item><item><title><![CDATA[A Wordle for the Worldle]]></title><description><![CDATA[
<p>Article URL: <a href="https://omarkamali.com/blog/worldle-for-the-world-wikilangs-launch">https://omarkamali.com/blog/worldle-for-the-world-wikilangs-launch</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47252571">https://news.ycombinator.com/item?id=47252571</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 04 Mar 2026 19:29:17 +0000</pubDate><link>https://omarkamali.com/blog/worldle-for-the-world-wikilangs-launch</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47252571</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47252571</guid></item><item><title><![CDATA[New comment by omneity in "Motorola announces a partnership with GrapheneOS"]]></title><description><![CDATA[
<p>Or your willingness to put up with power banks.</p>
]]></description><pubDate>Mon, 02 Mar 2026 10:48:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=47216256</link><dc:creator>omneity</dc:creator><comments>https://news.ycombinator.com/item?id=47216256</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47216256</guid></item></channel></rss>