<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: fzysingularity</title><link>https://news.ycombinator.com/user?id=fzysingularity</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 17 Sep 2026 14:11:51 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=fzysingularity" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by fzysingularity in "Xiaomi Mimo 2.6 live post-training dashboard"]]></title><description><![CDATA[
<p>Very cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.</p>
]]></description><pubDate>Wed, 16 Sep 2026 22:15:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49733777</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49733777</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49733777</guid></item><item><title><![CDATA[New comment by fzysingularity in "So you want to use OpenRouter?"]]></title><description><![CDATA[
<p>text is mostly beaten to death, so you can expect good defaults to work for vllm. VLMs specifically are quite sensitive to quantization, especially if you want it to do fine-grained localization (time or spatial), and the vllm default params can be way off for your use-case.<p>For example, most vision models don't need 256K context length for modes like Qwen3.8-27B when all you care about is single-image captioning, so you can technically save on KV cache. Video reasoning does require that context length, so it's a different set of deployment parameters that need to be enabled.<p>All of this to say that the providers that offer these models, are simply using vLLM / SGLang, and mostly cater to the text inference use-case (coding, etc). Vision always seems to be a bit of an afterthought.</p>
]]></description><pubDate>Fri, 11 Sep 2026 21:12:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49665456</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49665456</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49665456</guid></item><item><title><![CDATA[New comment by fzysingularity in "So you want to use OpenRouter?"]]></title><description><![CDATA[
<p>The author’s comments on vision providers is especially interesting. We saw that most providers don’t provide native video url support, have high-variability in vision performance (likely due to the fact that they’re serving different quantization levels behind the same model id).<p>If you’re building vision-native apps, there are so many footguns in vLLM/SGLang serving configurations, let alone the routing/orchestration in providers like OR, that leave the user more confused about the model’s capabilities.</p>
]]></description><pubDate>Fri, 11 Sep 2026 15:07:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49659752</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49659752</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49659752</guid></item><item><title><![CDATA[New comment by fzysingularity in "How to run Qwen3.8-27B on a single 16GB card"]]></title><description><![CDATA[
<p>Have you tried any vision tasks with this model? We've been serving these on our gateway [1], and the quants are quite terrible for vision. Curious to hear your experience.<p>[1] <a href="https://www.vlm.run/gateway" rel="nofollow">https://www.vlm.run/gateway</a></p>
]]></description><pubDate>Fri, 04 Sep 2026 16:28:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49566820</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49566820</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49566820</guid></item><item><title><![CDATA[New comment by fzysingularity in "METR Report on OpenAI / Hugging Face Hacking Incident"]]></title><description><![CDATA[
<p>I'm surprised this post isn't getting as much attention as it should. Crazy times!</p>
]]></description><pubDate>Thu, 03 Sep 2026 02:31:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49545274</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49545274</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49545274</guid></item><item><title><![CDATA[New comment by fzysingularity in "Ask HN: Who is hiring? (September 2026)"]]></title><description><![CDATA[
<p>VLM Run (<a href="https://vlm.run" rel="nofollow">https://vlm.run</a>) | 1x Founding Infrastructure Engineer<p>We’re building the inference platform for visual intelligence. We’re a deeply technical team of veteran AI / computer-vision engineers (20+ years combined, MIT/CMU PhDs) who’ve shipped production ML infrastructure across autonomous driving and LLMs.<p>We just launched the VLM Run Gateway (<a href="https://vlm.run/gateway" rel="nofollow">https://vlm.run/gateway</a>), a unified API that serves open-weight VLMs, embodied VLAs, ViTs, served across a fleet of GPUs and clouds. That's exactly the infrastructure problem this role will own.<p>If you’re interested, email us at hiring@vlm.run with your GitHub profile and link to projects you’ve recently built - especially with docker, k8s, GPUs.</p>
]]></description><pubDate>Wed, 02 Sep 2026 19:33:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49541233</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49541233</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49541233</guid></item><item><title><![CDATA[Run GLM-OCR, DeepSeek-OCR-2, Dots.mocr with an OpenAI Compatible API]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.vlm.run/product/gateway">https://www.vlm.run/product/gateway</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49365625">https://news.ycombinator.com/item?id=49365625</a></p>
<p>Points: 6</p>
<p># Comments: 1</p>
]]></description><pubDate>Wed, 19 Aug 2026 18:49:17 +0000</pubDate><link>https://www.vlm.run/product/gateway</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=49365625</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49365625</guid></item><item><title><![CDATA[New comment by fzysingularity in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>I think we all ought to look at the ZDR fine-print here.<p>I get that in principle that there's no retention, but these are powerful models that can comprehend, paraphrase and summarize your logs for the sake of "product" improvement. Who knows what's collected here.</p>
]]></description><pubDate>Thu, 16 Jul 2026 23:46:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=48941764</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48941764</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48941764</guid></item><item><title><![CDATA[New comment by fzysingularity in "Mistral's Robostral Navigate: a state of the art robotics navigation model"]]></title><description><![CDATA[
<p>The ICP question was more around the model itself. Are they looking to license it to robotics companies? Do they imagine that devs at robotics companies would be willing to deploy these models as a black box?</p>
]]></description><pubDate>Wed, 08 Jul 2026 16:22:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48833898</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48833898</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48833898</guid></item><item><title><![CDATA[New comment by fzysingularity in "Mistral's Robostral Navigate: a state of the art robotics navigation model"]]></title><description><![CDATA[
<p>It’s unclear to me what their desired outcome for a blog post like this. If you’ve ever worked in a robotics setting, 80% implies that 20% of your autonomous actions are incorrect. Imagine if this were the case for autonomous driving where your car misbehaves 1 in every 5 actions it takes.<p>Posts like this just reminds me of the end to end demos AV companies built in the early days using a single camera - only to realize that it’s harder than it looks years later into development.</p>
]]></description><pubDate>Wed, 08 Jul 2026 16:18:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48833852</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48833852</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48833852</guid></item><item><title><![CDATA[New comment by fzysingularity in "Mistral's Robostral Navigate: a state of the art robotics navigation model"]]></title><description><![CDATA[
<p>Frontier labs are realizing that software/models themselves don’t have real moats and move to embodied ai.<p>SOTA 80%  means a practically useless robot. What are they really imagining their ICP to be here?</p>
]]></description><pubDate>Wed, 08 Jul 2026 14:58:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=48832841</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48832841</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48832841</guid></item><item><title><![CDATA[New comment by fzysingularity in "Claude-real-video － any LLM can watch a video"]]></title><description><![CDATA[
<p>It's live now, <a href="https://github.com/vlm-run/mm" rel="nofollow">https://github.com/vlm-run/mm</a>.</p>
]]></description><pubDate>Mon, 06 Jul 2026 20:24:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48810019</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48810019</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48810019</guid></item><item><title><![CDATA[New comment by fzysingularity in "Claude-real-video － any LLM can watch a video"]]></title><description><![CDATA[
<p>We were planning to open-source this soon, but jumped the gun and posted about the video encoders here since it seemed relevant.<p>In either case, here you go, it's public now: <a href="https://github.com/vlm-run/mm" rel="nofollow">https://github.com/vlm-run/mm</a>.</p>
]]></description><pubDate>Mon, 06 Jul 2026 16:18:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=48806763</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48806763</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48806763</guid></item><item><title><![CDATA[New comment by fzysingularity in "Claude-real-video － any LLM can watch a video"]]></title><description><![CDATA[
<p>Exactly! We experimented with a whole bunch of video encoding techniques for LLMs here: <a href="https://vlm-run.github.io/mm/encoders/#video" rel="nofollow">https://vlm-run.github.io/mm/encoders/#video</a></p>
]]></description><pubDate>Thu, 02 Jul 2026 23:46:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48768866</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48768866</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48768866</guid></item><item><title><![CDATA[New comment by fzysingularity in "Claude-real-video － any LLM can watch a video"]]></title><description><![CDATA[
<p>Pretty terribly expensive way to watch a video with Claude.<p>Use Gemini or some local VLM to do this way more efficiently. We spent quite a bit of time on video understanding, and Claude will just burn tokens.<p>Check out this library: <a href="https://vlm-run.github.io/mm/" rel="nofollow">https://vlm-run.github.io/mm/</a><p>You can swap models and try out different encoding methods for videos (<a href="https://vlm-run.github.io/mm/encoders/#video" rel="nofollow">https://vlm-run.github.io/mm/encoders/#video</a>)</p>
]]></description><pubDate>Thu, 02 Jul 2026 23:43:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48768847</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48768847</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48768847</guid></item><item><title><![CDATA[New comment by fzysingularity in "How to Passive-Aggressively Shame People Who Use LLMs Selfishly"]]></title><description><![CDATA[
<p>This is neat. I'd love to figure out a sequence of emojis that triggers the LLM in ways that puzzles a human.</p>
]]></description><pubDate>Wed, 24 Jun 2026 01:28:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48654009</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48654009</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48654009</guid></item><item><title><![CDATA[Omnigent: Meta-Harness for Coding Agents (Claude Code, Codex, Cursor, Pi)]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/omnigent-ai/omnigent">https://github.com/omnigent-ai/omnigent</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48575159">https://news.ycombinator.com/item?id=48575159</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 17 Jun 2026 19:03:09 +0000</pubDate><link>https://github.com/omnigent-ai/omnigent</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48575159</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48575159</guid></item><item><title><![CDATA[New comment by fzysingularity in "OpenCV 5 Is Here: The Biggest Leap in Years for Computer Vision"]]></title><description><![CDATA[
<p>That’s a pretty large binary for simply loading images.<p>In all honesty, opencv has stood the test of time and I’m certain newer LLMs will likely not attempt to rewrite it from scratch.<p>P.S. I’ve been a user since the IplImage days, circa 2007, and I’d still consider using it over most CV libraries today.</p>
]]></description><pubDate>Wed, 10 Jun 2026 05:54:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48472001</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48472001</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48472001</guid></item><item><title><![CDATA[New comment by fzysingularity in "Claude Fable 5"]]></title><description><![CDATA[
<p>I can’t help but think that there are so many astroturfed comments in here.<p>Seems like a concerted and distributed effort from the entire Anthropic team every time to get this on top of HN.</p>
]]></description><pubDate>Wed, 10 Jun 2026 04:23:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48471397</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48471397</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48471397</guid></item><item><title><![CDATA[New comment by fzysingularity in "Running Python code in a sandbox with MicroPython and WASM"]]></title><description><![CDATA[
<p>Kind of crazy how many bespoke python sandbox implementations have popped up in the past few months.<p>I’d love to see if we can get GPU access within these runtimes, that’d be awesome.</p>
]]></description><pubDate>Sun, 07 Jun 2026 03:34:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48431522</link><dc:creator>fzysingularity</dc:creator><comments>https://news.ycombinator.com/item?id=48431522</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48431522</guid></item></channel></rss>