<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: parthsareen</title><link>https://news.ycombinator.com/user?id=parthsareen</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Mon, 17 Aug 2026 04:39:44 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=parthsareen" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by parthsareen in "Qwen 3.8 27B"]]></title><description><![CDATA[
<p>Hi! From Ollama here - you can run:
ollama run qwen3.8 (or if on mac qwen3.8:27b-mlx)</p>
]]></description><pubDate>Fri, 14 Aug 2026 21:37:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49304892</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=49304892</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49304892</guid></item><item><title><![CDATA[New comment by parthsareen in "Claude Code: connect to a local model when your quota runs out"]]></title><description><![CDATA[
<p>Also recently added ollama launch claude if you want to connect to cloud models from there :)</p>
]]></description><pubDate>Thu, 05 Feb 2026 01:39:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=46894514</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=46894514</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46894514</guid></item><item><title><![CDATA[New comment by parthsareen in "Ask HN: A good Model to choose in Ollama to run on Claude Code"]]></title><description><![CDATA[
<p>Hey! One of the maintainers of Ollama. 8GB of VRAM is a bit tight for coding agents since their prompts are quite large. You could try playing with qwen3 and at least 16k context length to see how it works.</p>
]]></description><pubDate>Tue, 27 Jan 2026 00:49:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=46773970</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=46773970</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46773970</guid></item><item><title><![CDATA[Building Reliable AI Agents]]></title><description><![CDATA[
<p>Article URL: <a href="https://parthsareen.com/writings/building-reliable-ai-agents/">https://parthsareen.com/writings/building-reliable-ai-agents/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46558809">https://news.ycombinator.com/item?id=46558809</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Fri, 09 Jan 2026 20:22:59 +0000</pubDate><link>https://parthsareen.com/writings/building-reliable-ai-agents/</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=46558809</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46558809</guid></item><item><title><![CDATA[New comment by parthsareen in "A guide to local coding models"]]></title><description><![CDATA[
<p>How much ram are you running with? Qwen3 and gpt-oss:20b punch a good bit above their weight. Personally use it for small agents.</p>
]]></description><pubDate>Mon, 22 Dec 2025 05:14:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=46351544</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=46351544</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46351544</guid></item><item><title><![CDATA[New comment by parthsareen in "A guide to local coding models"]]></title><description><![CDATA[
<p>You're welcome to go through the source: <a href="https://github.com/ollama/ollama/" rel="nofollow">https://github.com/ollama/ollama/</a></p>
]]></description><pubDate>Mon, 22 Dec 2025 05:10:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=46351516</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=46351516</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46351516</guid></item><item><title><![CDATA[New comment by parthsareen in "A guide to local coding models"]]></title><description><![CDATA[
<p>Desktop app is open-source now.</p>
]]></description><pubDate>Mon, 22 Dec 2025 05:09:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=46351508</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=46351508</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46351508</guid></item><item><title><![CDATA[How to Train an LLM: Part 1]]></title><description><![CDATA[
<p>Article URL: <a href="https://omkaark.com/posts/llm-1b-1.html">https://omkaark.com/posts/llm-1b-1.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45891801">https://news.ycombinator.com/item?id=45891801</a></p>
<p>Points: 20</p>
<p># Comments: 3</p>
]]></description><pubDate>Tue, 11 Nov 2025 19:38:59 +0000</pubDate><link>https://omkaark.com/posts/llm-1b-1.html</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45891801</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45891801</guid></item><item><title><![CDATA[New comment by parthsareen in "Ollama Web Search"]]></title><description><![CDATA[
<p>Since we shipped web search with gpt-oss in the Ollama app I've personally been using that a lot more especially for research heavy tasks that I can shoot off. Plus with a 5090 or the new macs it's super fast.</p>
]]></description><pubDate>Thu, 25 Sep 2025 20:58:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=45378978</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45378978</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45378978</guid></item><item><title><![CDATA[New comment by parthsareen in "Ollama Web Search"]]></title><description><![CDATA[
<p>Hi - author of the post. Yes it does! The "build a search agent" example can be used with a local model. I'd recommend trying qwen3 or gpt-oss</p>
]]></description><pubDate>Thu, 25 Sep 2025 20:44:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=45378793</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45378793</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45378793</guid></item><item><title><![CDATA[New comment by parthsareen in "Ollama Web Search"]]></title><description><![CDATA[
<p>Hey! Author of the blogpost and I also work on Ollama's tool calling. There has been a big push on tool calling over the last year to improve the parsing. What's the issues you're running into with local tool use? What models are you using?</p>
]]></description><pubDate>Thu, 25 Sep 2025 20:43:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=45378778</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45378778</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45378778</guid></item><item><title><![CDATA[New comment by parthsareen in "Sampling and structured outputs in LLMs"]]></title><description><![CDATA[
<p>That's a great idea. Going to try this next :)</p>
]]></description><pubDate>Thu, 25 Sep 2025 06:23:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=45369806</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45369806</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45369806</guid></item><item><title><![CDATA[New comment by parthsareen in "Sampling and structured outputs in LLMs"]]></title><description><![CDATA[
<p>Hey! I'm the author of the post. We haven't optimized sampling yet so it's running linearly on the CPU. A lot of SOTA work either does this while the model is running the forward pass or does the masking on the GPU.<p>The greedy accept is so that the mask doesn't need to be computed. Planning to make this more efficient from either ends.</p>
]]></description><pubDate>Thu, 25 Sep 2025 06:22:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=45369805</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45369805</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45369805</guid></item><item><title><![CDATA[New comment by parthsareen in "Sampling and structured outputs in LLMs"]]></title><description><![CDATA[
<p>Thank you! Maybe not "perfect" but near-perfect is something we can expect. Models like the Osmosis structure which just structure data inspired some of that thinking (<a href="https://ollama.com/Osmosis/Osmosis-Structure-0.6B">https://ollama.com/Osmosis/Osmosis-Structure-0.6B</a>). Historically, JSON generation has been a latent capability of a model rather than a trained one, but that seems to be changing. gpt-oss was particularly trained for this type of behavior and so the token probabilities are heavily skewed to conform to JSON. Will be interesting to see the next batch of models!</p>
]]></description><pubDate>Tue, 23 Sep 2025 19:01:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=45351322</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45351322</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45351322</guid></item><item><title><![CDATA[New comment by parthsareen in "Sampling and structured outputs in LLMs"]]></title><description><![CDATA[
<p>Thanks for posting! Didn't expect this to get picked up – it was a bit of a draft haha. Happy to answer questions around structured outputs :)</p>
]]></description><pubDate>Tue, 23 Sep 2025 18:01:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=45350622</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=45350622</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45350622</guid></item><item><title><![CDATA[New comment by parthsareen in "Structured Outputs with Ollama"]]></title><description><![CDATA[
<p>Yes! I have checked guidance out, as well as a few others. Planning to refactor sampling in the near future which would include improving using grammars for sampling as well. Thanks for sharing!</p>
]]></description><pubDate>Mon, 09 Dec 2024 22:24:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=42371223</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=42371223</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42371223</guid></item><item><title><![CDATA[New comment by parthsareen in "Structured Outputs with Ollama"]]></title><description><![CDATA[
<p>The constraints will always be met. It’s the data inside that might be inaccurate. YMMV with smaller models in that sense.</p>
]]></description><pubDate>Sat, 07 Dec 2024 17:57:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=42351440</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=42351440</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42351440</guid></item><item><title><![CDATA[New comment by parthsareen in "Structured Outputs with Ollama"]]></title><description><![CDATA[
<p>Hey! Author of the blog here. The current implementation uses llama.cpp GBNF which has allowed for a quick implementation. The biggest value-add at this time was getting the feature out.<p>With the newer research - outlines/xgrammar coming out, I hope to be able to update the sampling to support more formats, increase accuracy, and improve performance.</p>
]]></description><pubDate>Sat, 07 Dec 2024 17:31:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=42351243</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=42351243</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42351243</guid></item><item><title><![CDATA[New comment by parthsareen in "Structured Outputs with Ollama"]]></title><description><![CDATA[
<p>Hey! Author of the post and one of the maintainers here. I agree - we (maintainers) got to this late and in general want to encourage more contributions.<p>Hoping to be more on top of community PRs and get them merged in the coming year.</p>
]]></description><pubDate>Sat, 07 Dec 2024 17:26:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=42351214</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=42351214</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42351214</guid></item><item><title><![CDATA[New comment by parthsareen in "Structured Outputs with Ollama"]]></title><description><![CDATA[
<p>This looks really useful. Thank you!</p>
]]></description><pubDate>Sat, 07 Dec 2024 04:02:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=42347160</link><dc:creator>parthsareen</dc:creator><comments>https://news.ycombinator.com/item?id=42347160</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42347160</guid></item></channel></rss>