<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: nonatofabio</title><link>https://news.ycombinator.com/user?id=nonatofabio</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 10 Oct 2026 05:13:53 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=nonatofabio" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by nonatofabio in "Microsoft-Decision-1, our model for fast decision-making"]]></title><description><![CDATA[
<p>Yep, Fabio here, and I agree!</p>
]]></description><pubDate>Fri, 09 Oct 2026 23:40:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=50027930</link><dc:creator>nonatofabio</dc:creator><comments>https://news.ycombinator.com/item?id=50027930</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50027930</guid></item><item><title><![CDATA[Show HN: Luna Agent – Custom AI agent in ~2300 lines of Python, no frameworks]]></title><description><![CDATA[
<p>Hey HN. I evaluated three agent frameworks for a homelab project, one had 400K lines of code (and 42K exposed instances on Shodan), one was 9 days old, and the third was so thin I'd rebuild most of it anyway. None worked for me.<p>So I built my own in ~2300 lines of Python. No frameworks, 8 runtime dependencies, 106 tests.<p>What it does:
- Persistent memory via SQLite (FTS5 keyword search + sqlite-vec embeddings + recency decay, fused with Reciprocal Rank Fusion)
- MCP tool integration — add capabilities by editing a JSON file
- Native tools with safety guardrails (bash blocklist, timeouts, output caps)
- Discord interface with session isolation
- Structured JSON logging for every operation
- Conversation compression for effectively infinite context<p>Runs locally on 2x RTX 3090 with Qwen3-Coder-Next via llama-server. No cloud APIs.<p>The design philosophy was: don't build what you don't need, but don't block the insertion points. For example, my AI firewall isn't built yet, but all LLM traffic goes through a single configurable URL, swapping in a filtering proxy is a config change I'll do later.<p>DESIGN.md documents the reasoning behind every architectural decision. Tests mock the LLM client so you can run them on a laptop.<p>GitHub: <a href="https://github.com/nonatofabio/luna-agent" rel="nofollow">https://github.com/nonatofabio/luna-agent</a>
Blog post with full technical deep-dive: <a href="https://nonatofabio.github.io/blog/post.html?slug=luna_agent" rel="nofollow">https://nonatofabio.github.io/blog/post.html?slug=luna_agent</a><p>Happy to answer questions about any of the design tradeoffs.</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47291343">https://news.ycombinator.com/item?id=47291343</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 07 Mar 2026 20:49:53 +0000</pubDate><link>https://nonatofabio.github.io/blog/post.html?slug=luna_agent</link><dc:creator>nonatofabio</dc:creator><comments>https://news.ycombinator.com/item?id=47291343</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47291343</guid></item><item><title><![CDATA[Show HN: Local_faiss_MCP – A tiny MCP server for local RAG (FAISS and MiniLM)]]></title><description><![CDATA[
<p>I built this because I got frustrated with the current state of "local" RAG. It felt like I had to spin up a Docker container, configure a vector DB, and manage an ingestion pipeline just to let Claude ask questions about a few PDFs in a folder.<p>We seem to have turned "grep with semantics" into a microservices architecture problem.<p>What this is: local_faiss_mcp is a minimal implementation of the Model Context Protocol (MCP) that wraps FAISS and sentence-transformers. It runs entirely locally (no API keys, no external services) and connects to Claude Desktop via stdio.<p>How it works:<p>You run server.py (Claude runs this automatically via config).<p>It uses all-MiniLM-L6-v2 (on CPU) to embed text.<p>It stores the vectors in a flat FAISS index on disk alongside a JSON metadata file.<p>It exposes two tools to the LLM: ingest_document and query_rag_store.<p>The stack:<p>Python<p>mcp (Python SDK)<p>faiss-cpu<p>sentence-transformers<p>It’s intended for personal workflows (notes, logs, specs) where you want persistent memory for an agent without the infrastructure overhead.<p>Repo: <a href="https://github.com/nonatofabio/local_faiss_mcp" rel="nofollow">https://github.com/nonatofabio/local_faiss_mcp</a><p>I’d love feedback on the implementation—specifically if anyone has ideas on better handling the chunking logic without bloating the dependencies, or if you run into performance issues with larger indices (10k+ vectors).</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46136319">https://news.ycombinator.com/item?id=46136319</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 03 Dec 2025 16:22:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=46136319</link><dc:creator>nonatofabio</dc:creator><comments>https://news.ycombinator.com/item?id=46136319</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46136319</guid></item></channel></rss>