<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: jakevoytko</title><link>https://news.ycombinator.com/user?id=jakevoytko</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 27 Sep 2026 16:08:50 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=jakevoytko" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by jakevoytko in "Ask HN: What are you working on? (September 2026)"]]></title><description><![CDATA[
<p>The only thing I would open source is the spec file from the Q/A so people can implement it if they'd like.<p>I find it extremely unlikely that I'd be able to sell it since its two main competitors are (1) literally free to use and has been polished for 16 years, and (2) a best-in-class native application that has been the lingua franca of business writing for over 40 years</p>
]]></description><pubDate>Sun, 20 Sep 2026 14:01:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49775983</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49775983</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49775983</guid></item><item><title><![CDATA[New comment by jakevoytko in "Astra for Law"]]></title><description><![CDATA[
<p>> all relevant facts will be cited and checked easily by humans<p>I've talked to a lawyer about how they handle this. They do indeed double-check everything, since it'd be embarrassing (or worse) to send hallucinated statements to opposing council or to the court. They still find the assembly a huge time saver<p>But based on stories in the news on the subject, not everyone has this same level of diligence</p>
]]></description><pubDate>Thu, 17 Sep 2026 21:54:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49747135</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49747135</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49747135</guid></item><item><title><![CDATA[New comment by jakevoytko in "Can we stop with the uptime percentages?"]]></title><description><![CDATA[
<p>These numbers are useful proxies for how likely you are to have your work disrupted outside of your own control.<p>If you do something 100 times a day against a four-nines service, you can reasonably expect that everything will succeed.<p>If you do something 10,000 times a day against a two-nines service, you can expect to hit a substantial number of errors during that day, or even have long periods where your work cannot happen at all.<p>People aren't frustrated with Github because Github has 98% uptime or whatever the specific number is. They're frustrated because it regularly interferes with their ability to work. The 98% number is just a concise way to say it.</p>
]]></description><pubDate>Wed, 16 Sep 2026 16:17:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49729295</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49729295</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49729295</guid></item><item><title><![CDATA[New comment by jakevoytko in "Salesforce Global Outage"]]></title><description><![CDATA[
<p>In my experience it’s a safe way to do something useful while everyone is getting their bearings. It immediately partitions the situation space between being persisted or systemic vs local or caused by long-running processes. Plus everyone’s going to ask if you’ve tried that already, so you might as well get it out of the way if it makes any amount of sense</p>
]]></description><pubDate>Wed, 16 Sep 2026 14:10:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49727314</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49727314</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49727314</guid></item><item><title><![CDATA[New comment by jakevoytko in "XCancel service is suspended until further notice"]]></title><description><![CDATA[
<p>I know there are like 5 trillion Instagram viewers. But I've never actually needed to use any of them because there is absolutely nothing so important on Instagram that I <i>must</i> see it. The entire site is optional.</p>
]]></description><pubDate>Tue, 15 Sep 2026 03:40:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49707426</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49707426</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49707426</guid></item><item><title><![CDATA[New comment by jakevoytko in "XCancel service is suspended until further notice"]]></title><description><![CDATA[
<p>Actual details.</p>
]]></description><pubDate>Tue, 15 Sep 2026 03:38:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49707419</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49707419</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49707419</guid></item><item><title><![CDATA[New comment by jakevoytko in "XCancel service is suspended until further notice"]]></title><description><![CDATA[
<p>I hate when companies do this. Instagram is a black box for me. Whenever someone sends me a link the web app is completely nonfunctional; clicking links just reload the page, it constantly tries opening the app, etc<p>The last time I signed up for an Instagram account (2016) it got immediately flagged as a bot account. I think I was supposed to appeal it to have it reinstanted, but why bother if this is how badly they treat users?<p>Can you imagine going to work as Instagram's web developer, spending your day finding new ways that the app accidentally works and disabling them?</p>
]]></description><pubDate>Mon, 14 Sep 2026 18:04:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49701202</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49701202</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49701202</guid></item><item><title><![CDATA[New comment by jakevoytko in "Ask HN: What are you working on? (September 2026)"]]></title><description><![CDATA[
<p>I built my own personal Google Docs clone to use as the editor for my blog<p>I used to work on the Google Docs team, so I basically did an 8-hour Q/A with Sol spanning like 400 questions. Did my best to reach back 10 years into my memories to see how things were implemented, and discussing architecture and tradeoffs. Then went through 5 ChatGPT resets implementing it on Astra to get a feel for the model.<p>I have a long way to go. But I'm writing my first post on it this week, and then I'll do a writeup of the editor project.</p>
]]></description><pubDate>Mon, 14 Sep 2026 14:12:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49697215</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49697215</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49697215</guid></item><item><title><![CDATA[New comment by jakevoytko in "Ask HN: What default model do you use and why?"]]></title><description><![CDATA[
<p>My workhorse is Opus orchestrating Sonnet or Sol orchestrating Luna, depending on which ecosystem I'm using this week. I'm a backend engineer, so for tasks that either interact with the machine learning or client codebases, I have Fable (Medium) review across the product spec, tech spec, and all the codebases after the implementation is done (haven't tried this workflow with Astra yet, don't know if it's a good reviewer).<p>As a side project, I'm doing an experimental task now (having Astra implement my own personal Google Docs clone, OT and everything, for my blog), and it feels more capable than Fable but needs to be watched closer than Fable. It is obsessed with verification and evidence to the point that I've needed to make some absolute rules to stop it from repeatedly e.g. making unit tests with 10,000 test cases with 5+ hours of runtime<p>For well-defined coding tasks, I'd default to either Luna or Sonnet, or hell, just writing it by hand.</p>
]]></description><pubDate>Sat, 12 Sep 2026 17:46:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49675002</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49675002</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49675002</guid></item><item><title><![CDATA[New comment by jakevoytko in "On the Navier–Stokes Millennium Prize Problem"]]></title><description><![CDATA[
<p>For full context, here's the HN thread from the other side of the "Concurrent Work" section: <a href="https://news.ycombinator.com/item?id=49605915">https://news.ycombinator.com/item?id=49605915</a><p>Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees</p>
]]></description><pubDate>Tue, 08 Sep 2026 17:22:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49613393</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49613393</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49613393</guid></item><item><title><![CDATA[New comment by jakevoytko in "How well do agents use test/verification techniques?"]]></title><description><![CDATA[
<p>As someone who also spent a bunch of time benchmarking various techniques, this is as good as you can do without publishing a formal versioned benchmark suite that you want to maintain and run forever at immense cost to yourself. If you actually go to benchmark your own flow as you mentioned, you will quickly run into like a dozen problems that discourage you from publishing.<p>- Are you sure that temperature and other nondeterminism isn't affecting your output?<p>- Are you sure you're not being routed through an A/B test at this moment?<p>- Are you sure there's not a bug affecting the model at this moment?<p>- Are you sure that you picked the right model and effort level?<p>- Are you sure that your result generalizes across providers?<p>- Are you sure that you set up the correct level of sandboxing and the agent can't e.g. look at a sister directory or git history in the current directory for answers?<p>- Are you sure that the agent isn't leaking answers in memory or its conversation history?<p>- Are you sure that tool calls aren't somehow affecting results?<p>- Are you sure that your results are robust, i.e. you see the same results with mild tweaks to the prompt?<p>- Are you comfortable keeping your blog post live when your results are invalidated next week with the next model launch?<p>And that's just a quick list off the top of my head.<p>I personally decided that it wasn't worth it, I'm glad that Dan decided to publish his. Frankly I think we could use a lot more of these "I ran these 2 techniques side by side and here's what I saw" anecdata, because most people who promote prompt techniques can't produce a single prompt they ran twice because they never actually tested it per se.<p>Edited to add: formatting + the word "promote"</p>
]]></description><pubDate>Tue, 08 Sep 2026 13:49:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49610342</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49610342</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49610342</guid></item><item><title><![CDATA[New comment by jakevoytko in "Finder is so frustrating and has been since day one"]]></title><description><![CDATA[
<p>This was my instant reaction when I started using a Mac. But 15 years later, even knowing more tricks for interacting with it, it still baffles me. Once you've used Windows Explorer and a few Linux distro navigators, there's no denying that Finder is unquestionably worst in class.</p>
]]></description><pubDate>Sun, 06 Sep 2026 19:53:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49590226</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49590226</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49590226</guid></item><item><title><![CDATA[New comment by jakevoytko in "GLP-1s are being linked to fewer serious infections, including TB"]]></title><description><![CDATA[
<p>Yeah this is what I pay for Zepbound through LillyDirect. I also eat less food and drink less alcohol than I used to. So the net loss is probably smaller, maybe $200/mo</p>
]]></description><pubDate>Fri, 04 Sep 2026 00:09:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558820</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49558820</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558820</guid></item><item><title><![CDATA[New comment by jakevoytko in "FBI Probes Service Selling 153M+ Drivers Licenses"]]></title><description><![CDATA[
<p>As always, friendly reminder to lock your credit and enable your mobile carrier's protections against SIM swapping</p>
]]></description><pubDate>Wed, 02 Sep 2026 02:52:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49531173</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49531173</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49531173</guid></item><item><title><![CDATA[New comment by jakevoytko in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>My experience with output styles for long-running sessions is that Claude starts to forget the terse output style by the middle of the context window. Obviously I don't know if 5.1 suffers the same fate but I ran into this issue with both Opus and Fable 5</p>
]]></description><pubDate>Tue, 01 Sep 2026 19:46:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49527090</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49527090</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49527090</guid></item><item><title><![CDATA[New comment by jakevoytko in "Bug Blindness"]]></title><description><![CDATA[
<p>This reminds me of my first job, where we did a lot of 3d modeling and mobile robotics research. When we were trying to reproduce motion bugs, a coworker of mine would track down our manager and put him at the controls. More times than not, the bug would surface and it'd trip our logging and we tracked it down.<p>I asked him why he does this. His explanation was really built on this operant conditioning idea: "we use this stuff for 8 hours a day and we train ourselves to avoid all of its little pitfalls. So I get the most available person who hasn't used it all day, which is our manager. He uses it differently than we do because he doesn't avoid all of its little problems. But if we have a really tricky problem, I get our manager's manager. I don't know if you've ever seen him try to use an xbox controller, but he has the spatial reasoning abilities of a goldfish. He's never failed to reproduce a really hard bug. If I ever needed an Einstein-level bug reproduction, I'd track down the head of the department and put him in front of it, but it's never come to that.</p>
]]></description><pubDate>Sun, 30 Aug 2026 02:32:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49495158</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49495158</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49495158</guid></item><item><title><![CDATA[New comment by jakevoytko in "Slack Code"]]></title><description><![CDATA[
<p>They see the entire universe building their own version of Claude Tag and custom internal bots connected to your internal ecosystem, and realizing that they can fight for some enterprise revenue within their own product that everyone else is currently extracting</p>
]]></description><pubDate>Thu, 20 Aug 2026 15:12:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49375798</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49375798</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49375798</guid></item><item><title><![CDATA[New comment by jakevoytko in "A 25-year-old video patent just expired, ending a legal headache for Linux"]]></title><description><![CDATA[
<p>I worked at Etsy when it expired and the execs got this question a lot. The TL;DR is that the expected outcome is that it massively increases the support burden (wait I didn't mean to click that; wait I actually need to send it to another address; wait I didn't realize shipping was a hundred dollars) without really enabling more sales. So it was a neat idea and worth trying when ecommerce was new, but now we know enough about ecommerce to know that the user has to confirm details of their purchase.</p>
]]></description><pubDate>Wed, 19 Aug 2026 13:50:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49361593</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49361593</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49361593</guid></item><item><title><![CDATA[New comment by jakevoytko in ""Solving a largely imaginary user goal""]]></title><description><![CDATA[
<p>Their argument is that the user almost always experiences the moment of interaction with the controls as a single decision with 2 branches, (a) the OS behavior matches my expectations and I won't change it, and (b) the OS behavior violates my expectations and I want it to be the other one.<p>I'd also rather just select from the full tristate diagram, but their framing also makes sense to me</p>
]]></description><pubDate>Fri, 14 Aug 2026 16:44:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49301247</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49301247</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49301247</guid></item><item><title><![CDATA[New comment by jakevoytko in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>It's clear that we're not the audience; it writes to be read by its training evaluator, not a professional software engineer. Professional software engineers can't read this word soup and are desperately trying to find ways to fix it.<p>It feels like it found a register that games the evaluator, where it can ramble forever and rarely be marked wrong while slowly racking up points as it talks more.</p>
]]></description><pubDate>Fri, 14 Aug 2026 15:49:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49300414</link><dc:creator>jakevoytko</dc:creator><comments>https://news.ycombinator.com/item?id=49300414</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49300414</guid></item></channel></rss>