<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mjb</title><link>https://news.ycombinator.com/user?id=mjb</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 11 Oct 2026 09:12:39 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mjb" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mjb in "Microsoft-Decision-1, our model for fast decision-making"]]></title><description><![CDATA[
<p>> This is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider<p>As of this afternoon, we also have Gemma4-based variants of strands-decider at 2B, 4B, 12B, and 26B: <a href="https://huggingface.co/StrandsAgents" rel="nofollow">https://huggingface.co/StrandsAgents</a><p>I don't know if Fabio (who's leading work on this model line) would agree, but if I was going to start over I'd probably pick Gemma4 as the base rather than Qwen3.5. But it's not surprising to see a lot of Qwen competition, and I totally agree it's nice to see the work open.</p>
]]></description><pubDate>Fri, 09 Oct 2026 23:17:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=50027784</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=50027784</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50027784</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>Oh, huh, that's an error in the image I didn't notice! Those comparisons are to a model called 'decider-2B', which isn't ours.<p>These benchmarks cover strands-decider v19. v21 is better calibrated, and we've got some new ones coming this week that move us further towards the bottom right.</p>
]]></description><pubDate>Wed, 07 Oct 2026 16:23:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49994988</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49994988</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49994988</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>No particularly principled reason, no. It's one of the (many) design variants we haven't had time to experiment with yet.</p>
]]></description><pubDate>Wed, 07 Oct 2026 15:41:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=49994409</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49994409</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49994409</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>One of the authors here.<p>I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)<p>As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).</p>
]]></description><pubDate>Wed, 07 Oct 2026 12:37:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991910</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991910</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991910</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>It's very doable, but we haven't done it yet (although some great community folks did an ONNX version of v21).<p>I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).</p>
]]></description><pubDate>Wed, 07 Oct 2026 12:34:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991885</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991885</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991885</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>One of the authors here. Thanks!<p>On humans-for-humans, blog posts with my name on aren't written by AI (<a href="https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html" rel="nofollow">https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html</a>) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.<p>There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.<p>We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.</p>
]]></description><pubDate>Wed, 07 Oct 2026 12:29:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991831</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991831</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991831</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>I wrote a slightly deeper blog post here: <a href="https://brooker.co.za/blog/2026/09/28/engineering-system-one.html" rel="nofollow">https://brooker.co.za/blog/2026/09/28/engineering-system-one...</a><p>On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.<p>The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.</p>
]]></description><pubDate>Wed, 07 Oct 2026 12:22:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991766</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991766</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991766</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>v21 supports vision: <a href="https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v21" rel="nofollow">https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...</a></p>
]]></description><pubDate>Wed, 07 Oct 2026 12:14:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991684</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991684</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991684</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>One of the authors here.<p>v21 supports vision: <a href="https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v21" rel="nofollow">https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...</a></p>
]]></description><pubDate>Wed, 07 Oct 2026 12:14:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991682</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991682</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991682</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>All are viable options.<p>(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.<p>(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.<p>(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.<p>The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.</p>
]]></description><pubDate>Wed, 07 Oct 2026 12:13:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991670</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991670</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991670</guid></item><item><title><![CDATA[New comment by mjb in "Strands Decider 2B: a small, open-source, decision model"]]></title><description><![CDATA[
<p>One of the authors here.<p>Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.<p>As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).<p>Everything you need to fine-tune, or even retrain from scratch, is in the github repo.</p>
]]></description><pubDate>Wed, 07 Oct 2026 12:05:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=49991592</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49991592</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49991592</guid></item><item><title><![CDATA[New comment by mjb in "Distributed Systems Classics (2017)"]]></title><description><![CDATA[
<p>Inside multiple AWS products (including DynamoDB, Kinesis, and Aurora DSQL) is a system called Journal that moves a ton of data. It uses a variant of chain replication.<p>EBS is also a chain replication variant at heart, and moves even more data.</p>
]]></description><pubDate>Tue, 15 Sep 2026 01:04:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49706431</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49706431</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49706431</guid></item><item><title><![CDATA[New comment by mjb in "Distributed Systems Classics (2017)"]]></title><description><![CDATA[
<p>Worth mentioning that Dynamo (the classic 2007 paper) and DynamoDB (the modern AWS product) have fairly little in common architecturally.<p>For DynamoDB, check out <a href="https://www.usenix.org/conference/atc22/presentation/elhemali" rel="nofollow">https://www.usenix.org/conference/atc22/presentation/elhemal...</a> <a href="https://www.usenix.org/conference/atc23/presentation/idziorek" rel="nofollow">https://www.usenix.org/conference/atc23/presentation/idziore...</a> and my analysis of the differences here <a href="https://brooker.co.za/blog/2025/08/15/dynamo-dynamodb-dsql.html" rel="nofollow">https://brooker.co.za/blog/2025/08/15/dynamo-dynamodb-dsql.h...</a></p>
]]></description><pubDate>Tue, 15 Sep 2026 01:02:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49706421</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49706421</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49706421</guid></item><item><title><![CDATA[New comment by mjb in "Distributed Systems Classics (2017)"]]></title><description><![CDATA[
<p>Basically, many people convinced themselves that the formalization of CAP means that systems need to choose between highly available and strongly consistent, and hence chose eventual consistency. This is, partially, because Gilbert and Lynch define "availability" to mean "available to all clients, even those on a minority side of a partition". The much more useful "available to a majority of clients" is achievable at the same time as strong consistency in presence of a single partition, and this is the common cloud failure mode.<p>See <a href="https://brooker.co.za/blog/2024/07/25/cap-again.html" rel="nofollow">https://brooker.co.za/blog/2024/07/25/cap-again.html</a> for a longer take.<p>Other, much more reasonable trade-offs, lead to eventual consistency too. Mostly latency optimizations, but many of those also lead to non-zero RPO and so are undesirable for multiple reasons. We discuss some of this in section 8 of the DSQL paper: <a href="https://arxiv.org/pdf/2607.13276" rel="nofollow">https://arxiv.org/pdf/2607.13276</a></p>
]]></description><pubDate>Mon, 14 Sep 2026 23:05:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49705479</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49705479</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49705479</guid></item><item><title><![CDATA[New comment by mjb in "Distributed Systems Classics (2017)"]]></title><description><![CDATA[
<p>This is not a bad list for sure. Here are some deeper cuts for those looking for something a bit less mainstream:<p>"The Maintenance of Duplicate Databases" <a href="https://datatracker.ietf.org/doc/html/rfc677" rel="nofollow">https://datatracker.ietf.org/doc/html/rfc677</a> (AFAIK the genesis of the use of logical clocks in distributed systems).<p>"Chain Replication for Supporting High Throughput and Availability" <a href="https://www.usenix.org/legacy/event/osdi04/tech/full_papers/renesse/renesse.pdf" rel="nofollow">https://www.usenix.org/legacy/event/osdi04/tech/full_papers/...</a> (Chain replication is how a huge percentage of real-world cloud-scale data replication is done).<p>"Brewer’s Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services" (The formalization of CAP, which caused a ton of very poor trade-off thinking in the decade that followed by defining Availability in a very goofy way. Still a classic.)<p>"Paxos Made Live" <a href="https://research.google/pubs/paxos-made-live-an-engineering-perspective-2006-invited-talk/" rel="nofollow">https://research.google/pubs/paxos-made-live-an-engineering-...</a> (Brought a much-needed engineering perspective to a conversation that was largely theoretical up until this time.)<p>"Practical Byzantine fault tolerance" (Moved the conversation on Byzantine faults forward significantly).<p>This is just a short selection. There's so much good stuff going back in the 70s and 80s distributed database literature, for example (and in the modern systems and DB literature too).</p>
]]></description><pubDate>Mon, 14 Sep 2026 17:30:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49700698</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49700698</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49700698</guid></item><item><title><![CDATA[New comment by mjb in "The Theater of Sport: Bill Buford Revisits Among the Thugs"]]></title><description><![CDATA[
<p>Among the thugs, Heat, and Dirt are all great reads. Heat is my favorite of them. Highly recommended if you've ever worked in a kitchen, or considered it.</p>
]]></description><pubDate>Sun, 02 Aug 2026 04:49:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49141215</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=49141215</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49141215</guid></item><item><title><![CDATA[New comment by mjb in "MicroVMs: Run isolated sandboxes with full lifecycle control"]]></title><description><![CDATA[
<p>You absolutely can run agents on a regular VM. But if you want to build multi-tenant and multi-agent systems with strong security boundaries, then having a VM or MicroVM per agent session (or session with a group of agents) really simplifies things.<p>When we did AWS AgentCore Runtime last year we introduced session isolation, with MicroVMs per session. You can think of Lambda MicroVMs as the same stack, but generalized to fit a larger number of application patterns.</p>
]]></description><pubDate>Fri, 26 Jun 2026 18:10:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=48689958</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=48689958</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48689958</guid></item><item><title><![CDATA[New comment by mjb in "MicroVMs: Run isolated sandboxes with full lifecycle control"]]></title><description><![CDATA[
<p>AWS AgentCore runtime has been around for about a year: <a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/agents-tools-runtime.html" rel="nofollow">https://docs.aws.amazon.com/bedrock-agentcore/latest/devguid...</a> (spoiler, it's the same underlying technology as the Lambda MicroVMs).</p>
]]></description><pubDate>Fri, 26 Jun 2026 18:05:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48689878</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=48689878</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48689878</guid></item><item><title><![CDATA[New comment by mjb in "Surprising economics of load-balanced systems"]]></title><description><![CDATA[
<p>A dead comment says:<p>> Of course, this assumes independent events. World Cup, super bowls, etc break these assumptions.<p>Yes, this is very true. The model here works for Poisson arrivals and exponential service time (the M/M), which are poor approximations of real-world traffic patterns (which tend to be non-stationary and non-ergodic, and include substantial seasonality). However, the frequency of that seasonality is typically rather low (e.g. daily cycles), and so these stronger assumptions are quite defensible for short time periods.<p>A better approach is to do simulation with real traffic patterns, or even with more sophisticated parametric models, and get better answers (e.g. <a href="https://stability-sim.systems/" rel="nofollow">https://stability-sim.systems/</a>). The good news is that kind of simulation is cheaper to do than ever before.</p>
]]></description><pubDate>Fri, 19 Jun 2026 23:41:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48604659</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=48604659</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48604659</guid></item><item><title><![CDATA[New comment by mjb in "Surprising economics of load-balanced systems"]]></title><description><![CDATA[
<p>One explanation would be that more load could mean higher (absolute) variance in queue length, and therefore higher latency especially at higher percentiles. It doesn't work out that way (for reasons that Erlang actually writes about in one of his original works), but it's not an entirely unreasonable intuition.</p>
]]></description><pubDate>Fri, 19 Jun 2026 23:37:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48604626</link><dc:creator>mjb</dc:creator><comments>https://news.ycombinator.com/item?id=48604626</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48604626</guid></item></channel></rss>