<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: florianstandhar</title><link>https://news.ycombinator.com/user?id=florianstandhar</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 24 Sep 2026 01:13:51 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=florianstandhar" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by florianstandhar in "Show HN: JevBench, a reproducible benchmark for typed decision models"]]></title><description><![CDATA[
<p>Happy to — please open an issue at <a href="https://github.com/fstandhartinger/jevbench/issues" rel="nofollow">https://github.com/fstandhartinger/jevbench/issues</a> with an endpoint or runnable code, the model and licence, and whether it was trained on the public items. Every entrant runs through the same harness, including the sealed set.
I definitely prefer open models I can run locally, because sending our private held-out set of tasks to an external API produces some headaches on ur side and basically means we have to rotate the test set a lot, to avoid contaimination.
We can call API hosted models though, we'll just flag the leaderboard entries appropriately.</p>
]]></description><pubDate>Wed, 23 Sep 2026 11:32:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49814465</link><dc:creator>florianstandhar</dc:creator><comments>https://news.ycombinator.com/item?id=49814465</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49814465</guid></item><item><title><![CDATA[New comment by florianstandhar in "Show HN: JevBench, a reproducible benchmark for typed decision models"]]></title><description><![CDATA[
<p>Thanks! Yeah, the code is open: <a href="https://github.com/fstandhartinger/who-is-right" rel="nofollow">https://github.com/fstandhartinger/who-is-right</a>.
Email classifying is definitely a good Jev usecase, I agree</p>
]]></description><pubDate>Wed, 23 Sep 2026 11:29:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49814443</link><dc:creator>florianstandhar</dc:creator><comments>https://news.ycombinator.com/item?id=49814443</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49814443</guid></item><item><title><![CDATA[New comment by florianstandhar in "Show HN: JevBench, a reproducible benchmark for typed decision models"]]></title><description><![CDATA[
<p>Good call... BART-large-MNLI is the classic zero-shot baseline and it should be there. Adding it to the next run.</p>
]]></description><pubDate>Wed, 23 Sep 2026 11:26:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49814408</link><dc:creator>florianstandhar</dc:creator><comments>https://news.ycombinator.com/item?id=49814408</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49814408</guid></item><item><title><![CDATA[New comment by florianstandhar in "OpenAI is about to eat Jev's lunch – Arcturus Labs"]]></title><description><![CDATA[
<p>maybe open source even eats Jevs lunch first<p>see here: <a href="https://news.ycombinator.com/item?id=49800574">https://news.ycombinator.com/item?id=49800574</a></p>
]]></description><pubDate>Tue, 22 Sep 2026 15:09:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49802610</link><dc:creator>florianstandhar</dc:creator><comments>https://news.ycombinator.com/item?id=49802610</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49802610</guid></item><item><title><![CDATA[Show HN: JevBench, a reproducible benchmark for typed decision models]]></title><description><![CDATA[
<p>Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects <i>really</i> perform in comparison.<p>Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.<p>JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.<p>A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.<p>Leaderboard right now:<p><pre><code>  #1 - Jev            74.4
  #2 - SemIf          73.1
  #3 - djev           73.0
  #4 - Winnow-12B Q8  71.2
  #5 reflex 4B        70.3.
</code></pre>
MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:<p><a href="https://github.com/fstandhartinger/jevbench" rel="nofollow">https://github.com/fstandhartinger/jevbench</a><p>Two no-signup demos:<p><a href="https://who-is-right.app.mintapis.com" rel="nofollow">https://who-is-right.app.mintapis.com</a><p><a href="https://is-it-ai-slop.app.mintapis.com" rel="nofollow">https://is-it-ai-slop.app.mintapis.com</a><p>Limitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.<p>Wdyt?</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49800574">https://news.ycombinator.com/item?id=49800574</a></p>
<p>Points: 139</p>
<p># Comments: 36</p>
]]></description><pubDate>Tue, 22 Sep 2026 13:01:03 +0000</pubDate><link>https://benchmarkheaven.com/jev-models</link><dc:creator>florianstandhar</dc:creator><comments>https://news.ycombinator.com/item?id=49800574</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49800574</guid></item></channel></rss>