<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: reexpressionist</title><link>https://news.ycombinator.com/user?id=reexpressionist</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 10 Oct 2026 04:15:00 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=reexpressionist" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by reexpressionist in "Typesafe AI raises $870M at $7.5B"]]></title><description><![CDATA[
<p>A major limitation of directly using "System 1 decision models" (a.k.a., logistic regression, and related uncalibrated classifiers over the output/logit space) in enterprise settings (or other high-stakes settings) is that such estimators are not reliable estimators of the predictive uncertainty in the presence of covariate shifts, and such estimators also lack a means of instance-wise data attribution (i.e., interpretability-by-exemplar), so they're not the ideal estimator for verification, routing, uncertainty over retrieval and tool-calls, etc.<p>For decision-making with neural networks, we instead need the older idea of estimators of the predictive uncertainty with constraints in the feature-representation space (over training/support of the estimator), as with Similarity-Distance-Magnitude estimators: <a href="https://pypi.org/project/reexpress-sdm/" rel="nofollow">https://pypi.org/project/reexpress-sdm/</a></p>
]]></description><pubDate>Sat, 10 Oct 2026 00:34:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=50028266</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=50028266</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50028266</guid></item><item><title><![CDATA[New comment by reexpressionist in "Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers"]]></title><description><![CDATA[
<p>The key properties for using such models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).<p>The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:<p>[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.<p>[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.</p>
]]></description><pubDate>Fri, 02 Oct 2026 19:54:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49937836</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=49937836</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49937836</guid></item><item><title><![CDATA[New comment by reexpressionist in "Clef: Open-weight decision models, and new RL fine-tuning platform"]]></title><description><![CDATA[
<p>> "This means that a human does not necessarily need to be in the loop for agentic decisions anymore"<p>That's only true in a practical sense if you can actually rely on the probabilities estimated by the model. That's a non-trivial problem for multiple reasons, among them: 1. What is the particular quantity you seek to estimate (marginal, approximately conditional, etc.)? 2. What is the reference class for that quantity? 3. What method are you going to use to estimate that quantity? 4. What is the error in your method to estimate that quantity (e.g., as via accounting for the effective sample size)?<p>A further practical challenge is that the output logits of neural networks are in effect a highly lossy compression of the epistemic (reducible) uncertainty. Even if your estimates are well-calibrated (for some definition of well-calibrated) on in-distribution data using the output logits, those estimates can be grossly uncalibrated in the presence of covariate shifts, and the logits themselves are not reliable signals of such shifts, nor of being out-of-distribution. Informally, the output logits themselves do not encode a good sense of what they do[n't] know.<p>Additionally, ideally the probability estimates are interpretable in the sense that there is some instance-wise connection to the training/calibration data. If the estimates are being used for decision-making, you need to be able to post-hoc audit the estimates to be able to modify the data for future decision-making, if needed.<p>Growing evidence in ML/NLP/Stats from the last few years is that with neural networks, as a starting point for constructing reliable estimates of the predictive uncertainty, we need to control for metric-learner signals over the support/training set (e.g., the L^2 distance to the nearest training instance and depth-matches into training). Once you have that, then you can choose your desired quantity of interest (e.g., class- and prediction-conditional accuracy at least some given value). Concretely, here's a tutorial (along with Apache-2.0 code) that steps through a simple, illustrative example: <a href="https://reexpressai.github.io/reexpress_sdm/tutorials/getting_started_tutorial.html" rel="nofollow">https://reexpressai.github.io/reexpress_sdm/tutorials/gettin...</a><p>More context is in the link, but at a high-level from an engineering perspective, just as dense vector matching is used by RAG for information retrieval, we can also use dense vector matching in this way to estimate the predictive uncertainty, getting around the limitations of the output logits.</p>
]]></description><pubDate>Fri, 02 Oct 2026 02:33:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49929308</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=49929308</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49929308</guid></item><item><title><![CDATA[How to Calibrate Jev (and Related)]]></title><description><![CDATA[
<p>Article URL: <a href="https://pypi.org/project/reexpress-sdm/">https://pypi.org/project/reexpress-sdm/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49915210">https://news.ycombinator.com/item?id=49915210</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 30 Sep 2026 22:18:37 +0000</pubDate><link>https://pypi.org/project/reexpress-sdm/</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=49915210</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49915210</guid></item><item><title><![CDATA[New comment by reexpressionist in "On Sleeper Agent LLMs"]]></title><description><![CDATA[
<p>This type of behavior (and related) would primarily only be an issue with unconstrained generative models. If you're the one deploying the model, or a downstream consumer, once trained, the neural network can be reexpressed (via an exogenous/secondary model/process) to derive reliable and interpretable uncertainty quantification by conditioning on reference classes (in a held-out Calibration set) formed by the Similarity to Training (depth-matches to training), Distance to Training, and a CDF-based per-class threshold on the output magnitude. If the prediction/output falls below the desired probability threshold, gracefully fail by rejecting the prediction, rather than allowing silent errors to accumulate.<p>For higher-risk settings, you can always turn the crank to be more conservative (i.e., more stringent parameters and/or requiring a larger sample size in the highest probability and reliability data partition).<p>For classification tasks, this follows directly. For generative output, this comes into play with the final verification classifier used over the output.</p>
]]></description><pubDate>Sat, 13 Jan 2024 18:58:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=38983268</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38983268</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38983268</guid></item><item><title><![CDATA[New comment by reexpressionist in "LLaMA-Pro-8B"]]></title><description><![CDATA[
<p>I believe we have two rather different settings in mind. My statement assumes the enterprise use-case, where having a verifier is required. (In this context, I'm also assuming the approach of constraining against the observed data.) In such a selective classification setting, the end-user need not be exposed to lower quality outputs, but rather null predictions if the model cascade has been exhausted (i.e., progressively moving to larger models until the probability is acceptable).<p>Hopefully in 2024 we can get at least one of the benchmarks to move to assessing non-parametric/distribution-free uncertainty for selective classification, reflecting more recent CS/Stats advances that should be used in practice. Working on it.</p>
]]></description><pubDate>Sun, 07 Jan 2024 18:25:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=38903690</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38903690</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38903690</guid></item><item><title><![CDATA[New comment by reexpressionist in "LLaMA-Pro-8B"]]></title><description><![CDATA[
<p>The alternative approach is to start with a small[er] model, but derive reliable uncertainty estimates, only moving to a larger model if necessary (i.e., if the probability of the predictions is lower than needed for the task).<p>And I agree that the leaderboards don't currently reflect the quantities of interest typically needed in practice.</p>
]]></description><pubDate>Sat, 06 Jan 2024 21:23:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=38895631</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38895631</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38895631</guid></item><item><title><![CDATA[New comment by reexpressionist in "Best 7B LLM on leaderboards made by an amateur following a medium tutorial"]]></title><description><![CDATA[
<p>If the end goal is document classification and/or semantic search, the Reexpress Fast I model (3.2 billion parameters) is a good choice. The key is that it produces reliable uncertainty estimates (for classification), so you know if you need a larger (or alternative) model. (In fact, an argument can be made that since the other models don't produce such uncertainty estimates, they are not ideal for serious use cases without adding an additional mechanism, such as ensembling with the Reexpress model.)</p>
]]></description><pubDate>Sat, 06 Jan 2024 21:04:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=38895449</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38895449</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38895449</guid></item><item><title><![CDATA[New comment by reexpressionist in "Efficient LLM fine-tuning for classification on Mac"]]></title><description><![CDATA[
<p>TL;DR: Reexpress makes it really easy (and inexpensive) to fine-tune a large language model (LLM) for typical document classification tasks. All of the processing happens on your Mac and you also get the indispensable additional advantages of uncertainty quantification, interpretability by example/exemplar, and semantic search capabilities.</p>
]]></description><pubDate>Fri, 05 Jan 2024 15:39:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=38880176</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38880176</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38880176</guid></item><item><title><![CDATA[Efficient LLM fine-tuning for classification on Mac]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial3_financial_sentiment_comparison_to_genai_finetuning/README.md">https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial3_financial_sentiment_comparison_to_genai_finetuning/README.md</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=38880175">https://news.ycombinator.com/item?id=38880175</a></p>
<p>Points: 1</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 05 Jan 2024 15:39:03 +0000</pubDate><link>https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial3_financial_sentiment_comparison_to_genai_finetuning/README.md</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38880175</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38880175</guid></item><item><title><![CDATA[How to locally run a semantic search with representations fine-tuned on your Mac]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial4_semantic_search_without_labels/README.md">https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial4_semantic_search_without_labels/README.md</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=38852637">https://news.ycombinator.com/item?id=38852637</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 03 Jan 2024 10:46:39 +0000</pubDate><link>https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial4_semantic_search_without_labels/README.md</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38852637</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38852637</guid></item><item><title><![CDATA[Show HN: On-device, no-code LLMs with guardrails (for Apple Silicon)]]></title><description><![CDATA[
<p>We've been working to make uncertainty quantification and interpretability first-class properties of LLMs. Reexpress one, a macOS app, is our first effort to make these properties widely available.<p>Perhaps counter-intuitively, and contrary to common wisdom, LLMs can in fact be transformed to generate very reliable uncertainty estimates (i.e., "knowing what they do and don't know" by assigning a probability to the output).<p>Getting there is a bit complicated, with vector matching/databases, prediction-time data dependencies, complicated inference, and multiple models flying all over the place.<p>We've made it simple and efficient to use in practice with an on-device, no-code approach. Common document classification tasks can be handled with the on-device models (up to 3.2 billion parameters). Additionally, you can add these capabilities to another LLM (e.g., for QA or more complicated tasks) by connecting your existing model by simply uploading the output logits into the app. For example, if you're using an on-device Mistral AI model, or cloud-based genAI model, just upload the output logits into the app.<p>Would be great to get feedback. Also, if you have another use case with a scale that doesn't fully fit into the on-device setting, happy to discuss and collaborate for your setting.<p>And if anyone finds this interesting and wants to get involved more in building reliable AI, let us know!<p>(Note that an Apple silicon Mac is required; ideally M1 Max or better with 64gb of RAM. You train the model yourself, which requires labeled data. The tutorial 1 video has a link to sentiment data in the JSON lines format; it's a good place to start: <a href="https://github.com/ReexpressAI/Example_Data/blob/main/tutorials/tutorial1_sentiment/README.md">https://github.com/ReexpressAI/Example_Data/blob/main/tutori...</a>)</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=38649799">https://news.ycombinator.com/item?id=38649799</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Fri, 15 Dec 2023 01:07:41 +0000</pubDate><link>https://re.express</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38649799</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38649799</guid></item><item><title><![CDATA[New comment by reexpressionist in "AI and Trust"]]></title><description><![CDATA[
<p>Important essay and points. I want to mention that there exist now practical technical approaches that can be used to create trustworthy AI...and such approaches can be run on local models, as this comment suggests.<p>> "[...] [AI] will act trustworthy, but it will not be trustworthy. We won’t know how they are trained. We won’t know their secret instructions. We won’t know their biases, either accidental or deliberate. [...]"<p>I agree that this is true with standard deployments of the generative AI models, but we can instead reframe networks as a direct connection between the observed/known data and new predictions, and to tightly constrain predictions against the known labels. In this way, we can have controllable oversight of biases, out-of-distribution errors, and more broadly, a clear relation to the task-specific training data.<p>That is to say, I believe the concerns in the essay are valid in that they reflect one possible path in the current fork in the road, but it is not inevitable, given the potential of reliable, on-device, personal AI.</p>
]]></description><pubDate>Tue, 05 Dec 2023 01:54:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=38526023</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38526023</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38526023</guid></item><item><title><![CDATA[New comment by reexpressionist in "LLM Visualization"]]></title><description><![CDATA[
<p>Ditto. This is the most sophisticated viz of parameters I've seen...and it's also an interactive, step-through tutorial!</p>
]]></description><pubDate>Sun, 03 Dec 2023 20:06:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=38510263</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38510263</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38510263</guid></item><item><title><![CDATA[New comment by reexpressionist in "Good old-fashioned AI remains viable in spite of the rise of LLMs"]]></title><description><![CDATA[
<p>"As with most software development, modern AI work is all about knowing your tools and when it's appropriate to use them." 100% agree. Ditto with just using the easiest to access models as initial proof-of-concept/dev/etc. to get started.<p>(I do agree with the overall sentiment of the TC article, although as noted by others below, there's some mashing of terminology in the article. E.g., I, too, associate GOFAI with symbolic AI and planning.)<p>There's another dimension, too, not mentioned in the article: Even with general purpose LLMs, for production applications, it's still required to have labeled data to produce uncertainty estimates. (There's a sense in which any well-defined and tested production application is a 'single-task' setting, in it its own way.) One of the reasons on-device/edge AI has gotten so interesting, in my opinion, is that we now know how to derive reliable uncertainty estimates with the neural models (more or less independent of scale). As long as prediction uncertainty is sufficiently low, there's no particular reason to go to a larger model. That can lead to non-trivial cost/resource savings, as well as the other benefits of keeping things on-device.</p>
]]></description><pubDate>Sun, 03 Dec 2023 04:42:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=38504892</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38504892</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38504892</guid></item><item><title><![CDATA[New comment by reexpressionist in "Are Open-Source Large Language Models Catching Up?"]]></title><description><![CDATA[
<p>I like the analogy to a router and local Mixture of Experts; that's basically how I see things going, as well. (Also, agreed that Huggingface has really gone far in making it possible to build such systems across many models.)<p>There's also another related sense for which we want routing across models for efficiency reasons in the local setting, even for tasks for the same input modalities:<p>First, attempt prediction on small(er) models, and if the constrained output is not sufficiently high probability (with highest calibration reliability), route to progressively larger models. If the process is exhausted, kick it to a human for further adjudication/checking.</p>
]]></description><pubDate>Sat, 02 Dec 2023 00:03:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=38494312</link><dc:creator>reexpressionist</dc:creator><comments>https://news.ycombinator.com/item?id=38494312</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38494312</guid></item></channel></rss>