<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: c_moscardi</title><link>https://news.ycombinator.com/user?id=c_moscardi</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 11 Aug 2026 19:20:11 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=c_moscardi" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by c_moscardi in "How I use LLMs to learn complex topics"]]></title><description><![CDATA[
<p>Yeah, I am currently trying to work with Opus 5 to refresh myself on deep learning fundamentals, and... it's a mixed bag. I'm glad I already am familiar with the subject matter, as I can prompt for refinement and improvement. It is kinda following the Karpathy videos so far (a couple lessons in) but adding more math/derivations, which was what I asked for. It has trouble staying on topic, presenting information in a coherent/meaningful order, and providing all the context necessary to move through steps in its "course notes".<p>Like I said, I'm essentially continually prompting to refine the material. LLMs certainly continue to append, and never cut back. It just keeps spitting out additional content at me. So that's a bit annoying too. But I can basically get figure out what's going on with a few extra promps.<p>If youre curious what i've got so far... just be warned it is quite literally AI slop plus me continually prompting for clarification/cleanup etc. : <a href="https://github.com/cmoscardi/ai-for-ai" rel="nofollow">https://github.com/cmoscardi/ai-for-ai</a></p>
]]></description><pubDate>Sun, 09 Aug 2026 21:15:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49235981</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=49235981</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49235981</guid></item><item><title><![CDATA[New comment by c_moscardi in "Privately-Owned Rail Cars"]]></title><description><![CDATA[
<p>Riding in the family rail car like it’s 1895 (and you’re a robber baron)</p>
]]></description><pubDate>Thu, 21 Aug 2025 12:53:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=44972147</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=44972147</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44972147</guid></item><item><title><![CDATA[New comment by c_moscardi in "Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast"]]></title><description><![CDATA[
<p>We chatted a few months back -- congrats on launch! Looks like a great UX.</p>
]]></description><pubDate>Mon, 23 Jun 2025 18:39:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=44358706</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=44358706</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44358706</guid></item><item><title><![CDATA[New comment by c_moscardi in "Show HN: OpenTimes – Free travel times between U.S. Census geographies"]]></title><description><![CDATA[
<p>Amazing! GitHub actions to compute a giant skim matrix is an incredible hack.<p>I pretty regularly work with social science researchers who have a need for something like this... will keep it in mind. For a bit we thought of setting something like this up within the Census Bureau, in fact. I have some stories about routing engines from my time there...</p>
]]></description><pubDate>Mon, 17 Mar 2025 21:12:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=43392785</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=43392785</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43392785</guid></item><item><title><![CDATA[New comment by c_moscardi in "Mark Cuban offers to fund former 18f employees"]]></title><description><![CDATA[
<p>Worth noting that there already exists an ecosystem of these sorts of contracting firms (nava, skylight, truss I believe, forgetting even more) that, basically, pitch themselves as the antidote to beltway banditry<p>People from the federal civic tech nexus started them up over the past decade as they termed out of 18F, USDS, and PIF</p>
]]></description><pubDate>Sun, 02 Mar 2025 16:20:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=43231976</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=43231976</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43231976</guid></item><item><title><![CDATA[New comment by c_moscardi in "AI makes tech debt more expensive"]]></title><description><![CDATA[
<p>Came to post this — it’s the same underlying technology, just a lot more compute now.</p>
]]></description><pubDate>Thu, 14 Nov 2024 21:36:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=42141483</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=42141483</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42141483</guid></item><item><title><![CDATA[New comment by c_moscardi in "Fine-tuning OCR works really well: Statistical Abstracts of the United States"]]></title><description><![CDATA[
<p>Hi HN! I've spent a couple of months fiddling with OCR and wanted to share some of my findings.<p>The approach I share here (fine-tuning recent deep learning models) is the first one that's gotten me anything resembling high-quality OCR on these particular noisy historical documents. OCRing these has been something of a white whale for me for several years (except, a white whale that I have spent comparatively little time on).<p>At this point I think I am reasonably competent in OCR, but no expert... Curious for your thoughts.</p>
]]></description><pubDate>Thu, 03 Oct 2024 16:30:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=41732253</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=41732253</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=41732253</guid></item><item><title><![CDATA[Fine-tuning OCR works really well: Statistical Abstracts of the United States]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.christianmoscardi.com/blog/2024/10/03/digitizing-us-statistical-abstracts.html">https://www.christianmoscardi.com/blog/2024/10/03/digitizing-us-statistical-abstracts.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=41732252">https://news.ycombinator.com/item?id=41732252</a></p>
<p>Points: 2</p>
<p># Comments: 1</p>
]]></description><pubDate>Thu, 03 Oct 2024 16:30:01 +0000</pubDate><link>https://www.christianmoscardi.com/blog/2024/10/03/digitizing-us-statistical-abstracts.html</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=41732252</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=41732252</guid></item><item><title><![CDATA[New comment by c_moscardi in "The waiting time paradox: why is my bus always late? (2018)"]]></title><description><![CDATA[
<p>Related reading; explains the same concept quite well IMO with NYC subway data. This is where I learned about this concept.<p>[1] <a href="https://erikbern.com/2016/04/04/nyc-subway-math" rel="nofollow">https://erikbern.com/2016/04/04/nyc-subway-math</a><p>[2] <a href="https://erikbern.com/2016/07/09/waiting-time-math.html" rel="nofollow">https://erikbern.com/2016/07/09/waiting-time-math.html</a></p>
]]></description><pubDate>Tue, 20 Aug 2024 16:29:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=41301492</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=41301492</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=41301492</guid></item><item><title><![CDATA[New comment by c_moscardi in "AI for Data Journalism: demonstrating what we can do with this stuff"]]></title><description><![CDATA[
<p>Yeah, I think MS' is the best out there, but agree that the usability leaves something to be desired. 2 thoughts:<p>1. I believe the IR jargon for getting a JSON of this form is Key Information Extraction (KIE). MS has an out-of-the-box model for this. I just tried the screenshot and it did a pretty good (but not perfect) job. It didn't get every form field, but most. MS sort-of has a flow for fine-tuning, but it really leaves a lot to be desired IMO. Curious if this would be "good enough" to satisfy the use case.<p>2. In terms of just OCR (i.e. getting the text/numeric strings correct), MS is known to be the best on typed text at the moment [1]. Handwriting is a different beast... but it looks like MS is doing a very good job there (and SOTA on handwriting is very good). In particular, it got all the numbers in that screenshot correct.<p>If you want to see the results from MS on the screenshot in this blog post, here's the entire JSON blob. A bit of a behemoth but the key/value stuff is in there: <a href="https://gist.github.com/cmoscardi/8c376094181451a49f0c62406efc875c" rel="nofollow">https://gist.github.com/cmoscardi/8c376094181451a49f0c62406e...</a><p>[1] <a href="https://mindee.github.io/doctr/latest/using_doctr/using_models.html#end-to-end-ocr" rel="nofollow">https://mindee.github.io/doctr/latest/using_doctr/using_mode...</a></p>
]]></description><pubDate>Mon, 22 Apr 2024 15:59:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=40115600</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=40115600</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=40115600</guid></item><item><title><![CDATA[New comment by c_moscardi in "Show HN: Beyond text splitting – improved file parsing for LLMs"]]></title><description><![CDATA[
<p>Figures, too! Yeah you could write some logic essentially on top of a library like this, and tune based on optimizing for some notion of recall (grab more surrounding context) and precision (direct context around the word, e.g. only the paragraph or 5 surrounding table rows) for your specific application needs.<p>Using the models underlying a library like this, there's maybe room for fine-tuning as well if you have a set of documents with specific semantic boundaries that current approaches don't capture. (And you spend an hour drawing bounding boxes to make that happen).</p>
]]></description><pubDate>Mon, 08 Apr 2024 15:24:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=39970675</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=39970675</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39970675</guid></item><item><title><![CDATA[New comment by c_moscardi in "How to get every email returned"]]></title><description><![CDATA[
<p>Funnily enough, this is another great tactic for getting emails returned (looping in someone with more leverage than you or asking them to follow up for you)!</p>
]]></description><pubDate>Sun, 26 May 2019 22:34:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=20018008</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=20018008</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=20018008</guid></item><item><title><![CDATA[New comment by c_moscardi in "Ask HN: How to incorporate machine learning into day job?"]]></title><description><![CDATA[
<p>We should talk! I do work on automatically coding products for a shipping survey at the Census Bureau. One of the earliest production uses of ML here at Census :)<p>5 minute deck:
<a href="https://github.com/codingitforward/cdfdemoday2018/blob/master/Optimizing%20the%20Commodity%20Flow%20Survey%2C%20v2.pdf" rel="nofollow">https://github.com/codingitforward/cdfdemoday2018/blob/maste...</a><p>Feel free to shoot me a message.</p>
]]></description><pubDate>Mon, 10 Dec 2018 21:40:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=18651389</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=18651389</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=18651389</guid></item><item><title><![CDATA[New comment by c_moscardi in "The hacker's guide to uncertainty estimates"]]></title><description><![CDATA[
<p>Yep, my read too. The code he provides is definitely the Clopper-Pearson formula. [1]<p>This confused me as well, since I've always used the normal central-limit-based confidence interval for binary outcomes.<p>[1] <a href="https://en.wikipedia.org/wiki/Binomial_proportion_confidence_interval#Clopper%E2%80%93Pearson_interval" rel="nofollow">https://en.wikipedia.org/wiki/Binomial_proportion_confidence...</a><p>(I am no expert in the analytic underpinnings of the beta distribution or precisely how it is the conjugate prior to the binomial -- or, rigorously speaking, what conjugate prior means -- but the formula here lines up with his formula :P )</p>
]]></description><pubDate>Tue, 16 Oct 2018 15:54:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=18230360</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=18230360</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=18230360</guid></item><item><title><![CDATA[New comment by c_moscardi in "Mithril Capital Management Is Leaving the Bay Area"]]></title><description><![CDATA[
<p>I feel like they're learning from the mistakes of their forbears and getting out before the getting's truly bad...</p>
]]></description><pubDate>Sat, 22 Sep 2018 23:26:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=18048533</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=18048533</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=18048533</guid></item><item><title><![CDATA[New comment by c_moscardi in "The Second Avenue Subway and Property Values"]]></title><description><![CDATA[
<p>Author here! I'll keep an eye on this thread and would love to discuss anything in the post that's of interest to you.</p>
]]></description><pubDate>Thu, 18 Jan 2018 16:17:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=16178232</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=16178232</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=16178232</guid></item><item><title><![CDATA[The Second Avenue Subway and Property Values]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.christianmoscardi.com/blog/2017/12/18/SAS-real-estate-values.html">https://www.christianmoscardi.com/blog/2017/12/18/SAS-real-estate-values.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=16178220">https://news.ycombinator.com/item?id=16178220</a></p>
<p>Points: 2</p>
<p># Comments: 2</p>
]]></description><pubDate>Thu, 18 Jan 2018 16:16:23 +0000</pubDate><link>https://www.christianmoscardi.com/blog/2017/12/18/SAS-real-estate-values.html</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=16178220</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=16178220</guid></item><item><title><![CDATA[New comment by c_moscardi in "Ask HN: Ubuntu Desktop Default Apps"]]></title><description><![CDATA[
<p>Awesome idea! Let me make a shameless plug for my web app that does something like this for ubuntu LTS distros:<p><a href="https://www.lilite.co/" rel="nofollow">https://www.lilite.co/</a><p>I'd love for it to be made obsolete though :)</p>
]]></description><pubDate>Fri, 21 Jul 2017 19:14:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=14823099</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=14823099</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=14823099</guid></item><item><title><![CDATA[New comment by c_moscardi in "Show HN: Lilite, a Linux package autoinstaller"]]></title><description><![CDATA[
<p>Thanks for the feedback! I wasn't particularly dedicating this to new users. This solves a problem for me, an experienced *nix user. I'm not even sure it /would/ solve a problem for new users (arguably ninite doesn't either) - you need to know what packages you care about in order to have an interest in downloading them.<p>This site gives me an interface to say, "oh yeah, I /do/ need VLC on this laptop, come to think of it", rather than having to lazy-evaluate that when I have a video downloaded and realize I don't currently have a means of playing it.<p>I spend a fair bit of time doing DevOps work, and this isn't meant to be some sort of configuration management or provisioning tool. This is for desktop / dev users, which is, I think, quite a different use case (but an increasingly relevant one). I have a dotfiles repo, which is to say I'm familiar with the "document things in scripts and repos" game - but it's not clear to me that something like a script-as-documentation is even a useful process to go through when any given computer I set up is going to be used differently. That laptop playing VLC is very different than the web dev box I spin up to mess around with a new project.</p>
]]></description><pubDate>Sun, 11 Jun 2017 10:39:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=14531468</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=14531468</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=14531468</guid></item><item><title><![CDATA[New comment by c_moscardi in "Show HN: Lilite, a Linux package autoinstaller"]]></title><description><![CDATA[
<p>Hi all, I created this site!<p>The motivation was basically a version of ninite[1] for Linux. So very desktop-focused. I actually got in touch with the ninite folks, who used to run a Linux installer [2]. They said there wasn't enough interest and package managers are "good enough." I respectfully disagree on that point - I think that having the ability to point-and-click for the core, most important packages you want to install is really useful when you're spinning up a new desktop from scratch.<p>I was running into this problem after basically doing ad-hoc setup for every new machine I turned on. Sure - I could write a shell script / list of packages and save it to my dotfiles repo, but where would the fun be in that? :)<p>(Plus, package lists on every machine are different)<p>[1] <a href="https://ninite.com/" rel="nofollow">https://ninite.com/</a>
[2] <a href="https://ninite.com/linux" rel="nofollow">https://ninite.com/linux</a></p>
]]></description><pubDate>Sat, 10 Jun 2017 21:33:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=14529305</link><dc:creator>c_moscardi</dc:creator><comments>https://news.ycombinator.com/item?id=14529305</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=14529305</guid></item></channel></rss>