Hacker News: sidmanchkanti21

New comment by sidmanchkanti21 in "[dead]"

sidmanchkanti21 — Thu, 23 Apr 2026 19:59:56 +0000

We've been building table extraction at Pulse and evaluated four benchmarks: OmniDocBench, SCORE-Bench, ParseBench, and RD-TableBench. None of them fully reflect the enterprise document workflows we've encountered in production.

TEDS (OmniDocBench) penalizes HTML formatting differences that don't affect the actual table, so the same 3x3 grid scores differently depending on whether headers use vs , and the benchmark only covers English and Chinese plus a small mixed category.

SCORE-Bench's spatial tolerance parameter can mask real failures, because if you drop a header row and shift all data up by one with delta=1, the benchmark reports high accuracy even though the column labels are gone.

ParseBench generates its ground truth with frontier VLMs (Claude Opus for tables), which introduces hallucination risk, and its TableRecordMatch metric treats tables as unordered bags of key-value records, so it doesn't penalize column transposition or row reordering. The table set is also 503 pages, English-only, with over half from a single source.

RD-TableBench linearizes tables into 1D sequences, losing horizontal vs vertical adjacency.

The RD-TableBench ground truth audit is what concerned us most. We went through all 1,000 ground truth files against the source images, and the errors consisted of scrambled text and wrong structure, garbled OCR on CJK and Arabic, and buffer artifacts where random digit sequences got appended to real numeric values. Dozens of ground truth files are byte-for-byte identical to one provider's output, and in a subset of the error cases the ground truth and that provider share the exact same specific error (same wrong word order in headers, same watermark text pulled into cells, same garbled CJK characters) while independent providers don't produce those errors.

This also motivated us to build PulseBench-Tab, a benchmark of 1,820 human-annotated tables across 9 languages and 4 scripts, with graph-based evaluation via T-LAG that operates on the parsed grid rather than the DOM tree, and fully open ground truth, scoring code, and provider outputs. Arabic and Korean both show 75+ point spreads across providers, and everything is available on HuggingFace and GitHub.

Computational complexity of schema-guided document extraction

sidmanchkanti21 — Mon, 12 Jan 2026 15:13:31 +0000

Article URL: https://www.runpulse.com/blog/computational-complexity-of-schema

Comments URL: https://news.ycombinator.com/item?id=46589576

Points: 13

# Comments: 1

New comment by sidmanchkanti21 in "Launch HN: Pulse (YC S24) – Production-grade unstructured document extraction"

sidmanchkanti21 — Thu, 18 Dec 2025 17:37:44 +0000

Thanks for testing! Glad the results work well for you

Launch HN: Pulse (YC S24) – Production-grade unstructured document extraction

sidmanchkanti21 — Thu, 18 Dec 2025 15:35:52 +0000

Hi HN, we’re Sid and Ritvik, co-founders of Pulse (https://www.runpulse.com/). Pulse is a document extraction system to create LLM-ready text using hybrid VLM + OCR models.

Here’s a demo video: https://video.runpulse.com/video/pulse-platform-walkthrough-....

Later in this post, you’ll find links to before-and-after examples on particularly tricky cases. Check those out to see what Pulse can really do! Modern vision language models are great at producing plausible text, but that makes them risky for OCR and data ingestion. Plausibility isn’t good enough when you need accuracy.

When we started working on document extraction, we assumed the same thing many teams do: foundation models are improving quickly, multi-modal systems appear to read documents well, what’s not to like? And indeed, for small or clean inputs, those assumptions mostly give good results. However, limitations show up once you begin processing real documents in volume. Long PDFs, dense tables, mixed layouts, low-fidelity scans, and financial or operational data expose errors that are subtle, hard to detect, and expensive to correct. Outputs look reasonable even though they contain small but important mistakes, especially in tables and numeric fields.

Running into those challenges got us working. We ran controlled evaluations on complex documents, fine tuned vision models, and built labeled datasets where ground truth actually matters. There have been many nights where our team stayed up hand-annotating pages, drawing bounding boxes around tables, labeling charts point by point, or debating whether a number was unreadable or simply poorly scanned. That process shaped our intuition far more than benchmarks.

One thing became clear quickly. The core challenge is not extraction itself, but confidence. Vision language models embed document images into high-dimensional representations optimized for semantic understanding, not precise transcription. That process is inherently lossy. When uncertainty appears, models tend to resolve it using learned priors instead of surfacing ambiguity. This behavior can be helpful in consumer settings. In production pipelines, it creates verification problems that do not scale well. Pulse grew out of our trying to address this gap through system design rather than prompting alone.

Instead of treating document understanding as a single generative step, our system separates layout analysis from language modeling. Documents are normalized into structured representations that preserve hierarchy and tables before schema mapping occurs. Extraction is constrained by schemas defined ahead of time, and extracted values are tied back to source locations so uncertainty can be inspected rather than guessed away. In practice, this results in a hybrid approach that combines traditional computer vision techniques, layout models, and vision language models, because no single approach handles these cases reliably on its own.

We are intentionally sharing a few documents that reflect the types of inputs that motivated this work. These are representative of cases where we saw generic OCR or VLM-based pipelines struggle.

Here is a financial 10K: https://platform.runpulse.com/dashboard/examples/example1

Here is a newspaper: https://platform.runpulse.com/dashboard/examples/example2

Here is a rent roll: https://platform.runpulse.com/dashboard/examples/example3

Pulse is not perfect, particularly on highly degraded scans or uncommon handwriting, and we’re working on improvements. However, our goal is not to eliminate errors entirely, but to make them visible, auditable, and easier to reason about.

Pulse is available via usage-based access to the API and platform You can sign up to try it at https://platform.runpulse.com/login. API docs are at https://docs.runpulse.com/introduction.

We’d love to hear how others here evaluate correctness for document extraction, which failure modes you have seen in practice, and what signals you rely on to decide whether an output can be trusted.

We will be around to answer questions and are happy to run additional documents if people want to share examples. Put links in the comments and we’ll plug them in and get back to you.

Looking forward to your comments!

Comments URL: https://news.ycombinator.com/item?id=46313930

Points: 40

# Comments: 37

Why Tables Are the Hardest Problem in Document AI

sidmanchkanti21 — Mon, 22 Sep 2025 13:06:05 +0000

Article URL: https://www.runpulse.com/blog/the-geometry-problem-why-tables-are-the-hardest-problem-in-document-ai

Comments URL: https://news.ycombinator.com/item?id=45332890

Points: 2

# Comments: 1

Evaluating Document Extraction Accuracy

sidmanchkanti21 — Tue, 12 Aug 2025 17:42:45 +0000

Article URL: https://www.runpulse.com/blog/evaluating-document-extraction

Comments URL: https://news.ycombinator.com/item?id=44879562

Points: 2

# Comments: 0

New comment by sidmanchkanti21 in "Beyond the Hype: Real-World Tests of Mistral's OCR"

sidmanchkanti21 — Fri, 07 Mar 2025 00:10:50 +0000

shoot us any questions you may have, we’ll be monitoring this closely

New comment by sidmanchkanti21 in "Putting Andrew Ng's OCR models to the test"

sidmanchkanti21 — Fri, 28 Feb 2025 07:21:19 +0000

appreciate it!

New comment by sidmanchkanti21 in "Why LLMs still have problems with OCR"

sidmanchkanti21 — Sat, 08 Feb 2025 18:14:55 +0000

thanks, glad to hear it.

New comment by sidmanchkanti21 in "Why LLMs still have problems with OCR"

sidmanchkanti21 — Sat, 08 Feb 2025 17:42:34 +0000

thanks! re: 18-19th century cursive, while we handle historical handwriting, we can't guarantee specific error rates. each document's accuracy varies based on condition, writing style, and preservation. happy to run test samples to check.

feel free to send over sample docs: sid [at] trypulse [dot] ai

New comment by sidmanchkanti21 in "Why LLMs still have problems with OCR"

sidmanchkanti21 — Fri, 07 Feb 2025 22:23:19 +0000

our current customers are both enterprises and individuals.

pricing page is here https://www.runpulse.com/pricing-studio-pulse

New comment by sidmanchkanti21 in "Why LLMs still suck at OCR"

sidmanchkanti21 — Fri, 07 Feb 2025 21:57:54 +0000

fixed the year, good catch