<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: technoabsurdist</title><link>https://news.ycombinator.com/user?id=technoabsurdist</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 04 Aug 2026 02:26:16 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=technoabsurdist" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by technoabsurdist in "Running Kimi K3 on MI355X at Better Performance per Dollar Than B300"]]></title><description><![CDATA[
<p>hi I work at wafer. yes we ran benchmarks. for example our Kimi K3 is live on open router and in order to host there you have to run accuracy checks. like tau/gpqa.<p>and then u must pass test regarding thinking, coherency, and tool calling</p>
]]></description><pubDate>Sun, 02 Aug 2026 16:29:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49145961</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=49145961</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49145961</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>sounds good feedback taken, thanks beffjezos</p>
]]></description><pubDate>Sat, 04 Jul 2026 05:52:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=48782959</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=48782959</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48782959</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>hi yes it’s not optimized for single stream it’s optimized for total node throughput</p>
]]></description><pubDate>Sat, 04 Jul 2026 03:47:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48782484</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=48782484</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48782484</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>this is exactly our thesis at wafer :) thank you for the support</p>
]]></description><pubDate>Sat, 04 Jul 2026 00:22:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48781541</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=48781541</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48781541</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>AMD MI355X uses 1,400W per GPU and NVIDIA B200 uses 1,200W. So AMD uses about 16% more power.</p>
]]></description><pubDate>Sat, 04 Jul 2026 00:21:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48781534</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=48781534</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48781534</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>yes it is 213 tok/s single stream (so per user)</p>
]]></description><pubDate>Fri, 03 Jul 2026 23:53:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48781393</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=48781393</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48781393</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>hi i work at wafer. no the margins are lower averaging at about ~40%. utilization is one of the highest order bits in determining margins here, yes.</p>
]]></description><pubDate>Fri, 03 Jul 2026 23:52:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=48781390</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=48781390</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48781390</guid></item><item><title><![CDATA[AI Could Democratize One of Techs Most Valuable Resources]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.wired.com/story/ai-could-democratize-one-of-techs-most-valuable-resources/">https://www.wired.com/story/ai-could-democratize-one-of-techs-most-valuable-resources/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47783621">https://news.ycombinator.com/item?id=47783621</a></p>
<p>Points: 6</p>
<p># Comments: 2</p>
]]></description><pubDate>Wed, 15 Apr 2026 18:59:53 +0000</pubDate><link>https://www.wired.com/story/ai-could-democratize-one-of-techs-most-valuable-resources/</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=47783621</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47783621</guid></item><item><title><![CDATA[Show HN: Wafer – Profile, inspect assembly, and iterate on CUDA within your IDE]]></title><description><![CDATA[
<p>Hi HN, I’m Emilio. We’re launching the Wafer extension for the popular IDEs (VS Code, Cursor and Antigravity).<p>Wafer exists to make performance engineers more efficient. Most of the work perf engs do is extracting signal and turning it into the next experiment. You spend hours per kernel doing interpretation and bookkeeping: which counters matter, what changed, what hypothesis you’re testing, what to try next.<p>Wafer is building an environment where profiling, compiler analysis, and docs are first-class context in your workflow, so iteration is cheap. long-term, that same structured context becomes the interface for an automation layer that can read the evidence, propose a change, and rerun the loop.<p>NVIDIA has poured an insane amount of truth into their tooling. NCU, compiler output, SASS, the counters, the sections, the warnings, the “this is why you’re slow” breadcrumbs. Serious perf engineers already live in this stuff. The real problem is that it’s still not packaged as a tight loop. You run a profile, you get a giant report, then you spend a bunch of time translating it into a plan, mapping it back to the right lines of code, deciding what to ignore, deciding what to try next, and keeping track of what you’ve already tested. That translation step is where a ton of time goes, and it’s also the part that doesn’t scale.<p>We're just starting out and today, Wafer makes that translation step cheaper by keeping the evidence and the code in one place. You can run Nsight Compute profiling from your editor and view results where you’re editing, so you’re not flipping between terminals, report viewers, and screenshots. You can compile CUDA and inspect PTX and SASS mapped back to your source, so “what did the compiler actually do” is something you can answer in seconds and iterate on quickly. And you can query GPU documentation from inside the editor with the exact context you’re working in.<p>What we’re adding and moving towards is making that loop not just faster, but more automatic and more reproducible. We’re rolling out GPU Workspaces, where you keep a persistent CPU environment for your repo and dependencies, and only spin up GPU execution when you actually run something. A lot of GPU dev time is editing, debugging, and iterating on hypotheses, not burning GPU cycles - but today the workflow forces you to keep a GPU box alive just to preserve state. We want the “run the experiment” part to be on-demand and reliable, without killing your environment.<p>The bigger direction is the same theme: take the evidence perf engineers already use and make it machine-legible, so an automation layer can actually act on it. We're working on tool-driven loops: read the profile, identify the highest leverage bottleneck, propose a concrete code change, run the diff, re-profile, and keep a history of what worked and what didn’t.<p>If you’ve ever wished you could hand an agent your kernel plus the profiler and compiler evidence and have it do real work instead of vibes, that’s what we’re building towards.<p>You can see more about us here: <a href="https://wafer.ai">https://wafer.ai</a><p>Or download directly from here: 
VS Code: <a href="https://marketplace.visualstudio.com/items?itemName=Wafer.wafer" rel="nofollow">https://marketplace.visualstudio.com/items?itemName=Wafer.wa...</a>
Cursor: <a href="https://open-vsx.org/extension/wafer/wafer" rel="nofollow">https://open-vsx.org/extension/wafer/wafer</a><p>Would love feedback from anyone doing CUDA, CUTLASS/CuTe, Triton, training or inference perf. If you try it and something feels slow, confusing, or missing, email me at emilio@wafer.ai</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46367310">https://news.ycombinator.com/item?id=46367310</a></p>
<p>Points: 3</p>
<p># Comments: 1</p>
]]></description><pubDate>Tue, 23 Dec 2025 17:43:14 +0000</pubDate><link>https://www.wafer.ai/</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=46367310</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46367310</guid></item><item><title><![CDATA[Show HN: GPU Profiling That's Useful in 60 Seconds]]></title><description><![CDATA[
<p>Hey HN! We're building a profiler for ML inference that actually shows what's happening at the hardware level without having to manually parse through flame graphs, or set up nsys and ncu.<p>The problem: Current ML profilers either dump too much data (torch.profiler) or abstract away the details you need. You can't see why your model is actually slow - is it memory bandwidth? Kernel launch overhead? Cache misses?<p>Our approach: We're reverse engineering GPU execution to trace from Python ops down to PTX instructions. One decorator gives you the full execution graph with actual bottlenecks highlighted.<p>Technical details:
- Traces Python → CUDA kernels → PTX with timing breakdowns
- Shows memory access patterns and bandwidth utilization  
- Kernel occupancy and scheduling analysis
- Works with PyTorch/JAX, TensorFlow coming<p>We used this to optimize Llama inference and found bottlenecks we couldn't see before - got 50%+ speedup: <a href="https://www.herdora.com/blog/the-overlooked-gpu">https://www.herdora.com/blog/the-overlooked-gpu</a><p>Free beta with 10 hours of profiling: <a href="https://keysandcaches.com" rel="nofollow">https://keysandcaches.com</a>
Github: <a href="https://github.com/Herdora/kandc" rel="nofollow">https://github.com/Herdora/kandc</a>
Docs: <a href="https://www.keysandcaches.com/docs" rel="nofollow">https://www.keysandcaches.com/docs</a><p>Curious what inference bottlenecks others are hitting that current tools can't diagnose. What's your experience with existing profilers? Would be very useful to hear thoughts from the community :)</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44850506">https://news.ycombinator.com/item?id=44850506</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 09 Aug 2025 21:37:39 +0000</pubDate><link>https://www.keysandcaches.com/</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44850506</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44850506</guid></item><item><title><![CDATA[Show HN: We made PyTorch profiling usable for ML engineers]]></title><description><![CDATA[
<p>If you've ever tried to profile PyTorch or other Python code, you know that it can be painful to setup right. Nsight profiler is great, but it feels like using a sledgehammer to hang a picture frame when you're coming from a Python background.<p>We're building Chisel to solve this. Instead of wrestling with nvtx events and complex CUDA interfaces, you just add decorators to the functions you want to profile. As simple as:<p>app = ChiselApp("my-app", gpu=GPUType.A100_80GB_1)<p>@app.capture_trace(trace_name="gpu_task")<p>These 2 lines will get you the profiling trace for your code on an A100.<p>Run your code normally, then view results in a clean dashboard or locally.<p>Documentation: <a href="https://herdora.mintlify.app/" rel="nofollow">https://herdora.mintlify.app/</a><p>If you sign up this week, there's $50 in free credits to try it out (our profiling runs are billed by the second, so this is a ton of free profiling runs!)<p>We'd love feedback from anyone who gives it a shot. We're also currently developing a C++ SDK for teams that need to profile lower-level code. Let us know if this would be of interest to you. We want to learn what profiling pain points have you run into with ML codebases. Please reach out if you have any fun stories to tell :). We very much welcome contributions!<p>The github repo is: <a href="https://github.com/Herdora/chisel">https://github.com/Herdora/chisel</a></p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44779115">https://news.ycombinator.com/item?id=44779115</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 03 Aug 2025 19:40:29 +0000</pubDate><link>https://herdora.mintlify.app/</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44779115</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44779115</guid></item><item><title><![CDATA[Chip Benchmark: Hardware-Centric Performance Insights for AI Workloads]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.herdora.com/blog/chip-benchmark">https://www.herdora.com/blog/chip-benchmark</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44613348">https://news.ycombinator.com/item?id=44613348</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 19 Jul 2025 07:30:14 +0000</pubDate><link>https://www.herdora.com/blog/chip-benchmark</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44613348</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44613348</guid></item><item><title><![CDATA[New comment by technoabsurdist in "ChipBenchmark: Open-Source Benchmarking for LLM Performance Across Hardware"]]></title><description><![CDATA[
<p>^ We currently just have llama3.1-8b, so we'll be working on adding more models across more hardware options!</p>
]]></description><pubDate>Tue, 15 Jul 2025 02:56:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=44567461</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44567461</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44567461</guid></item><item><title><![CDATA[New comment by technoabsurdist in "ChipBenchmark: Open-Source Benchmarking for LLM Performance Across Hardware"]]></title><description><![CDATA[
<p>We just launched Chip Benchmark, an open-source tool for hardware-centric benchmarking of open-weight LLMs across accelerators like NVIDIA A100/H100/L40S and AMD MI300X. It measures throughput, latency, and time-to-first-token with transparent scripts and an interactive web dashboard—making apples-to-apples comparisons easier.<p>We're actively welcoming contributions, new hardware support, and benchmark requests.<p>Repo here: <a href="https://github.com/Herdora/chip-benchmark">https://github.com/Herdora/chip-benchmark</a>
Dashboard: <a href="https://herdora.com/benchmark">https://herdora.com/benchmark</a><p>Feedback and contributions welcome! We made it super easy to add other architectures by including the script we used for benchmarking.</p>
]]></description><pubDate>Tue, 15 Jul 2025 02:55:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=44567455</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44567455</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44567455</guid></item><item><title><![CDATA[ChipBenchmark: Open-Source Benchmarking for LLM Performance Across Hardware]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.chipbenchmark.com/">https://www.chipbenchmark.com/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44567454">https://news.ycombinator.com/item?id=44567454</a></p>
<p>Points: 3</p>
<p># Comments: 2</p>
]]></description><pubDate>Tue, 15 Jul 2025 02:55:32 +0000</pubDate><link>https://www.chipbenchmark.com/</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44567454</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44567454</guid></item><item><title><![CDATA[Using AMD MI300X for High-Throughput, Low-Cost LLM Inference]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.herdora.com/blog/the-overlooked-gpu">https://www.herdora.com/blog/the-overlooked-gpu</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44544923">https://news.ycombinator.com/item?id=44544923</a></p>
<p>Points: 8</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 12 Jul 2025 20:31:54 +0000</pubDate><link>https://www.herdora.com/blog/the-overlooked-gpu</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44544923</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44544923</guid></item><item><title><![CDATA[New comment by technoabsurdist in "Profile CUDA kernels with one command, zero GPU setup"]]></title><description><![CDATA[
<p>We've been doing lots of GPU kernel profiling and optimization on cloud infrastructure, but without local GPU hardware, that meant constant SSH juggling: upload code, compile remotely, profile kernels, download results, repeat. Or, work entirely on cloud which is expensive, slow, and annoying. We were spending more time managing infrastructure than writing the kernels we wanted to optimize.<p>So we built Chisel: one command to run profiling commands on any kernel. Zero local GPU hardware required.<p>Next up we're planning to build a web dashboard for visualizing results, simultaneous profiling across multiple GPU types, and automatic resource cleanup. But please let us know what you would like to see in this project.<p>Available via PyPI: pip install chisel-cli<p>Github: <a href="https://github.com/Herdora/chisel">https://github.com/Herdora/chisel</a><p>We're actively developing and would love community feedback. Feature requests and contributions always welcome!</p>
]]></description><pubDate>Fri, 04 Jul 2025 09:52:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=44463018</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44463018</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44463018</guid></item><item><title><![CDATA[Profile CUDA kernels with one command, zero GPU setup]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/Herdora/chisel">https://github.com/Herdora/chisel</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44463017">https://news.ycombinator.com/item?id=44463017</a></p>
<p>Points: 3</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 04 Jul 2025 09:52:48 +0000</pubDate><link>https://github.com/Herdora/chisel</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44463017</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44463017</guid></item><item><title><![CDATA[Show HN: Profile GPU Kernels with One Command, Zero GPU Setup]]></title><description><![CDATA[
<p>We've been doing lots of GPU kernel profiling and optimization on cloud infrastructure, but without local GPU hardware, that meant constant SSH juggling: upload code, compile remotely, profile kernels, download results, repeat. We were spending more time managing infrastructure than writing optimized kernels.<p>So we built Chisel: one command spins up a cloud instance, syncs your GPU code, runs profiling (supports Python/PyTorch/TensorFlow and AMD rocprofv3), and automatically pulls results back. Zero local GPU hardware required.<p>Next up: Web dashboard for visualizing results, simultaneous profiling across multiple GPU types, and automatic resource cleanup.<p>Available via PyPI: pip install chisel-cli<p>Github: <a href="https://github.com/Herdora/chisel">https://github.com/Herdora/chisel</a><p>We're actively developing and would love community feedback, especially from GPU devs exploring AMD alternatives to NVIDIA. Feature requests and contributions welcome!</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44416042">https://news.ycombinator.com/item?id=44416042</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 29 Jun 2025 20:15:30 +0000</pubDate><link>https://github.com/Herdora/chisel</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44416042</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44416042</guid></item><item><title><![CDATA[Show HN: Chisel – Profile GPU Kernels Without a GPU (Nvidia and AMD)]]></title><description><![CDATA[
<p>We built Chisel to make GPU kernel profiling hardware-free. It lets you run chisel profile kernel.cu and get full Nsight/Ncompute or rocprofv3 reports without a GPU needed.<p>It spins up remote H100, L40S, or MI300X machines (via DigitalOcean for now, but gonna expand backends soon), runs your code, and gives you back detailed traces (kernel timings, memory transfers, API calls, etc). Everything is CLI-based and designed for iterative dev—profiling takes \~1–2 minutes per run.<p>For example:<p># Profile a PyTorch training script on H100 with Nsight Systems
chisel profile --nsys train.py<p># Profile a HIP kernel on MI300X with system trace
chisel profile --rocprofv3="--sys-trace" matrix_add.cpp<p>Repo: <a href="https://github.com/Herdora/chisel">https://github.com/Herdora/chisel</a>
PyPI: pip install chisel-cli<p>Would love feedback! especially from anyone building custom kernels, ML layers, or low-level GPU ops.</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44391255">https://news.ycombinator.com/item?id=44391255</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Thu, 26 Jun 2025 20:55:17 +0000</pubDate><link>https://github.com/Herdora/chisel</link><dc:creator>technoabsurdist</dc:creator><comments>https://news.ycombinator.com/item?id=44391255</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44391255</guid></item></channel></rss>