Hacker Newsnew | past | comments | ask | show | jobs | submit | kkm's submissionslogin
1.How to serve trillions of tokens for trillion-parameter coding agents (modal.com)
1 point by kkm 10 hours ago | past | discuss
2.Unlocking parallel test-time scaling for long-horizon agents (doubleword.ai)
1 point by kkm 16 hours ago | past | discuss
3.Claude.md is good for taste and project context.It's a weak place for invariants (tesseracted-labs-blog.vercel.app)
3 points by kkm 7 days ago | past | discuss
4.Moving coding-agent guardrails from prompts to hooks (tesseracted-labs-blog.vercel.app)
4 points by kkm 8 days ago | past | discuss
5.Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads (j9s.io)
3 points by kkm 9 days ago | past | 1 comment
6.Guide to the Kimi DeltaNet Family of linear attention (doubleword.ai)
3 points by kkm 58 days ago | past
7.Forensic Analysis of Container Snapshot Chains for Post-Event Reconstruction [pdf] (radostin.io)
2 points by kkm 64 days ago | past
8.The Agent swarm that designs itself (peterbhabra.com)
1 point by kkm 85 days ago | past
9.Don't Build a Router. Train the Small Model to Know When to Defer (distillabs.ai)
2 points by kkm 85 days ago | past | 1 comment
10.The gap between open weights LLMs and closed source LLMs (doubleword.ai)
306 points by kkm 89 days ago | past | 250 comments
11.InfiniBand, RoCE, and All That (fergusfinn.com)
5 points by kkm 3 months ago | past
12.2678x Faster Matrix Multiplication with a GPU (0mean1sigma.com)
2 points by kkm 3 months ago | past
13.UCCL-EP: DeepEP-style expert parallelism on any NIC, no GPU-initiated comms (fergusfinn.com)
9 points by kkm 3 months ago | past
14.Hacking Google with A.I. For $500k (brutecat.com)
1 point by kkm 3 months ago | past
15.How to setup a local coding agent on macOS (ikyle.me)
507 points by kkm 3 months ago | past | 127 comments
16.Anatomy of a high-performance EP kernel (fergusfinn.com)
16 points by kkm 3 months ago | past | 1 comment
17.No Token Left Behind: Demystifying Token-in-Token-Out in Miles (lmsys.org)
2 points by kkm 3 months ago | past
18.MoE expert co-activations: Reordering inputs yields easy throughput gains (doubleword.ai)
2 points by kkm 3 months ago | past
19.The Economics of Speculative Decoding (fergusfinn.com)
30 points by kkm 3 months ago | past | 6 comments
20.Speculative KV coding: losslessly compressing KV cache by up to ~4× (fergusfinn.com)
155 points by kkm 3 months ago | past | 48 comments
21.70x faster cold(ish) starts for SGLang (fergusfinn.com)
1 point by kkm 3 months ago | past
22.Bringing Up DeepSeek-V4-Flash on AMD MI300X (fergusfinn.com)
120 points by kkm 3 months ago | past | 25 comments
23.Brave AI privacy:LLMs on NEAR AI Nvidia-Backed Trusted Execution Environments (brave.com)
1 point by kkm 10 months ago | past
24.How fast can an LLM go? (fergusfinn.com)
2 points by kkm 10 months ago | past
25.FHE can be leveraged for LLMs such as ChatGPT in a privacy-preserving manner (huggingface.co)
4 points by kkm on Aug 13, 2024 | past
26.Harnessing the Power of Large Language Models for Insightful Review Analysis (holidaycheck.com)
1 point by kkm on April 16, 2024 | past
27.A Privacy-First approach to use AI for understanding our Customers Better (holidaycheck.com)
1 point by kkm on March 12, 2024 | past
28.How to make LLMs go fast (vgel.me)
2 points by kkm on Dec 19, 2023 | past
29.Leveraging Large Language Models for Sentiment Classification in Hotel Reviews (holidaycheck.com)
1 point by kkm on Nov 7, 2023 | past
30.Managers Should Think More Like Hackers (hbr.org)
3 points by kkm on April 9, 2023 | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: