Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
kkm's submissions
login
1.
How to serve trillions of tokens for trillion-parameter coding agents
(
modal.com
)
1 point
by
kkm
10 hours ago
|
past
|
discuss
2.
Unlocking parallel test-time scaling for long-horizon agents
(
doubleword.ai
)
1 point
by
kkm
16 hours ago
|
past
|
discuss
3.
Claude.md is good for taste and project context.It's a weak place for invariants
(
tesseracted-labs-blog.vercel.app
)
3 points
by
kkm
7 days ago
|
past
|
discuss
4.
Moving coding-agent guardrails from prompts to hooks
(
tesseracted-labs-blog.vercel.app
)
4 points
by
kkm
8 days ago
|
past
|
discuss
5.
Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads
(
j9s.io
)
3 points
by
kkm
9 days ago
|
past
|
1 comment
6.
Guide to the Kimi DeltaNet Family of linear attention
(
doubleword.ai
)
3 points
by
kkm
58 days ago
|
past
7.
Forensic Analysis of Container Snapshot Chains for Post-Event Reconstruction [pdf]
(
radostin.io
)
2 points
by
kkm
64 days ago
|
past
8.
The Agent swarm that designs itself
(
peterbhabra.com
)
1 point
by
kkm
85 days ago
|
past
9.
Don't Build a Router. Train the Small Model to Know When to Defer
(
distillabs.ai
)
2 points
by
kkm
85 days ago
|
past
|
1 comment
10.
The gap between open weights LLMs and closed source LLMs
(
doubleword.ai
)
306 points
by
kkm
89 days ago
|
past
|
250 comments
11.
InfiniBand, RoCE, and All That
(
fergusfinn.com
)
5 points
by
kkm
3 months ago
|
past
12.
2678x Faster Matrix Multiplication with a GPU
(
0mean1sigma.com
)
2 points
by
kkm
3 months ago
|
past
13.
UCCL-EP: DeepEP-style expert parallelism on any NIC, no GPU-initiated comms
(
fergusfinn.com
)
9 points
by
kkm
3 months ago
|
past
14.
Hacking Google with A.I. For $500k
(
brutecat.com
)
1 point
by
kkm
3 months ago
|
past
15.
How to setup a local coding agent on macOS
(
ikyle.me
)
507 points
by
kkm
3 months ago
|
past
|
127 comments
16.
Anatomy of a high-performance EP kernel
(
fergusfinn.com
)
16 points
by
kkm
3 months ago
|
past
|
1 comment
17.
No Token Left Behind: Demystifying Token-in-Token-Out in Miles
(
lmsys.org
)
2 points
by
kkm
3 months ago
|
past
18.
MoE expert co-activations: Reordering inputs yields easy throughput gains
(
doubleword.ai
)
2 points
by
kkm
3 months ago
|
past
19.
The Economics of Speculative Decoding
(
fergusfinn.com
)
30 points
by
kkm
3 months ago
|
past
|
6 comments
20.
Speculative KV coding: losslessly compressing KV cache by up to ~4×
(
fergusfinn.com
)
155 points
by
kkm
3 months ago
|
past
|
48 comments
21.
70x faster cold(ish) starts for SGLang
(
fergusfinn.com
)
1 point
by
kkm
3 months ago
|
past
22.
Bringing Up DeepSeek-V4-Flash on AMD MI300X
(
fergusfinn.com
)
120 points
by
kkm
3 months ago
|
past
|
25 comments
23.
Brave AI privacy:LLMs on NEAR AI Nvidia-Backed Trusted Execution Environments
(
brave.com
)
1 point
by
kkm
10 months ago
|
past
24.
How fast can an LLM go?
(
fergusfinn.com
)
2 points
by
kkm
10 months ago
|
past
25.
FHE can be leveraged for LLMs such as ChatGPT in a privacy-preserving manner
(
huggingface.co
)
4 points
by
kkm
on Aug 13, 2024
|
past
26.
Harnessing the Power of Large Language Models for Insightful Review Analysis
(
holidaycheck.com
)
1 point
by
kkm
on April 16, 2024
|
past
27.
A Privacy-First approach to use AI for understanding our Customers Better
(
holidaycheck.com
)
1 point
by
kkm
on March 12, 2024
|
past
28.
How to make LLMs go fast
(
vgel.me
)
2 points
by
kkm
on Dec 19, 2023
|
past
29.
Leveraging Large Language Models for Sentiment Classification in Hotel Reviews
(
holidaycheck.com
)
1 point
by
kkm
on Nov 7, 2023
|
past
30.
Managers Should Think More Like Hackers
(
hbr.org
)
3 points
by
kkm
on April 9, 2023
|
past
More
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: