Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
YetAnotherNick
4 months ago
|
parent
|
context
|
favorite
| on:
Surpassing vLLM with a Generated Inference Stack
For compute bound region(high batch size) yes, but for low batch size it could improve the throughput.
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: