We got 207 tok/s with Qwen3.5-27B on an RTX 3090
207 tok/s on a $300 GPU. Here's what optimization just unlocked for inference on consumer hardware.

Why it matters
Real-world inference optimization benchmarks matter to engineers deploying models at scale. A 27B parameter model hitting 207 tokens/second on consumer-grade GPUs (RTX 3090) signals meaningful progress in making frontier-class model inference accessible without enterprise infrastructure.
The key facts
12 to knowQwen3.5-27B achieved 207 tokens/second throughput
RTX 3090 GPU (consumer-grade, ~$300-400 MSRP)
GitHub repo: Luce-Org/lucebox-hub
Posted April 20, 2026 on Hacker News (137 points, 37 comments)
Inference optimization technique demonstrated on public hardware
207 tokens/second throughput achieved
Hardware: RTX 3090 (consumer-grade GPU)
Model: Qwen3.5-27B
137 points on Hacker News (strong community interest)
37 comments indicating active technical discussion
Published Apr 20, 2026
Source: GitHub repository (Luce-Org/lucebox-hub)
Go to the source
Hacker Newsgithub.com
Publisher excerpt: Article URL: Comments URL: Points: 137 # Comments: 37