ChipsThe story, in brief

We got 207 tok/s with Qwen3.5-27B on an RTX 3090

207 tok/s on a $300 GPU. Here's what optimization just unlocked for inference on consumer hardware.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Real-world inference optimization benchmarks matter to engineers deploying models at scale. A 27B parameter model hitting 207 tokens/second on consumer-grade GPUs (RTX 3090) signals meaningful progress in making frontier-class model inference accessible without enterprise infrastructure.

The key facts

12 to know
  1. Qwen3.5-27B achieved 207 tokens/second throughput

  2. RTX 3090 GPU (consumer-grade, ~$300-400 MSRP)

  3. GitHub repo: Luce-Org/lucebox-hub

  4. Posted April 20, 2026 on Hacker News (137 points, 37 comments)

  5. Inference optimization technique demonstrated on public hardware

  6. 207 tokens/second throughput achieved

  7. Hardware: RTX 3090 (consumer-grade GPU)

  8. Model: Qwen3.5-27B

  9. 137 points on Hacker News (strong community interest)

  10. 37 comments indicating active technical discussion

  11. Published Apr 20, 2026

  12. Source: GitHub repository (Luce-Org/lucebox-hub)

Go to the source

Hacker Newsgithub.com

Publisher excerpt: Article URL: Comments URL: Points: 137 # Comments: 37
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips