RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
80 tokens/second. That's what you get pairing RTX 5080 + RTX 3090 on Qwen 3.6 27B—real-world inference benchmarks leaders need to see.

Why it matters
Real-world GPU inference performance data is critical for infrastructure planning. This hands-on benchmark demonstrates practical throughput achievable with consumer/prosumer hardware, informing cost-per-token calculations and deployment decisions for mid-size models.
The key facts
10 to knowRTX 5080 + RTX 3090 GPU pairing
80 tokens/second throughput achieved
Qwen 3.6 27B model
Q8 quantization
Inference benchmark (real hardware)
Published June 2026
80 tokens/second throughput
RTX 5080 + RTX 3090 dual-GPU setup
Published June 13, 2026
Low engagement (13 HN points, 0 comments) suggests niche but technical audience
Go to the source
Hacker Newsimil.net
Publisher excerpt: Article URL: Comments URL: Points: 13 # Comments: 0