Skip to main content

Category

Self-hosting quantized LLMs: Qwen3-8B-AWQ, AWQ and vLLM in production

Quantization is what lets a large, capable model earn its keep on one affordable GPU. At 4-bit AWQ the weights compress to roughly a third, which puts a reasoning-capable model within reach of a 24GB card — your data never leaves, and the cost is fixed rather than per token. This cluster uses Qwen3-8B-AWQ as its worked example: how to choose a quantization format, how to get type-safe structured output from vLLM, and how to build on a model that thinks before it answers.

6 articles in total

Foundational guide

Foundational guide (start here)

Qwen
AWQ
量子化
vLLM
生成AI

Qwen3-8B-AWQ practical guide: self-hosting a 'reasoning LLM' on a single GPU with 4-bit quantization

Explaining Qwen3-8B-AWQ faithful to the official documentation. With AWQ 4-bit quantization, compress the weights to about 6GB and run in production on a single 24GB GPU. Switching hybrid thinking (thinking/non-thinking), OpenAI-compatible serving with vLLM, the recommended sampling per mode, 131K extension with YaRN, tool calling, and quantization-specific pitfalls (presence_penalty / greedy forbidden), all in real code.

15 min read

Related practical articles