weicj/vLLM-2080Ti-Definitive
A vLLM fork tailored for dual RTX 2080 Ti GPUs with NVLink, enabling local LLM inference on Turing hardware.

This repository is a hardware-focused fork of vLLM that preserves SM75-specific source changes, launcher profiles, and validation evidence needed to run LLM inference on dual RTX 2080 Ti 22GB cards connected via NVLink, as well as other Turing GPUs. It targets serving 27B and 35B-class models such as Qwen 27B with FP8 weight support and reports single-request decode throughput above 200 tokens per second. The project provides a reproducible runtime stack based on upstream vLLM, with documentation and benchmarks comparing performance against higher-end consumer GPUs.
Frequently asked
- What is weicj/vLLM-2080Ti-Definitive?
- A vLLM fork tailored for dual RTX 2080 Ti GPUs with NVLink, enabling local LLM inference on Turing hardware.
- Is vLLM-2080Ti-Definitive open source?
- Yes — weicj/vLLM-2080Ti-Definitive is open source, released under the Apache-2.0 license.
- What language is vLLM-2080Ti-Definitive written in?
- weicj/vLLM-2080Ti-Definitive is primarily written in Python.
- How popular is vLLM-2080Ti-Definitive?
- weicj/vLLM-2080Ti-Definitive has 1k stars on GitHub.
- Where can I find vLLM-2080Ti-Definitive?
- weicj/vLLM-2080Ti-Definitive is on GitHub at https://github.com/weicj/vLLM-2080Ti-Definitive.