x64 desktop. Use a supported, updated Windows installation.
Windows.
NVIDIA RTX. Your choice.
NInferEZ supports NVIDIA GeForce RTX 3000, 4000 and 5000 series on Windows x64. Choose a model that fits your card’s memory.
A growing model collection, with memory-fit guidance in Manager.
Install a current driver compatible with your card and the downloaded engine.
Support across three RTX series.
Manager recommends the matching runtime. For manual setup, use the corresponding series below.
| NVIDIA GeForce series | Model starting point | Runtime package |
|---|---|---|
| RTX 3000 series | Supported groupwise integer models, such as OrcaRouter or Heretic. | Ampere · sm86 |
| RTX 4000 series | Supported groupwise integer models, such as OrcaRouter or Heretic. | Ada · sm89 |
| RTX 5000 series | Supported integer models, or Swift NVFP4 when graphics memory allows. | Blackwell · sm120a |
Series support does not mean every model fits every card. For example, a larger card can accommodate models that exceed the memory of a smaller card in the same series.
Enough room for the whole job.
The model file is only one part of what the GPU needs.
Model weights
These are the parameters loaded for inference. Download size is a useful guide, not an exact VRAM total.
Conversation & work buffers
Longer inputs and saved conversation context need more memory, alongside the engine’s working space.
Images & acceleration
Vision inputs and additional draft components can increase the requirement beyond a text-only profile.
A comfortable starting point for the model and a useful conversation.
Limited headroom. Use a smaller context or a smaller model.
The model and a reduced context are estimated not to fit.
Manager shows these recommendations for your GPU before downloading or selecting a model. Other GPU apps reduce free memory. Close them when you need the space; do not assume a bigger context will fit just because the model file is small.
There is no single memory minimum for every model. Choose a compatible artifact and profile for your GPU. New catalogue entries can expand the choice without changing the hardware-series support.
Platform and format boundaries.
Windows version and installation
The desktop runtime targets Windows x64, with a minimum platform baseline of Windows 10 build 19041. Prefer a currently supported Windows version with updates installed. Manager is self-contained; the app runtime is included. Keep all Portable files together in a writable folder.
What about NVFP4 on RTX 3000 or 4000?
Native NVFP4 acceleration belongs to the Blackwell target used for RTX 5000-series packages. For RTX 3000/4000, start with a supported integer model instead. Do not assume the Swift NVFP4 artifact has the same execution path on older series.
Does any NVIDIA or other GPU work?
The published GPU packages target GeForce RTX 3000, 4000 and 5000 series. They are not AMD, Intel or macOS inference packages. Professional NVIDIA cards have their own architecture identities; do not select a GeForce package just because a product name contains a similar number.
Will any NInfer v3 model work?
No. The container version is not a universal architecture promise. The runtime must support the model geometry, encodings and components. Inspect the artifact and its capabilities, then choose a profile that matches.
Read supported format details →What else should I leave room for?
Allow disk space for the app, engine package, model files and downloads during extraction. Converting models also needs source and output storage, plus path-dependent RAM. No fixed system-RAM figure would accurately describe every workflow.
Your next step.
Your own local AI.
One Windows app. A matching GPU engine. Your choice of model.
