Speculative decoding
MTP, DFlash, and DFlash2 when the model contains compatible components. Select the mode that fits your artifact and profile.
Standalone NInfer serving for Windows, with GPU-specific packages and model inspection. Let Manager handle setup, or run the engine your way.
Engine 1.0.0 · Windows x64 · Apache 2.0
SERVE · ACCELERATE · INSPECTManager can download the appropriate package. For standalone use, choose the compatible target below.
| GPU family | Target | Native NVFP4 | Qualification | Download |
|---|---|---|---|---|
| RTX 3000 / Ampere | sm86 | No | Community Preview | Download ZIP ↓1274.8 MiB |
| RTX 4000 / Ada | sm89 | No | Community Preview | Download ZIP ↓1266.7 MiB |
| RTX 5000 / Blackwell | sm120a | Yes | RTX 5090 measured; other cards Preview | Download ZIP ↓1062.0 MiB |
A matching target does not mean every GPU/model combination was tested. A compatible NVIDIA driver and enough available VRAM are required. Native NVFP4 weights are Blackwell-specific.
MTP, DFlash, and DFlash2 when the model contains compatible components. Select the mode that fits your artifact and profile.
OpenAI-compatible Chat Completions and Responses, plus supported Messages serving, for compatible clients and workloads.
Host-only structural inspection and machine-readable identity and capabilities make integration more transparent.
Supported KV formats, prefix caching, structured output, and configurable context, subject to runtime and artifact support.
Packages include serving and inspection tools, runtime libraries, manifests, checksums, and license notices. Model weights are separate.
NInferEZ Engine is an independent distribution of ninfer-all, based on NInfer. Original authorship and upstream licenses are retained.
Extract the complete package, inspect your supported artifact, and start a conservative serving profile. Manager offers a more guided starting point.
Runtime behavior depends on architecture, encoded tensors, optional components, GPU target, and memory.
ninfer-serve.exe --version-json
ninfer-serve.exe --capabilities-json
ninfer-inspect.exe --model model.ninfer
--json --estimate-memory
ninfer-serve.exe model.ninfer
--max-context 32768
--kv-capacity 32768
--kv-dtype rk8v4 --max-concurrency 1Your models, your workspace, your next idea.
Start with NInferEZ Manager.