NINFEREZ ENGINE

The engine behind
your local AI.

Standalone NInfer serving for Windows, with GPU-specific packages and model inspection. Let Manager handle setup, or run the engine your way.

Engine 1.0.0 · Windows x64 · Apache 2.0

NInferEZ Engine logoSERVE · ACCELERATE · INSPECT
EXPLORE THE PROJECTEngine on GitHub ↗Models on Hugging Face ↗
MATCH YOUR HARDWARE

The right target for your GPU.

Manager can download the appropriate package. For standalone use, choose the compatible target below.

GPU engine package compatibility
GPU familyTargetNative NVFP4QualificationDownload
RTX 3000 / Amperesm86NoCommunity PreviewDownload ZIP ↓1274.8 MiB
RTX 4000 / Adasm89NoCommunity PreviewDownload ZIP ↓1266.7 MiB
RTX 5000 / Blackwellsm120aYesRTX 5090 measured; other cards PreviewDownload ZIP ↓1062.0 MiB

A matching target does not mean every GPU/model combination was tested. A compatible NVIDIA driver and enough available VRAM are required. Native NVFP4 weights are Blackwell-specific.

UNDER THE HOOD

Built for the NInfer workflow.

Speculative decoding

MTP, DFlash, and DFlash2 when the model contains compatible components. Select the mode that fits your artifact and profile.

Compatible API serving

OpenAI-compatible Chat Completions and Responses, plus supported Messages serving, for compatible clients and workloads.

Artifact inspection

Host-only structural inspection and machine-readable identity and capabilities make integration more transparent.

Runtime controls

Supported KV formats, prefix caching, structured output, and configurable context, subject to runtime and artifact support.

GPU-specific distribution

Packages include serving and inspection tools, runtime libraries, manifests, checksums, and license notices. Model weights are separate.

Independent & attributed

NInferEZ Engine is an independent distribution of ninfer-all, based on NInfer. Original authorship and upstream licenses are retained.

FOR HANDS-ON WORKFLOWS

From the terminal.
Or through Manager.

Extract the complete package, inspect your supported artifact, and start a conservative serving profile. Manager offers a more guided starting point.

Runtime behavior depends on architecture, encoded tensors, optional components, GPU target, and memory.

POWERSHELL · EXTRACTED ENGINE FOLDER
ninfer-serve.exe --version-json
ninfer-serve.exe --capabilities-json

ninfer-inspect.exe --model model.ninfer 
  --json --estimate-memory

ninfer-serve.exe model.ninfer 
  --max-context 32768 
  --kv-capacity 32768 
  --kv-dtype rk8v4 --max-concurrency 1
Example text profile; choose settings for your model.
YOUR NEXT STEP

Put your GPU to work.

Your models, your workspace, your next idea.
Start with NInferEZ Manager.

Windows x64 · Installer or Portable