NINFEREZ MANAGER

Local AI.
All together.

A convenient Windows workspace built around NInfer. Bring model libraries, GPU engines, inference profiles, and API access into one place.

Manager workspace · Windows x64 · MIT licensed
NInferEZ Manager
YOUR WORKSPACE

Everything in one place.

Ready when you are.

Local first
Model libraryLocal files & linked models
GPU engineA package for your GPU
Your models NInfer v3
S
Swift 1.527B · NVFP4 · Text + Vision
Vision
O
OrcaRouter27B · IQ3_XXS · MTP + DFlash2
Text
127.0.0.1:8173/v1YOUR LOCAL API
Explore the workspace Interactive product illustration
BUILT AROUND YOUR WORKFLOW

One workspace.
The control you need.

Your library, your way

Download recommendations, keep local files, or link existing artifacts without making another copy. Give each model an editable API alias.

Profiles with a purpose

Managed settings coordinate dependent controls. Advanced manual settings stay available, with clear compatibility errors.

Engine management

GPU-specific packages, SHA-256 verification, activation, and update checks. Models and engines are managed independently.

An API that is ready

A local API starts with Manager. Load models on request, cancel loading, or adjust idle unloading to reclaim GPU memory.

Integrated conversion

The guided and manual workflows use the same backend as the standalone Converter, with progress and cancellation.

See what is happening

Request metrics, searchable logs, tray operation, and single-instance activation keep the application practical for daily use.

A LITTLE MORE HEADROOM

Know what fits.
Before you load.

Manager estimates model memory with a practical context allowance, so you can choose an appropriate starting point for your available VRAM.

One model occupies the inference GPU at a time. By default, an idle model unloads after three minutes; adjust that to suit your workflow.

Fits

Estimated room for model and context.

Close

Review your context and GPU headroom.

Not recommended

Choose a smaller artifact or reduce memory use.

Estimates are guidance. Actual usage depends on artifact, engine, context, and runtime state.
CONNECT YOUR FAVORITE CLIENT

Your local model.
Your familiar tools.

Use the exact API name shown in Manager to connect an OpenAI-compatible client. Listing models does not load their weights.

Connection guide →
CLIENT CONFIGURATION
// Base URL
http://127.0.0.1:8173/v1

// Model
<your-model-api-name>
Localhost by default. LAN access is opt-in.

A complete app, with separate model downloads.

The current Manager workspace includes its CPU conversion runtime. Download a GPU engine and model weights separately. Windows releases are currently unsigned.

YOUR NEXT STEP

Put your GPU to work.

Your models, your workspace, your next idea.
Start with NInferEZ Manager.

Windows x64 · Installer or Portable