GETTING STARTED

Make yourself
at home.

From a fresh download to your first local request.
A practical guide to NInferEZ Manager.

STEP 01

Install or extract Manager.

Download the Windows installer or Portable ZIP. The installer creates shortcuts for your Windows account. For Portable, extract the complete archive to a folder you can write to and run NInferEZ Manager.exe.

The app and its CPU conversion runtime are included. You do not need to install a separate Python runtime to use the bundled converter.

Choose a download →
STEP 02

Select your GPU engine.

Open Engine in Manager and download the package recommended for your compatible NVIDIA GPU. Manager verifies integrity before activation.

  • sm86: compatible RTX 3000 / Ampere.
  • sm89: compatible RTX 4000 / Ada.
  • sm120a: compatible RTX 5000 / Blackwell, with native NVFP4.

Keep your NVIDIA driver compatible with the engine. Published measurements used RTX 5090; other GPU/model combinations remain Community Preview.

Full compatibility table →
STEP 03

Make a model yours.

In Models, download a recommendation, put a supported .ninfer artifact in the Models folder, or link an existing file without copying it. External links need their original targets to stay available.

Choose the model API name you want to use in clients. Leave managed settings enabled for a coordinated starting profile, and check the estimated memory fit.

Explore our model collection →
STEP 04

Connect an OpenAI-compatible client.

The API starts with Manager. Configure your client with the base URL below and use the exact model API name displayed in the app.

BASE URL
http://127.0.0.1:8173/v1

You can list available model names with the example below. Listing them does not load their weights.

POWERSHELL
Invoke-RestMethod http://127.0.0.1:8173/v1/models

Models load when inference is requested. Localhost is the default. If you choose optional LAN access, configure an API key and review network access in Manager’s guide.

Keep room for your workflow.

Model file size is not total VRAM usage. Context, KV cache, draft components, and runtime workspace need additional memory. Manager’s Fits, Close, and Not recommended labels are estimates.

One model occupies the inference GPU at a time. The default policy unloads it after three idle minutes; you can adjust the timeout. Managed settings are the recommended starting point for ordinary use.

When something needs attention.

A model is missing.

Check the supported artifact format, its Models-folder location, or the external link target. Confirm its API name in Manager.

Loading runs out of memory.

Close other GPU applications, reduce context or optional components, or choose a smaller supported artifact. Review the memory estimate and model card.

A client cannot connect.

Keep Manager running, check the base URL and model API name, and look at the searchable logs for details. Remote access requires explicit LAN setup.

Report an issue on GitHub

Frequently asked questions.

What do I need to get started?

Windows x64, a compatible NVIDIA GPU and driver, and enough VRAM for the model and context you choose. Install Manager, download the matching engine, and add a supported NInfer model. Model weights and GPU engines are separate downloads.

Installer or Portable: which should I choose?

The installer sets up Manager and shortcuts for your Windows account. Portable keeps the application and its data together: extract the entire ZIP into a writable folder, then run NInferEZ Manager.exe. Both offer the same Manager features.

Can I use my existing AI client?

Yes. Connect an OpenAI-compatible client to http://127.0.0.1:8173/v1 and use the exact model API name shown in Manager. The API starts with Manager and loads a model when inference is requested.

Do my prompts leave my computer?

Inference through the default localhost API runs on your machine. Downloading models or engines and checking updates require network access. Your chosen client may have its own network behavior; optional LAN access must be enabled explicitly.

Can I run any GGUF or Hugging Face model?

The engine runs supported NInfer artifacts. The Converter can import supported GGUF encodings or compatible checkpoint geometry with matching configuration and tokenizer resources. A file extension alone does not establish compatibility.

Are all RTX cards tested?

Engine packages target compatible Ampere (sm86), Ada (sm89), and Blackwell (sm120a) GPUs. Published measurements are from RTX 5090. Other combinations remain Community Preview, and native NVFP4 weights require a compatible Blackwell target.

YOUR NEXT STEP

Put your GPU to work.

Your models, your workspace, your next idea.
Start with NInferEZ Manager.

Windows x64 · Installer or Portable