A home for your models.
Download a recommendation, import a supported artifact, or link a model you already have. Give each one an API name and a profile that makes sense.
NInfer, made simple on Windows. NInferEZ Manager brings your models, engines, and local AI API together in one convenient app.
Get to know Manager ↗Ready when you are.
127.0.0.1:8173/v1YOUR LOCAL APIDownload, import, or link an existing file.
The same backend as the standalone Converter.
Choose a package for your NVIDIA GPU.
From choosing a model to connecting your favorite client, Manager keeps the moving parts within reach.
Download a recommendation, import a supported artifact, or link a model you already have. Give each one an API name and a profile that makes sense.
Managed profiles adapt dependent controls to your model and engine. Go manual whenever you want more control.
Connect an OpenAI-compatible app to Manager. Models load on demand and can unload after an adjustable idle period.
127.0.0.1:8173/v1Download the GPU package that matches your hardware. Manager checks integrity before activating it.
Explore engine compatibilityPrepare supported checkpoints in Manager’s guided converter, or open up manual controls for advanced plans.
Meet the ConverterA complete desktop experience, with independent tools for the way you like to work.

The desktop workspace for models, inference profiles, GPU engines, API access, and conversion.
Explore Manager PREPARETurn supported checkpoints into NInfer artifacts with a guided wizard or detailed manual controls.
Explore Converter RUNStandalone serving, GPU-specific packages, and model inspection. The engine behind the experience.
Explore EngineInstall it on Windows or extract the Portable ZIP into a writable folder.
Choose a download →Select the package for your GPU, then download or link a supported NInfer artifact.
Find your model →Point your AI client at the local API. Manager loads the selected model when you need it.
Connect a client →Explore native NInfer conversions, with artifact details, component support, and original model cards.
A complete text and vision artifact, with native NVFP4 weights and a bundled speculative draft for compatible Blackwell GPUs.
Adds a DFlash2 draft to OrcaRouter, with MTP available as an alternative. Explore the published long-context measurements.
Our smallest featured 27B artifact. A text-only IQ2_XS conversion for a more compact model library.
Weights are downloaded separately. Each model keeps its own license.
More guidance in our getting started guide.
Open the guideWindows x64, a compatible NVIDIA GPU and driver, and enough VRAM for the model and context you choose. Install Manager, download the matching engine, and add a supported NInfer model. Model weights and GPU engines are separate downloads.
The installer sets up Manager and shortcuts for your Windows account. Portable keeps the application and its data together: extract the entire ZIP into a writable folder, then run NInferEZ Manager.exe. Both offer the same Manager features.
Yes. Connect an OpenAI-compatible client to http://127.0.0.1:8173/v1 and use the exact model API name shown in Manager. The API starts with Manager and loads a model when inference is requested.
Inference through the default localhost API runs on your machine. Downloading models or engines and checking updates require network access. Your chosen client may have its own network behavior; optional LAN access must be enabled explicitly.
Your models, your workspace, your next idea.
Start with NInferEZ Manager.