HELP & TROUBLESHOOTING

A clear next step.
When something gets stuck.

Check these common causes before changing advanced settings. Detailed errors and progress are available in Manager’s Logs.

The model is loading but not ready yet

Cold loading may include artifact preparation and GPU-kernel calibration, not just copying a file. Follow progress and Logs. Allow it to finish if progress continues; use Cancel loading if you need to stop. GPU utilization is not a reliable percent-complete indicator.

If it exits, read the full final error. A timeout, memory reservation failure and incompatible kernel are different problems and should not all be treated as “wait longer.”

The model cannot fit in graphics memory

Close other GPU workloads, choose a smaller model or reduce context using managed settings. Check the Fits / Close / Not recommended indicator. Drafts and Vision can add memory beyond the weight file.

Reducing context cannot fix a model whose weights and required runtime buffers already exceed memory. Do not increase host-cache settings at random.

The client cannot connect to the API

Keep Manager running with its API enabled. On the same PC, use http://127.0.0.1:8173/v1 unless you changed the port. A phone’s 127.0.0.1 points at the phone, not your PC.

For another device, enable network access, configure the real key and check the Windows machine’s address, firewall and VPN permissions. Check the client fields.

I get an authentication or model-name error

Use the API key configured in Manager, not a placeholder such as “ninfer.” If local authentication is off but the client demands a key, configure a real key and use it on both sides.

Copy the exact model API name shown in Models. Avoid file extensions or provider prefixes unless they are genuinely part of the configured ID. Confirm the linked file still exists.

A setting is greyed out or rejected

The model or installed engine may not support that feature. Text-only files cannot enable Vision, and MTP/DFlash requires its matching components. Start with managed settings or restore the recommended profile, then make one manual change at a time.

A download or update failed

Check the connection and free disk space, then retry from Manager. Allow room for both download and extraction. A checksum failure means the package should not be activated. Use the official downloads and keep the complete folder structure.

If a model is already on disk, link the existing compatible file instead of downloading another copy. Do not delete your workspace data to fix an app update.

The Converter says the source is unsupported

Check the model architecture and quantization path, and supply the matching configuration, tokenizer and complete shards. The Converter cannot accept every checkpoint just because it has a supported file extension.

Keep the source and review the conversion log. Do not turn on optional GPU validation while another workload needs that GPU. Read supported input paths.

IF YOU NEED TO REPORT IT

Useful details.
No private data.

Include the essentials

  • Manager and engine version.
  • Windows version, GPU series/model, VRAM and driver.
  • Model identity, selected profile and whether managed settings are enabled.
  • Exact steps and the relevant error lines from Logs.

Review before sharing

Remove API keys, personal paths, private prompts and unrelated logs. Do not upload your settings/data folder or model weights as a troubleshooting shortcut.

Request timings can help diagnose performance, but a screenshot of GPU usage alone is not enough to identify a failed load.

GitHub issue reporting requires a GitHub account. Reading these guides and downloading the software does not.

START WITH MANAGER

Your next step.
Your own local AI.

One Windows app. A matching GPU engine. Your choice of model.

Windows x64 · NVIDIA RTX 3000 / 4000 / 5000 series