Local API and model names.
Manager exposes an OpenAI-compatible API at http://127.0.0.1:8173/v1 by default. Opening the app starts the API; the first inference request loads a model. Listing models does not require loading weights into the GPU.
| Operation | Endpoint | Behavior |
|---|---|---|
| Discover models | GET /v1/models | Available local/linked model artifacts and their configured API names. |
| Chat generation | POST /v1/chat/completions | Send messages and the exact model ID. Streaming responses are available when requested. |
Use the model’s configured API name, not its long file name or an invented provider prefix. You can rename the API alias in Models. The application can advertise several available models, but keeps one model resident and runs one generation at a time. Waiting requests are queued; a model switch may incur another load.
curl.exe http://127.0.0.1:8173/v1/modelsIf authentication is enabled, send Authorization: Bearer <your-configured-key>. Do not put a real secret in a screenshot, issue report or shared command history.
Generation controls supported by the engine and template include output-token limits and sampling choices. The request still must fit the model’s configured context. A large output limit does not create extra context.
Client connection example →Localhost, LAN and VPN.
Local-only access is the default. Network access is opt-in in Settings, and Manager requires an API key for it. Binding to 0.0.0.0 means listening on reachable interfaces; it is not the address to paste into a client.
On another device, use the Windows machine’s LAN or VPN address with port 8173 and the /v1 suffix. VPN routing, access rules and Windows Firewall must permit the connection. Enabling network access does not automatically grant every device access or create a secure public internet service.
Use a strong key and a trusted network. Plain HTTP does not encrypt the key; use an appropriately secured private connection or a separately managed TLS layer when needed. Never port-forward an unauthenticated endpoint to the internet.
For browser-based clients, CORS is a separate setting. Allow only origins needed by your client; CORS is not authentication and does not replace firewall or key checks.
Manager’s internal engine connection is distinct from its user-facing API. Configure the public API key in Manager rather than editing engine process secrets.
Managed settings and advanced controls.
Managed mode derives compatible launch settings from the selected artifact, available runtime and your choices. Unsupported acceleration options are disabled instead of being blindly passed to the engine. It is the recommended starting point, not a guarantee against every resource shortage.
Manual mode restores the saved advanced controls and gives you responsibility for their combination. Keep a known-working preset before changing numerical or memory settings.
| Category | What it changes | Practical guidance |
|---|---|---|
| Context and prefill | How much conversation fits, and how input is processed in chunks. | A larger context consumes more cache. Chunk size can affect working memory and prefill speed. |
| KV precision and capacity | Conversation-cache representation, allocation policy and reserves. | Weight quantization is separate. Match capacity to context unless you understand the policy. |
| Speculative acceleration | MTP or DFlash/DFlash2 proposals and draft settings. | Requires matching artifact components and a compatible engine; faster output depends on the workload. |
| State precision and thinking | Numerical state representation and template-dependent thinking behavior. | Can change output quality or token count. It is not merely a cosmetic performance switch. |
| Vision and media | Image input, projector availability and media resource budgets. | Requires Vision components. Text-only artifacts cannot gain Vision by checking a box. |
| API and idle policy | Port, key, network reachability and when weights unload. | Default local port is 8173; the default idle unload is three minutes. |
LM-head drafting and n-gram proposals are acceleration details, not independent substitutes for a missing MTP or draft component. Host cache/state settings can consume system RAM; they are not a promise of zero-VRAM inference or a free way to make any model fit.
Ordinary settings save automatically. Model profiles and explicit connection changes retain their own Save/Apply actions. The runtime’s reported capabilities are authoritative; do not assume every upstream feature is present in every installed package.
Context capacity is not filled context.
A 256K or 512K configured limit is the size of the permitted token envelope, not evidence that a request contained that many tokens. An actual long prompt has to be processed before decoding begins. This can take substantially longer than loading weights or answering a short prompt.
Reserve room for output: prompt tokens, retained messages and the generated answer share the context limit. The practical maximum also depends on attention geometry, KV precision, drafts, Vision and runtime work buffers.
Prefill processes the input. Time to first token includes the waiting work before output begins. Decode tokens/s measures generated tokens after generation is underway. Comparing these as if they were the same speed is misleading.
Prefix reuse can shorten later requests when the engine retains a compatible prefix; it does not eliminate all long-input processing. Lower KV precision reduces cache size but may have quality or hardware trade-offs. Keep the recommended profile unless you validate changes for your use.
Plain-language memory recommendations →Artifacts and supported architecture.
NInfer v3 describes the artifact container and component layout. The runtime also needs to support the model’s architecture, tensor shapes and encoded formats. “v3” alone does not mean any architecture can run.
The shipped engine line includes supported Qwen3.5-family hybrid dense and MoE geometries and compatible derivatives. Familiar model names or extensions do not replace structural inspection. The presence of text, Vision, MTP and DFlash components is checked separately.
For supported GGUF imports, the Converter preserves encoded weight blocks rather than requantizing them. A Safetensors checkpoint needs the correct config, tokenizer, tensor files and supported conversion path. Unsupported quantization or geometry cannot be fixed by renaming the file.
Conversion and repacking do not reconstruct precision lost in the original weights. Missing drafts or a Vision projector are not synthesized. Keep original source and output licenses; the model’s ownership and terms do not change.
Converter workflows and inputs →Updates, presets and reproducibility.
The app, engine packages and model catalogue update independently. Refreshing the catalogue can show newly published models without a new Manager release. Downloading a new package is not the same as silently replacing a running engine.
Manager offers app updates and a Check updates control. Apply an update when you choose. Keep the workspace data and use the supported update flow; do not delete your models or settings just because the executable changes.
Presets pair a model identity with settings. Portable and installed data locations differ. An external model link depends on the original file still being present. Import only trusted presets; unsupported settings or version requirements should produce an explanatory error rather than launch arbitrary commands.
For reproducible measurements, record the engine version, artifact hash, GPU/driver, actual input tokens, sampling, draft mode, KV type, context limit, prefill chunk and prefix-reuse policy. Preserve request-level timings, not just a best result.
Inspect engine and artifact metadata →Published measurements, in context.
These are historical measurements of our converted artifacts on an RTX 5090. They are not promises for every GPU, not newly rerun tests of the current release, and not interchangeable workloads.
| Artifact and workload | Requests / context | Mean decode | Qualification |
|---|---|---|---|
| Swift NVFP4 · ordinary short requests | 6 requests · 128K configured capacity | 197.3 tokens/s | Short input, not a filled 128K prompt. |
| Swift NVFP4 · structured-output fixtures | 15 requests · 128K capacity · DFlash2 K7 | 361.2 ± 71.2 tokens/s | Text-only profile. Quality of generated structured output was not checked; no aggregate-throughput claim. |
| OrcaRouter DFlash2 · standard SpecBench prompts | 15 turns · 256K capacity · prefix reuse off | 258.50 tokens/s | Configured capacity, not a full 256K input. |
| OrcaRouter DFlash2 · near-full extended context | 15 turns · 510,035–511,099 input tokens · prefix reuse on | 148.03 tokens/s | Extended-context experiment; not a comprehensive long-context quality validation. |
Different prompts, seeds, profiles, prefix-reuse settings and output lengths change the result. Do not rank models by the fastest isolated request. The extended-context test also has substantial prefill time; its decode rate is not an instant-response claim.
Swift quality measurements
Published converted-artifact results include AIME 2026 29/30 (96.67%), GPQA-Diamond 90.4%, IFBench 78.33%, ERQA 63.25%, C-Eval 90.56% and HMMT 93.33%. These are benchmark-specific results, not a universal quality guarantee. Reproduction settings, sample counts and comparison caveats are in the model card.
Swift model card and benchmark details ↗OrcaRouter variants and acceleration
In a five-request short-input comparison, MTP3 plus n-gram 15 averaged 203.41 tokens/s and DFlash2 K4 plus proposal head averaged 203.61 tokens/s — effectively a practical tie for that workload. Full-prefill retrieval tests and SpecBench runs are different experiments, with different rates and timing.
OrcaRouter measurements and exact artifacts ↗Heretic validation
Structural conversion checks were performed. No complete quality suite or directly comparable throughput benchmark is published for this artifact. Source-model results are not presented as measurements of the converted file.
Heretic artifact and validation notes ↗Source code and component terms.
Technical use does not require reading code, but all three software projects are public. Original runtime, conversion and model work remain attributed.