# ninfer-cmp170hx **Repository Path**: python395118/ninfer-cmp170hx ## Basic Information - **Project Name**: ninfer-cmp170hx - **Description**: No description available - **Primary Language**: Unknown - **License**: Apache-2.0 - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-10-08 - **Last Updated**: 2026-10-08 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # NInfer-CMP170HX NInfer-CMP170HX is the `sm_80` port of the NInfer single-GPU inference engine for an **unlocked NVIDIA CMP 170HX**. The qualified host has compute capability 8.0, 70 SMs, and 64 GiB of visible HBM. This project is not a generic support claim for stock 8 or 10 GiB CMP cards. The qualification environment used NVIDIA Linux driver **610.43.03**. The CUDA 13.1 container requires Linux driver **580 or newer**. Forward-compatibility `libcuda` is intentionally removed because it is not appropriate for CMP/GeForce- class hosts. See the [CMP 170HX Linux guide](docs/cmp-170hx-linux.md) for the complete qualified scope, setup details, and limitations. ## Qualification evidence | Gate | Result | |---|---| | Fresh CUDA 13.1 SM80 build | **245** build/link steps completed | | Compose health check | Healthy, using an explicit `/usr/bin/bash` healthcheck | | HTTP routes | `/health` and `/v1/models` passed | | Real model completion | Qwen3.8-27B OpenAI-compatible completion returned `CMP170HX_OK` with HTTP 200 | | Home Assistant route | Validated through Home Assistant after its Local OpenAI LLM integration was switched from the llama.cpp-specific configuration to its Generic OpenAI-compatible backend | The Home Assistant backend selection is an integration configuration requirement, not an NInfer code change. ## Observed CMP performance These are qualification and operator observations, not a normalized benchmark suite. ### Controlled qualification A Qwen3.8-27B Compose smoke test produced **51.7 generated tok/s** during the exact runtime validation. It used 25 prompt tokens and 7 generated tokens, so this is a smoke result rather than a full benchmark. ### Field observations Qwen3.6-35B-A3B runs from the NInfer/llama-swap logs showed **20.82 GiB** of weights, **4.10 - 6 GiB** of KV, depending on KV Quantization while samples were around **1,580–2,347 prompt tok/s** and **93–126 generated tok/s**. Recent Home Assistant calls through the Generic OpenAI-compatible path recorded these dashboard observations. The `Cached` field showed no cached-token value (`—`) for each row: | Request | Prompt tokens | Generated tokens | Prompt speed | Generated speed | Duration | |---:|---:|---:|---:|---:|---:| | 48 | 14,148 | 72 | **4,299.62 tok/s** | **218.20 tok/s** | 3.63 s | | 46 | 13,876 | 65 | **4,278.49 tok/s** | **217.55 tok/s** | 3.55 s | | 44 | 13,720 | 122 | **4,263.00 tok/s** | **217.16 tok/s** | 3.79 s | | 42 | 12,815 | 84 | **4,292.47 tok/s** | **206.95 tok/s** | 3.40 s | These are field observations, not a reproducible benchmark. All figures depend on the model artifact, context, MTP/graphs/KV settings, prompt, output length, and concurrency. ## Validation screenshots ### Home Assistant / llama-swap activity ![llama-swap dashboard showing recent successful Jarvis API requests and throughput](docs/images/llama-swap-jarvis-metrics.png) ### Recent NInfer request logs ![llama-swap NInfer logs showing model load, OpenAI-compatible requests, and throughput](docs/images/cmp170hx-ninfer-llama-swap-logs.png) ### GPU residency: Qwen3.6-35B-A3B ![NVTOP showing Qwen3.6-35B-A3B NInfer alongside the Jarvis runtime on the 64 GiB CMP 170HX](docs/images/cmp170hx-qwen36-35b-ninfer-nvtop.png) ### GPU residency: Qwen3.8-27B ![NVTOP showing the Qwen3.8-27B NInfer runtime on the CMP 170HX](docs/images/cmp170hx-qwen38-27b-ninfer-nvtop.png) ## RTX 3090 comparison context The direct parent is [`Don-Chad/ninfer-3090`](https://github.com/Don-Chad/ninfer-3090). Its prior README published the RTX 3090 figures below. They are useful context, not a head-to-head result: | Workload or resource | CMP 170HX qualification context | RTX 3090 figures published by the parent | |---|---|---| | Visible memory | 64 GiB HBM on the qualifying CMP host | 24 GiB VRAM baseline | | Qwen3.8-27B | 51.7 generated tok/s controlled Compose smoke | C1: 861.51 tok/s prefill and 71.00 tok/s decode; C8: 165.33 tok/s aggregate decode | | Qwen3.6-35B-A3B | Field samples: 93–126 generated tok/s | C1: 162.7 aggregate tok/s; C2: 267.9; C6: 383.4; longer 512-token C2: 399.1 aggregate tok/s | These are **different controlled runs, artifacts, workload shapes, and configurations**. Do not read the table as a normalized benchmark or as a universal speedup claim. The CMP's practical advantage is additional memory headroom plus validated Linux and OpenAI-compatible serving—not a declaration of victory over the RTX 3090. ## Quick start Requirements: x86-64 Linux, NVIDIA Linux driver 580 or newer, Docker, NVIDIA Container Toolkit, and a supported `.ninfer` model artifact. ```bash docker build \ --build-arg NINFER_CUDA_ARCHITECTURES=80 \ --tag ninfer-cmp170hx:sm80 . export NINFER_MODEL_DIR=/path/to/models docker compose up ``` The Compose profile publishes the API on `http://127.0.0.1:8080/v1` and mounts the model directory read-only. For the full standalone setup and qualified operating details, use the [CMP 170HX Linux guide](docs/cmp-170hx-linux.md). ### Home Assistant For this NInfer API path, configure Home Assistant's Local OpenAI LLM integration with its **Generic OpenAI-compatible backend**. Do not use the llama.cpp-specific backend configuration. ## Scope and limitations - One model per GPU process, with bounded concurrency fixed at startup. - No multi-GPU execution or CPU offload. - Blackwell-only NVFP4/W4A4 execution is unavailable on SM80. - Tool calls are returned to the client but are not executed by NInfer. - This README does not claim support for generic stock CMP cards. ## Lineage and credits - [`Neroued/ninfer`](https://github.com/Neroued/ninfer) is the original NInfer engine. - [`Don-Chad/ninfer-3090`](https://github.com/Don-Chad/ninfer-3090) is the direct parent and provided the SM86 compatibility, Linux/Docker, and Qwen3.8 foundations used by this port. ## License Apache License 2.0. See [LICENSE](LICENSE).