Nagame

AI model hosting and inference infrastructure.

Infrastructure

Purpose-built inference

Nagame operates an OpenAI-compatible inference service on dedicated NVIDIA infrastructure. Requests pass through an HTTPS edge and a health-aware routing layer to vLLM inference backends. Weighted routing and backend failover provide controlled redundancy, while streaming responses and early rate limiting help avoid excessive queueing.

Models available through the Nagame API

Model API model ID Serving Status
Qwen3.6 35B A3B qwen3.6-35b-a3b vLLM ยท NVFP4 Initial trial

Qwen3.6 is the initial trial offering. Additional models are planned through the same API. The service supports streaming chat completions and tool calling. Private encrypted connectivity links infrastructure components. Under saturation, requests may receive an early HTTP 429 response rather than remain in a long queue.

Contact