Skip to main content

AI Inference

Maps to console Artificial Intelligence → Inference. Used to manage LLM inference services: image and model preparation, template specs, deployment runtime, and optional benchmarks. Engine type is determined by the template (e.g. vLLM, SGLang).

  1. Quick Start: import model → create deployment → API Key → chat / curl.
  2. Inference Deployments: deployment concepts, detail tabs, and operations.
  3. Inference Templates, Model Files, Inference Images: specs and resource preparation.
  4. (Optional) Benchmark: load-test an existing deployment.
  5. (Optional) Accelerator Scenarios: inference on Ascend, Hygon, and other accelerators.
Console menuDescription
Inference DeploymentsPrimary entry; the old "Inference Instances" menu is hidden and redirected here
Inference TemplatesReusable specs
Model FilesMountable model data (formerly "inference model library")
Inference ImagesInference runtime images
BenchmarkLoad-test deployments