AI Inference
Maps to console Artificial Intelligence → Inference. Used to manage LLM inference services: image and model preparation, template specs, deployment runtime, and optional benchmarks. Engine type is determined by the template (e.g. vLLM, SGLang).
Recommended Reading Order
- Quick Start: import model → create deployment → API Key → chat / curl.
- Inference Deployments: deployment concepts, detail tabs, and operations.
- Inference Templates, Model Files, Inference Images: specs and resource preparation.
- (Optional) Benchmark: load-test an existing deployment.
- (Optional) Accelerator Scenarios: inference on Ascend, Hygon, and other accelerators.
Menu Mapping
| Console menu | Description |
|---|---|
| Inference Deployments | Primary entry; the old "Inference Instances" menu is hidden and redirected here |
| Inference Templates | Reusable specs |
| Model Files | Mountable model data (formerly "inference model library") |
| Inference Images | Inference runtime images |
| Benchmark | Load-test deployments |
快速开始
从导入模型到 chat 验证的端到端路径:在推理模板侧 导入模型 生成模板 → 基于模板创建部署 → 用 API Key 做 chat测试(或 curl)。
推理部署
概念介绍
Inference Templates
Overview
Model Files
Overview
推理镜像
概念介绍