Skip to main content

Product Introduction

Cloudpods AI Cloud is a unified management platform for large language model (LLM) inference and AI container applications, helping enterprises deploy, schedule, and operate AI workloads on a single platform, seamlessly integrated with the Cloudpods private cloud / multi-cloud resource ecosystem.

Core Capabilities​

  • AI Inference Services: Deploy and manage LLM inference deployments with GPU scheduling, model mounting, and inference service address allocation.
  • AI Application Management: One-stop deployment of LLM application orchestration, agent assistants, image generation, and other AI container applications.
  • Model Library: Unified management of model sources, versions, and caches, supporting multi-instance reuse, offline distribution, and avoiding redundant downloads.
  • Templates and Images: Define resource specifications such as CPU/memory/GPU through templates, and manage container runtime environments through images for standardized delivery.
  • GPU Operations: Automatic GPU device detection and registration, with unified configuration and management of NVIDIA/CUDA environments.

Console Features​

After entering the Artificial Intelligence section of the console, there are three main modules:

  • Applications: Manage AI application instances and application templates.
  • Inference: Manage inference deployments, inference templates, model files, inference images, and benchmarks.
  • Images: Manage container images used by AI applications and inference services.

Instance = Image + Template Spec + (optional) Model.

Supported Applications​

AI Inference​

  • vLLM: A high-throughput, low-latency inference service that requires GPU, suitable for OpenAI-compatible APIs and higher concurrency scenarios.
  • SGLang: A high-performance inference engine that requires GPU, suitable for online inference scenarios similar to vLLM.

AI Applications​

  • Dify: An LLM application development and orchestration platform that supports building conversational, RAG, Agent, and workflow applications. Does not require GPU.
  • OpenClaw: An open-source self-hosted personal agent assistant. Does not require GPU.
  • HermesAgent: An open-source self-hosted personal agent assistant. Does not require GPU.
  • ComfyUI: An image generation and visual workflow application with node-based workflow orchestration. Requires GPU.

Getting Started​