Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -51,3 +51,46 @@ Each OpenNebula deployment is different but they all have to provide a basic set
| SC20 | Create and assign quota to group(s) | In an existing group, create a quota and showcase how it functions. | ▶️ [Enforcing Resource Quotas in OpenNebula](https://www.youtube.com/watch?v=wAd9YpFJnq4) | [Quotas]({{% relref "product/cloud_system_administration/capacity_planning/quotas/" %}}) |

## AI Factory Success Criteria (Optional)

### GPU Instances

| **ID** | **Task** | **Validation description** | **Video tutorials** | **Documentation** |
| ------ | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------- |
| AI-SC1 | Configure GPU Passthrough | Configure a host so that a physical GPU can be assigned directly to a virtual machine. Validate that the GPU is exposed by the hypervisor and available for PCI passthrough. | | [NVIDIA GPU Passthrough]({{% relref "product/cluster_configuration/pci_passthrough_sriov/nvidia_gpu_passthrough/" %}}) |
| AI-SC2 | Deploy a VM with GPU resources | Create or update a VM template with a GPU device, instantiate the VM successfully, and verify that the guest operating system detects and can use the assigned GPU. | | [NVIDIA GPU Passthrough]({{% relref "product/cluster_configuration/pci_passthrough_sriov/nvidia_gpu_passthrough/" %}}) |
| AI-SC3 | Configure MIG Partitioning | Partition a MIG-capable GPU and assign separate MIG instances to independent workloads. Validate that each workload sees only its allocated MIG resources. | | [NVIDIA vGPU and MIG-backed vGPU]({{% relref "product/cluster_configuration/pci_passthrough_sriov/nvidia_mig_passthrough/" %}}) |
| AI-SC4 | Configure GPU Quotas | Define GPU resource limits for a group or VDC and verify that users cannot allocate GPU resources beyond the configured quota. | | [Usage Quotas]({{% relref "product/cloud_system_administration/capacity_planning/quotas/" %}})
| AI-SC5 | Validate GPU Tenant Isolation | Run GPU workloads for separate tenants or VMs and validate that assigned GPU resources are isolated and cannot be accessed by other tenants. | | [Multi-tenancy Models and User Roles]({{% relref "getting_started/understand_opennebula/opennebula_concepts/cloud_access_model_and_roles/" %}}) |
| AI-SC6 | Monitor GPU Resources | Validate visibility of GPU utilization, memory consumption, health and other relevant GPU metrics through the monitoring stack. | | [Prometheus and Grafana]({{% relref "product/cloud_system_administration/prometheus/overview/" %}}) |

### Managed Kubernetes

| **ID** | **Task** | **Validation description** | **Video tutorials** | **Documentation** |
| ------ | ------------------------ | --------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- |
| AI-SC7 | Deploy a GPU-enabled Kubernetes Cluster | Provision a Kubernetes cluster with GPU-capable worker nodes and verify that the cluster is operational and the GPU nodes are registered correctly. | | [Deployment of AI-ready Kubernetes]({{% relref "solutions/ai_factory_blueprints/containerized_ai_execution/ai_ready_k8s/" %}})
| AI-SC8 | Schedule a Kubernetes GPU Workload | Submit a pod that requests GPU resources and verify that Kubernetes schedules it onto an appropriate GPU node and that the GPU is visible inside the container. | | [Deployment of AI-ready Kubernetes]({{% relref "solutions/ai_factory_blueprints/containerized_ai_execution/ai_ready_k8s/" %}})
| AI-SC9 | Validate NVIDIA GPU Operator Integration | Deploy or validate the NVIDIA GPU Operator and confirm that GPU drivers, device plugins and related components are healthy and advertise GPU resources to Kubernetes. | | [Deployment of AI-ready Kubernetes]({{% relref "solutions/ai_factory_blueprints/containerized_ai_execution/ai_ready_k8s/" %}})
| AI-SC10 | Run Kubernetes Workloads on MIG Resources | Configure MIG resources and verify that separate Kubernetes workloads can request and consume individual MIG instances. | | [NVIDIA vGPU and MIG-backed vGPU]({{% relref "product/cluster_configuration/pci_passthrough_sriov/nvidia_mig_passthrough/" %}})
| AI-SC11 | Validate Manual Elastic GPU Capacity | Manually add GPU worker capacity to a Kubernetes cluster when additional resources are required and verify that the new capacity becomes available to workloads. Then remove the additional GPU worker capacity when demand decrease. | | [Deployment of AI-ready Kubernetes]({{% relref "solutions/ai_factory_blueprints/containerized_ai_execution/ai_ready_k8s/" %}})
| AI-SC12 | Validate Advanced GPU Scheduling with Run:ai (if required) | Submit GPU workloads through Run:ai and validate project allocation, quotas and scheduling policies against the OpenNebula-provisioned Kubernetes infrastructure. | | [Deployment of AI-ready Kubernetes]({{% relref "solutions/ai_factory_blueprints/containerized_ai_execution/ai_ready_k8s/" %}})


### Token Factory

| **ID** | **Task** | **Validation description** | **Video tutorials** | **Documentation** |
| ------ | ------------------------ | --------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- |
| AI-SC13 | Deploy an LLM Inference Service | Deploy an approved language model as an inference service on GPU infrastructure and verify that the service reaches a healthy state. | | [Inferencing with vLLM]({{% relref "solutions/ai_factory_blueprints/direct_ai_execution/llm_inference_certification/" %}})
| AI-SC14 | Access the Inference API | Expose the deployed model through an inference API endpoint and validate successful request and response processing from an authorized client. | | [Inferencing with vLLM]({{% relref "solutions/ai_factory_blueprints/direct_ai_execution/llm_inference_certification/" %}})
| AI-SC15 | Validate Tenant Isolation for Inference Services | Provide independent access to inference services for different tenants and verify that each tenant can access only its authorized endpoints and resources. | | [Multi-tenancy Models and User Roles]({{% relref "getting_started/understand_opennebula/opennebula_concepts/cloud_access_model_and_roles/" %}})
| AI-SC16 | Validate GPU-backed Inference | Run inference requests and verify that the model executes on the assigned GPU resources and that GPU utilization is observable during execution. | | [Inferencing with vLLM]({{% relref "solutions/ai_factory_blueprints/direct_ai_execution/llm_inference_certification/" %}})
| AI-SC17 | Measure Token Usage | Record input and output token consumption per tenant, project or API consumer so that usage can be attributed for operational, showback or billing purposes. | | [AI factory deployment and validation.]({{% relref "solutions/ai_factory_blueprints/overview/" %}})
| AI-SC18 | Monitor the Inference Service | Validate monitoring of service availability, request performance, model-serving metrics and underlying GPU utilization. | | [Prometheus and Grafana]({{% relref "product/cloud_system_administration/prometheus/overview/" %}})

### Elastic Slurm / HPC

| **ID** | **Task** | **Validation description** | **Video tutorials** | **Documentation** |
| ------ | ------------------------ | --------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- |
| AI-SC19 | Deploy a GPU-enabled Slurm Cluster | Provision a Slurm cluster with GPU-capable compute nodes and verify that Slurm detects and manages the GPU resources correctly. | | [Elastic Slurm Overview]({{% relref "platform_services/slurm/overview/" %}})
| AI-SC20 | Schedule a Slurm GPU Job | Submit a Slurm job that requests GPU resources and verify that it is scheduled onto an appropriate GPU node and completes successfully. | | [Elastic Slurm Overview]({{% relref "platform_services/slurm/overview/" %}})
| AI-SC21 | Run a Multi-GPU Slurm Job | Execute a workload that requests multiple GPUs and verify that all assigned GPUs are available to the job and operate as expected.| | [Elastic Slurm Overview]({{% relref "platform_services/slurm/overview/" %}})
| AI-SC22 | Validate Manual Elastic Slurm Capacity | Manually add compute capacity in response to queued workload demand and verify that the new nodes are provisioned, join the Slurm cluster, execute the workload, and can be manually released when no longer required. | | [Elastic Slurm Overview]({{% relref "platform_services/slurm/overview/" %}})
Loading