LAS VEGAS – Broadcom today announced VMware AI Factory, the software-defined foundation of VMware Private AI Cloud, at the company’s VMware Explore 2026 conference. VMware AI Factory provides customers a simplified path to production AI with new automation innovations for deploying AI-ready infrastructure and supporting Day 2 operations. With VMware AI Factory, customers can achieve faster time to first model deployment and better manage AI tokenomics.
“Enterprises want to run AI where their data lives, but the journey from metal to model is slow, complex, and expensive,” said Paul Turner, chief product officer, VMware Cloud Foundation Division, Broadcom. “VMware AI Factory changes that. We give customers a software-defined foundation that automates infrastructure deployment, unifies lifecycle management, and lets them choose their preferred hardware and vetted models. The result is faster time to first model, predictable private cloud costs, and better control over AI tokenomics.”
VMware AI Factory Accelerates Time to First Model
VMware AI Factory brings AI applications directly to enterprise private data within a secure private cloud environment. VCF’s unique infrastructure automation capabilities can reduce the time from bare metal server deployment to serving the first AI model from weeks to a matter of hours. VMware AI Factory streamlines AI infrastructure management by fully automating hardware provisioning, software stack enablement, and end-to-end lifecycle management. By integrating hardware and software operations into a unified, automated solution, organizations can rapidly scale AI workloads while minimizing operational complexity.
“Organizations are shifting their focus back on on-premises private clouds, and they are repatriating their production workflows,” Prashanth Shenoy, VP of Product Marketing for the VMware Cloud Foundation Divisoin, told SD Times in an interview.
As part of the VMware AI Factory, private AI services help make AI operational, governable, and cost-effective. VCF pools and shares GPU resources across the organization so teams can run multiple models on shared hardware instead of dedicating infrastructure to each workload. A unified model gallery gives IT and data science teams a single interface for deploying and managing model inference and RAG workflows across VMs, containers, and GPU resources, with built-in observability into token throughput, latency, and compute and memory utilization.
Enterprises can pivot to new models while keeping costs low through shared infrastructure and governed models-as-a-service. New and forthcoming private AI services include:
● Multi-tenant Model Sharing: Model Runtime now supports secure sharing of AI models between tenants or lines of business through isolated namespaces, maintaining data privacy while eliminating redundant model deployments that waste GPU allocation and infrastructure resources.
● AI Gateway: Unified model governance between on-premises and cloud environments through a single consumption interface, enabling access to locally hosted models with enhancements like intelligent prompt routing, token and usage rate-limiting, and application authorization.
● Secure AI Sandboxes and Governance: Secure virtualized container spaces will isolate dynamic agent-generated code execution, with a control layer defining how agents are invoked, what tools they access, and how their outputs are validated before being acted upon.
To further streamline VCF AI Factory deployment, Broadcom is announcing a new partnership with MetalSoft to deliver integrated heterogeneous bare metal automation for VCF that drops bare-metal provisioning time from weeks to minutes. The integration will help IT provision or repave physical servers from multiple vendors directly through the VCF management console, unifying the software and hardware lifecycle into a single operational model and eliminating the need for vendor-specific tools for hardware and firmware management.
VMware AI Factory combines VMware Cloud Foundation (VCF) with certified VCF AI ReadyNodes from Cisco, Dell Technologies, Lenovo, Supermicro and others and customers’ preferred AI software and accelerator architectures. Broadcom and AMD are collaborating to deliver a VMware AI Factory that pairs VCF with AMD Instinct GPUs and the open AMD ROCm software ecosystem. Zero-touch provisioning will orchestrate the end-to-end deployment of the entire stack, from vSphere and vSAN through Kubernetes and the AMD GPU operator, and the AMD DVX driver can attach GPUs to large VMs consumed by a VMware vSphere Kubernetes Service cluster.
VMware AI Factory gives enterprises a production-ready path to running leading AI models on-premises. VCF customers can run more than 150 open source and commercial models, including Nemotron 3, Gemma 4, cotomi, Qwen 3.7-Max, and GLM 5.2.
Broadcom is working with the world’s leading AI model providers to give organizations a clear path to data sovereignty and cost-effective AI at scale, with leading models available securely and delivered as a service to their user community through VCF’s built-in services.

