Take open AI from model to production on premises.
Discover validated models, deployable applications, and hardware-aware configurations engineered for secure AI on Dell infrastructure.
dell-ai brings catalog discovery, configuration, and deployment
to your terminal. Explore the CLI & SDK Latest optimized models
Recently added open models, ready for Dell infrastructure.
DeepSeek V4 Flash Vision Exp
New DeepSeek V4 Flash Vision Exp is an experimental multimodal MoE model built for efficient visual reasoning, tool use, and agentic workloads. It matches the performance of DeepSeek V4 Flash while heavily increasing multimodal agent capabilities.
Qwen3.8 Flash Next FP8
New Qwen3.8-Flash-Next-FP8 is a native multimodal MoE model with 125B language-model parameters and 6B activated parameters, plus efficient n-gram embeddings and MTP for long-context agentic workloads.
GLM 5.3 Flash
New GLM 5.3 Flash is a native multimodal 320B-parameter MoE model with 18B activated parameters, combining sparse and linear attention for efficient long-context coding and agentic workloads.
Qwen3.8 27B
New Qwen3.8-27B is a compact 27B dense vision-language model with native image and video understanding, flexible thinking control, and stronger coding, research, and long-horizon agentic performance. Context length is 262K natively and extensible to 1M tokens.
DeepSeek V4 Pro 0813
New DeepSeek V4 Pro 0813 is an updated release of the 1.6T-parameter MoE model with 49B activated parameters and a one-million-token context window for advanced long-context reasoning and agentic tasks.
NVIDIA Nemotron 3.5 Lightning 30B A3B NVFP4
New NVIDIA Nemotron 3.5 Lightning 30B A3B is a 30B-parameter sparse language model, provided in an NVIDIA-optimized NVFP4 format for efficient enterprise inference.
Latest updates
New models, platform support, and practical deployment guidance.
Qwen3.8 Flash Next and GLM 5.3 Flash land on Dell Enterprise Hub
Two new multimodal models are now available on Dell Enterprise Hub. Qwen3.8 Flash Next FP8, a MoE model with 125B parameters and just 6B active parameters, is the preview version of Qwen4, pairing efficient sparse attention with FP8 weights for long context workloads. Meanwhile, GLM 5.3 Flash, a 320B MoE model with 18B active parameters, is built for coding and agentic tasks. Both frontier open models achieve great performance and you already can deploy them on premise on Dell platforms. Deploy Qwen3.8 Flash Next FP8 → · Deploy GLM 5.3 Flash →
Announcing dell-ai v1.0: a CLI and SDK for Dell Enterprise Hub
dell ai v1.0 is a Python SDK and CLI for Dell Enterprise Hub. It lets you discover validated models, generate Docker, Kubernetes, or Helm configs, and deploy locally with automatic GPU allocation. It even comes with an agent skill so tools like Cursor and Claude Code can run the same workflow from natural language. Go take a look at the release announcement and start using it today! Read the announcement → · View on GitHub →
Qwen3.8 27B is now available on Dell Enterprise Hub
Qwen3.8 27B is now available on Dell Enterprise Hub. A compact 27B dense vision language model with native image and video understanding, flexible thinking control, and stronger coding, research, and long horizon agentic performance. Deploy it with vLLM 0.27.1 on Dell Pro Max with GB10, 4×L40S, or 1×, 2×, and 4× H100 and H200 systems. Deploy Qwen3.8 27B →
DeepSeek V4 Pro 0813 lands on Dell Enterprise Hub
DeepSeek V4 Pro 0813 is now available on Dell Enterprise Hub. This updated 1.6T parameter Mixture of Experts model with 49B activated parameters delivers advanced long context reasoning and agentic capabilities. Deploy it on prem with vLLM 0.27.1 on H200 and B300 systems. Deploy DeepSeek V4 Pro 0813 →
Deploy complete AI applications
Go beyond endpoints with production-ready AI experiences.
Super Analyzer
Super Analyzer is an application that identifies and evaluates common anti-patterns in C++, Java, Python, and Rust code. It uses a multi-agent architecture with three coordinated agent types—Primary, Fixer, and Chat—to deliver accurate analysis and iterative improvements. The Web UI supports natural multi-turn interactions and uses a PostgreSQL database to persist in user accounts and conversation history across sessions. Super Analyzer also provides a REST/API layer and a Python interface, all secured through user/password authentication backed by PostgreSQL. Access is protected with JWTs signed using RSA-256, ensuring a consistent and secure experience across all entry points.
OpenWebUI
OpenWebUI is an extensible, feature-rich, and user-friendly self-hosted AI platform designed to operate entirely offline. It provides a modern interface for interacting with various LLM runners including Ollama and OpenAI-compatible APIs. With built-in RAG (Retrieval Augmented Generation) capabilities, OpenWebUI offers a comprehensive solution for AI deployment and interaction in Kubernetes environments.
AnythingLLM
AnythingLLM is an open-source, all-in-one AI chat application with built-in RAG (Retrieval Augmented Generation) capabilities. It provides a user-friendly interface for document management, LLM integration, and AI agent creation - all without complex setup. The application supports both local and cloud-based deployments, making it ideal for teams and organizations seeking a private, customizable AI solution.
Agentic Smart Router
In an Agentic AI application, not all prompts are the same, different prompts need to have different requirements in terms of complexity of the prompt - some need reasoning, some need access to RAG system and some need tool calling capabilities. In this custom-built Agentic Smart Router, we demonstrate that capability by leveraging NVIDIA Agent Intelligence Toolkit, NVIDIA NIMs, NVIDIA LLM Router blueprint and stitching them together to build an Agentic Smart Router application.
Optimized for Dell platforms
From AI PCs to accelerated PowerEdge servers, deploy with configurations tested for the hardware you already run.
H200NVIDIA H200 141GB HBM3e
RTX PRO 6000NVIDIA RTXPRO6000 96GB GDDR7
L40SNVIDIA L40S
MI355XAMD MI355X 288GB HBM3E
GB10NVIDIA GB10 Blackwell
H100NVIDIA H100 PCIe
B300NVIDIA B300 288GB HBM3e
MI300XAMD MI300X 192GB HBM3
Gaudi 3INTEL GAUDI3 192GB HBM3
L40SNVIDIA L40S
H200NVIDIA H200 141GB HBM3e
MI300XAMD MI300X 192GB HBM3
H100NVIDIA H100 PCIe
MI355XAMD MI355X 288GB HBM3E
RTX PRO 6000NVIDIA RTXPRO6000 96GB GDDR7
Gaudi 3INTEL GAUDI3 192GB HBM3
B300NVIDIA B300 288GB HBM3e
GB10NVIDIA GB10 Blackwell