Prebuilt, GPU-optimised inference containers — deploy any model anywhere with one command.
Reasoning Mode
A deliberate "thinking" mode that works through multi-step problems before answering.
Multilingual Support
Understands and replies in dozens of languages, including translation between them.
Text Generation
Drafts, rewrites and edits long-form and short-form copy.
Image Generation
Creates original images from a text prompt.
Image Understanding
Reads photos, screenshots, diagrams and handwriting and answers questions about them.
Audio Transcription
Turns recorded speech into searchable, summarised text.
Data Connectors
RAG blueprintsConnects to external stores such as Drive, SharePoint or a code host.
Public API
OpenAI-compatibleA documented API for building the same capability into your own product.
Code Generation & Debugging
Writes, explains, reviews and fixes code across common languages.
Custom Assistants
Fine-tuned modelsBuild and share a configured assistant with its own instructions and knowledge.
Task Automation
Schedules recurring jobs or runs multi-step tasks without supervision.
Third-Party Integrations
Connects to other SaaS products through apps, connectors or an integration platform.
Team Workspaces
Shared workspace with pooled seats, billing and admin management.
SSO & SAML
EnterpriseSingle sign-on through your identity provider.
Admin Controls & Roles
Central console for seats, roles, permissions and usage.
Data-Retention Controls
Self-hostedConfigure how long conversation data is kept, or turn training off.
SOC 2 Compliance
Independently audited against the SOC 2 security framework.
GDPR Compliance
Offers a data-processing agreement and EU data-protection commitments.
Encryption at Rest & in Transit
Traffic and stored data are encrypted with industry-standard ciphers.
Prices are set by the vendor and may change.
$0 /month
What's included
$0
What's included
Hourly usage
What's included
Per GPU /year
What's included
NVIDIA NIM (NVIDIA Inference Microservices) packages models as prebuilt, GPU-optimised containers with a standard OpenAI-compatible API. Instead of assembling CUDA versions, inference engines and serving code yourself, you pull a container and run it.
The point is portability. The same NIM container runs on a workstation, in your datacentre, or on any cloud — with NVIDIA's inference optimisations already applied, typically delivering far better throughput than a naive deployment of the same model.
You can try every model free on build.nvidia.com before deploying anywhere.
What it removes from an MLOps team's job:
NIM is enterprise infrastructure.
It is not for individuals — for casual use, a hosted API is far simpler.
More open models packaged as ready-to-run NIM containers.
Reference architectures for retrieval-augmented generation deployments.
Hourly pay-as-you-go through the major cloud marketplaces.
Inference engine updates raising tokens per second on supported GPUs.
Join the users who trust NVIDIA NIM for their daily tasks.