singletenant.ai

Vendors

16 providers of dedicated or single-tenant LLM hosting, ordered alphabetically. Click a vendor for the full evidence-linked profile.

Vendor Category Deployment Regions Compliance Pricing Data residency
Amazon Bedrock Provisioned Throughput Hyperscaler dedicated Provisioned throughput, Custom model (provisioned) US, Europe, Asia-Pacific SOC 2 Type II, SOC 3, ISO 27001, HIPAA BAA, C5, FedRAMP, GDPR DPA Per Model Unit (MU) per hour; no-commitment, 1-month, or 6-month terms Guaranteed (US parent)
Azure OpenAI Provisioned Throughput (PTU) Hyperscaler dedicated Global provisioned, Data-zone provisioned, Regional provisioned Global, US, Europe SOC 2 Type II, ISO 27001, C5, FedRAMP, HIPAA BAA, GDPR DPA Per PTU per hour; Azure Reservations discount with 1-month or 1-year terms Guaranteed (US parent)
Baseten Dedicated GPU Dedicated deployment, Self-hosted, Hybrid US, UK, Europe, Asia-Pacific SOC 2 Type II, SOC 3, HIPAA BAA, GDPR DPA Per 1M tokens (Model APIs); per-minute GPU/CPU instances for dedicated deployments Not verified (US parent)
Civo UK sovereign Bare metal, Managed Kubernetes, Dedicated GPU inference UK ISO 27001, SOC 2 Type II Per-GPU-hour on demand, lower rates on 6-36 month commitments Guaranteed (UK parent)
CoreWeave Dedicated GPU Bare metal, Managed Kubernetes, On-demand, Spot, Reserved capacity US, UK, Europe SOC 2 Type II, GDPR DPA Per-GPU-hour on-demand and spot; reserved capacity at up to 60% discount Not verified (US parent)
Databricks Model Serving Managed self-host Custom model serving, Provisioned throughput, Pay-per-token endpoints US, Europe, Asia-Pacific SOC 2 Type II, ISO 27001, HIPAA BAA, FedRAMP, GDPR DPA Per DBU per hour by GPU instance size (e.g. A10G 20 DBU/hour); per-token Foundation Model APIs Not verified (US parent)
Fireworks AI Dedicated GPU Serverless, On-demand dedicated, Reserved capacity Global, US, Europe, Asia-Pacific SOC 2 Type II, ISO 27001, HIPAA BAA Per 1M tokens (serverless); per-GPU-second on-demand deployments; reserved capacity Not verified (US parent)
Hugging Face Inference Endpoints Managed self-host Dedicated endpoints, Public endpoint, Protected endpoint, Private endpoint (PrivateLink) US, Europe SOC 2 Type II, GDPR DPA Per instance-hour (CPU, GPU, and accelerator instances); dedicated inference from $0.033/hour Not verified (US parent)
Modal Dedicated GPU Serverless US, UK, Europe, Asia-Pacific SOC 2 Type II, HIPAA BAA Per-second GPU, CPU core, and memory billing Guaranteed (US parent)
Nebius EU sovereign Serverless API, Dedicated endpoints Europe, US SOC 2 Type II, SOC 3, ISO 27001, GDPR DPA Per 1M input/output tokens; shared and dedicated endpoint tiers Not verified (NL parent)
NVIDIA NIM Managed self-host Self-hosted containers, Dedicated endpoints, NVIDIA-hosted API, DGX Cloud Not verified ISO 27001, GDPR DPA NVIDIA AI Enterprise per GPU ($4,500/GPU/year self-managed; $1/GPU/hour in cloud marketplaces plus CSP instance costs) Not verified (US parent)
OVHcloud EU sovereign Serverless API, Managed model deployment Europe ISO 27001, C5, SecNumCloud, GDPR DPA Per 1M tokens (AI Endpoints, input/output priced separately) Not verified (FR parent)
RunPod Dedicated GPU Secure Cloud, Community Cloud, Reserved clusters US, Europe, Asia-Pacific SOC 2 Type II, SOC 3, HIPAA BAA, GDPR DPA Per GPU-hour (per-second display available); reserved clusters via sales Not verified (US parent)
Scaleway EU sovereign Serverless API, Dedicated GPU inference Europe ISO 27001, GDPR DPA Per-hour dedicated GPU (Managed Inference); per token (Generative APIs) Guaranteed (FR parent)
Together AI Dedicated GPU Serverless, Dedicated endpoint, GPU clusters Not verified SOC 2 Type II, HIPAA BAA Per 1M tokens serverless; per GPU-hour for dedicated endpoints and clusters No guarantee (US parent)
Vertex AI Provisioned Throughput Hyperscaler dedicated Provisioned throughput, Single-zone provisioned throughput US, Europe, Asia-Pacific SOC 2 Type II, SOC 3, ISO 27001, GDPR DPA Per GSU (generative AI scale unit) fixed-cost term; 1-week terms for select models Not verified (US parent)

Category reflects a vendor's parent jurisdiction; a data-residency guarantee is a separate, per-vendor fact verified on its own evidence, so a sovereign-category vendor (EU or UK) can still read "Not verified" for residency (see why an EU region is not EU sovereign).