singletenant.ai

Vendors · Comparison

Together AI vs Vertex AI Provisioned Throughput

Side-by-side comparison of Together AI and Vertex AI Provisioned Throughput for single-tenant LLM hosting. Deployment options, compliance, pricing, and operational fit compared.

Together AI

Dedicated endpoints are explicitly single-tenant GPU instances with public per-hour pricing. Terms of service make no geographic storage commitment, so residency guarantee is recorded as none.

Vertex AI Provisioned Throughput

Provisioned Throughput reserves Gemini and partner-model capacity in generative AI scale units (GSUs) on fixed-cost terms, including 1-week options for select models. Google Cloud docs now brand the platform Gemini Enterprise Agent Platform. Vertex AI Platform appears in Google's SOC 1/2/3 and ISO 27001 scope; public GSU price list not verified at review.

From the data

Key differences

  • Category Together AI is a Dedicated GPU offering; Vertex AI Provisioned Throughput is Hyperscaler dedicated.
  • Compliance Both document SOC 2 Type II. Only Together AI documents HIPAA BAA. Only Vertex AI Provisioned Throughput documents SOC 3, ISO 27001 and GDPR DPA.
  • Data residency Together AI offers no residency guarantee (verified); Vertex AI Provisioned Throughput's residency guarantee is not verified.
  • Jurisdiction Both parent companies are under US jurisdiction.
  • Pricing Together AI: Per 1M tokens serverless; per GPU-hour for dedicated endpoints and clusters (published pricing). Vertex AI Provisioned Throughput: Per GSU (generative AI scale unit) fixed-cost term; 1-week terms for select models (publication not verified).

Generated from the verified vendor data below; sources and dates on each vendor profile.

Side by side

Capabilities compared

Together AI Vertex AI Provisioned Throughput
Category Dedicated GPU Hyperscaler dedicated
Deployment models ServerlessDedicated endpointGPU clusters Provisioned throughputSingle-zone provisioned throughput
Regions Not verified US, Europe, Asia-Pacific
Compliance SOC 2 Type IIHIPAA BAA SOC 2 Type IISOC 3ISO 27001GDPR DPA
Pricing model Per 1M tokens serverless; per GPU-hour for dedicated endpoints and clusters Per GSU (generative AI scale unit) fixed-cost term; 1-week terms for select models
Public pricing Yes Not verified
Residency guarantee No Not verified
Parent jurisdiction US US
Analyst note Dedicated endpoints are explicitly single-tenant GPU instances with public per-hour pricing. Terms of service make no geographic storage commitment, so residency guarantee is recorded as none. Provisioned Throughput reserves Gemini and partner-model capacity in generative AI scale units (GSUs) on fixed-cost terms, including 1-week options for select models. Google Cloud docs now brand the platform Gemini Enterprise Agent Platform. Vertex AI Platform appears in Google's SOC 1/2/3 and ISO 27001 scope; public GSU price list not verified at review.

Where they diverge

Deployment differentiation

Only Together AI

ServerlessDedicated endpointGPU clusters

Both

No overlap in deployment models.

Only Vertex AI Provisioned Throughput

Provisioned throughputSingle-zone provisioned throughput

Read the full profiles