Best platforms to host MCP servers and AI agents in 2026

Sep 18, 2026·10 min read
Photo of Rak Siva
Rak Siva

The Suga team makes a lot of outreach calls to understand what our users are working on, and agents come up more often than anything else. Hosting a typical web app is pretty well solved. Agents build on that, but they put different demands on the host: they can run for minutes instead of seconds, need somewhere to keep context between runs, and often depend on several MCP servers and other services working together.

The applications, models, and frameworks vary, but most of the agents we see on Suga look like the following:

  • MCP servers wired around a coding assistant.
  • Chatbots running against their own data.
  • Background workers triggered by webhooks or schedules.

This guide compares five platforms for hosting agents: Suga, Cloudflare Workers, Railway, Fly.io, and Vercel. We chose platforms that can handle the infrastructure around an agent: the runtime, state, MCP servers, and other services. The model itself runs at a provider such as Anthropic, OpenAI, or OpenRouter.

Pricing and feature availability were verified on September 18, 2026.

What matters when hosting an agent

The host runs the agent's code, its tools, its state, and the services it depends on. That's a different workload from a typical web application.

  • Long-running compute: An agent response can run for minutes instead of seconds. A runtime that caps requests or scales an idle process to zero can interrupt the job halfway through.
  • Persistent state: Chat history, checkpoints, embeddings, and other state need to live somewhere durable. Postgres and Redis are useful because the data isn't tied to one particular agent runtime or platform.
  • Recovery: An agent can fail halfway through a response. If you want it to pick up where it left off, you need checkpoints in a datastore or a workflow engine that knows how to resume the work.
  • Agent dependencies: An agent rarely runs by itself. It might need several MCP servers, a database, Redis, and a vector store alongside it. Keeping those services together on a private network becomes increasingly important as the application grows.

1. Suga

Suga was built around the idea that an agent is more than a container. Its runtime, MCP servers, datastores, and other services can run together in the same environment.

Long-running compute

Suga has no request timeout and doesn't scale to zero unless you turn it on. A streamed response from Anthropic can keep flowing until the client closes the connection.

Multi-container environments

On Suga, those services run as separate containers in the same environment. They share a private network and reach each other by internal hostname. Adding another MCP server or datastore means adding another container to the environment rather than creating a separate project or account.

State and recovery

Postgres, Redis, and vector stores run as their own containers in the same environment, using standard interfaces that any language can work with. Replicas and one-click rollbacks handle routine crashes, and because state lives outside the runtime, a restarted container can reload it from Postgres or Redis.

You can export the environment configuration and rebuild it on another Suga cluster, or on your own Kubernetes cluster with Enterprise. Teams can switch model providers, frameworks, or orchestration tools without rebuilding the whole application from scratch.

AI-driven config via Suga MCP

Suga also exposes its own MCP server at dashboard.suga.app/api/mcp, so an AI agent can set up a Suga environment for you. Point an editor's MCP client at Suga's URL and give it a prompt like this. The AI agent creates containers, sets environment variables, connects services, and stages the changes as a draft for you to review and apply.


Add a Qdrant vector store and a filesystem MCP server to my agent environment. Wire QDRANT_URL and MCP_FILESYSTEM_URL on the agent-runtime container so it can reach them, and set ANTHROPIC_API_KEY as a shared sensitive variable across the runtime and any MCP server that needs it.

Paste into an AI agent connected to the Suga MCP.

Environment forking for evaluation

Suga lets you fork a production environment into a separate environment and test changes against real production data without touching production. Non-sensitive variables copy over by default. For sensitive values, you decide whether to copy or reset each variable.

Pricing

  • Free: $0/month, limited resources.
  • Pro: $20/month per seat, including $20 in monthly hosting credits per seat.
  • Enterprise: custom pricing, including bring-your-own-cloud.

A small agent (0.25 vCPU, 256 MiB) and a Postgres instance for memory (0.125 vCPU, 512 MiB, 2 GB disk) cost about $17/month, which is covered by the Pro plan's credits.

Beyond the included credits, the rates are:

  • vCPU: $0.05/hour
  • Memory: $0.0054/GB-hour
  • Storage: $0.00015/GB-hour
  • Egress: $0.12/GB

See Suga's full pricing for volume discounts and Enterprise details.

Full walkthrough: Deploy an MCP server on Suga.

2. Cloudflare Workers

Cloudflare takes a different approach to agents. The request runtime is a Worker, Durable Objects provide per-session state, and the Agents SDK exposes an Agent class that combines them.

Each agent session becomes a Durable Object with its own SQLite database, addressable by ID.

The result is a stateful agent architecture that fits naturally with Cloudflare's edge platform but relies more heavily on Cloudflare-specific services.

Durable Objects for agent state

State lives inside each Durable Object as an on-instance SQLite database. Latency is low and consistency is strong, but this approach is specific to Cloudflare. Moving to another host means replacing the Durable Object state layer with a different database. R2 and D1 reduce dependence on the per-session Durable Object pattern for object and relational storage but are still Cloudflare services, so switching platforms means migrating them as well.

Long-running work

Workers on the Paid plan have a default CPU time limit of 30 seconds per invocation. This limit can be increased to 5 minutes through Wrangler configuration. CPU time measures active code execution, so network I/O such as fetch calls to Anthropic or database queries doesn't count against it. HTTP-triggered Workers have no separate wall-clock limit, so a long streaming response stays within the CPU budget as long as the code processing it does.

For work that outgrows the CPU budget, Durable Object alarms let an agent schedule a callback, return, and resume later. WebSockets can also hibernate for hours between messages without consuming CPU time.

Recovery

Durable Objects hibernate when idle and wake up with their state intact. Alarms can drive scheduled work through crashes and restarts. The Agents SDK also exposes Fibers for durable execution. A fiber records intermediate state to SQLite, so a Durable Object evicted mid-response can resume from the stored snapshot on its next activation. Because recovery is tied to the Durable Object runtime, moving that logic to another platform means replacing the Cloudflare-specific pieces.

Agent dependencies

Cloudflare's architecture centers on Workers, with Durable Objects, R2, and D1 available in the same account.

MCP servers can run elsewhere and be called over HTTP, or be rewritten for the Workers runtime. There isn't a way to run an arbitrary container next to a Worker as a private service.

Pricing

  • Free: 100,000 requests/day, 10 ms CPU per invocation.
  • Paid: $5/month base, including 10M Worker requests and 30M CPU-milliseconds.
  • Overage: $0.30 per million requests and $0.02 per million CPU-ms.
  • Durable Objects: the monthly allowance includes 1M requests, 400,000 GB-s of duration, and 5 GB of SQLite storage. Overage runs $0.15/million requests, $12.50/million GB-s, and $0.20/GB-month.

3. Railway

Railway looks familiar if you're coming from a traditional PaaS. You deploy an agent, add the services it needs, and put them together in a project. An agent runtime and its MCP servers can run as separate Railway services on the same private network.

Railway covers the basics: long-running compute, standard datastores, private networking, and straightforward service deployment. Resuming an agent after a crash is something you build yourself.

Projects and private networking

Railway projects can hold multiple services on the same private network, so an agent runtime, MCP server, and Postgres can communicate through internal DNS. Cross-service variable references also make it easier to share credentials between services.

Standard datastores

Postgres, Redis, MySQL, and MongoDB can be deployed from Railway templates as their own services. State lives in standard datastores, so the data can be moved independently of the agent runtime. Cross-service variable references such as ${{Postgres.DATABASE_URL}} reduce duplication when services need to share configuration.

Long-running compute and recovery

There's no request timeout and no scale-to-zero unless you opt in. Streaming responses pass through end to end.

Replicas and one-click rollbacks handle routine crashes. Because state sits in Postgres or Redis, a restarted service can reload it on startup. Railway doesn't provide a workflow engine for resuming an agent halfway through a job. That logic is something you build against your own database or workflow tool.

Pricing

  • Hobby: $5/month, including $5 in usage credits.
  • Pro: $20/month per workspace, applied as $20 in usage credits.

Usage beyond the included credits is charged at:

  • $20/vCPU-month
  • $10/GB memory-month
  • $0.05/GB egress

Usage is metered per second.

4. Fly.io

Fly.io takes a lower-level approach. Applications run in Firecracker microVMs, giving each Machine a full Linux userspace, block storage, and TCP/UDP support in the regions you choose.

That control is useful when an agent needs more than an HTTP endpoint, but there are more concepts to manage: Machines, Volumes, Apps, Organizations, and Networks.

Fly covers long-running compute and standard datastores, but several services take more setup to deploy and operate.

Full Linux environments

There is no request timeout, and sub-second boots let you wake a Machine on request and stop it when idle. Because Machines provide a full Linux userspace with TCP and UDP support, an agent can run services that don't fit neatly as serverless HTTP endpoints.

Persistent storage

Fly Managed Postgres and Upstash Redis are standard datastores accessed over their usual protocols. Fly Volumes provide attached storage for local state that needs to live on the Machine itself, such as an MCP server's knowledge graph. Moving off Fly can mean exporting Postgres and taking a Volume snapshot.

Recovery and multi-region deployment

Machines automatically restart on failure and can be replicated across regions behind an anycast IP. Volume snapshots can restore local state. Fly doesn't include a workflow engine, so resuming work after a mid-response failure is something you build against Postgres or an external workflow tool.

Agent dependencies

A Fly app can host multiple Machines, and separate apps can share an internal WireGuard network. The agent runtime can run as one Machine or app, MCP servers as others, and databases as separate apps or managed services. This works, but each app has its own deployment and configuration. Putting the whole stack together takes more manual setup than on a platform that groups services into an environment.

Pricing

Fly is usage-based and billed per second.

  • Shared CPU: from ~$0.0028/hour (shared-cpu-1x, 256 MB).
  • Performance CPU: from ~$0.044/hour (performance-1x, 2 GB).
  • Volumes: $0.15/GB/month.
  • Egress: $0.02 to $0.12/GB depending on region.

5. Vercel

Vercel is where a lot of JavaScript agents end up, largely because of the AI SDK. Fluid Compute handles requests, while Vercel Workflows handles work that needs to continue beyond a single function invocation.

Fluid Compute

Fluid Compute keeps function instances warm and can run multiple I/O-bound requests concurrently in the same instance.

On the Pro plan with Fluid Compute, function duration defaults to 300 seconds and can be raised to 800 seconds. An extended beta pushes the maximum to 1800 seconds (30 minutes) for specific Node.js, Bun, and Python runtime versions when configured per function rather than as a project default. The limit measures wall-clock time, while billing is based on Active CPU time that pauses during I/O, so a long streaming response spends most of that window waiting on the model provider rather than on billed compute. Work that needs to outlive a single function invocation belongs in Vercel Workflows.

Workflows for recovery

Vercel Workflows runs multi-step logic as durable code. A workflow can retry failed steps, wait for external events, pause for minutes or months, and resume across crashes and deployments through deterministic replay. An agent expressed as a Workflow can therefore resume after a crash without you writing the resume logic yourself.

Databases and external services

Marketplace-integrated Postgres, KV, and blob storage sit outside Vercel Functions and use standard interfaces. The databases themselves are portable, although moving the whole application still means migrating each service.

Vercel's compute centers on serverless functions. Vercel Sandbox extends that with Firecracker microVMs that run arbitrary OCI container images from Vercel Container Registry, which suits isolated agent execution and untrusted code. Sandbox is a separate runtime from Vercel Functions, not a shared private network where a function and a long-running MCP server sit side by side in the same environment. An MCP server that runs continuously alongside a Vercel Function typically runs elsewhere or gets rewritten as a function.

Pricing

  • Hobby: free for personal projects.
  • Pro: $20/month per member.
  • Enterprise: custom.

Fluid Compute is priced based on Active CPU time, provisioned memory, and invocations. See Vercel's pricing page for current rates.

Which one to pick

All five platforms can host an agent. The choice mostly comes down to how much infrastructure you want the platform to manage and how much of your agent you want tied to its runtime.

  • Suga — when the agent, MCP servers, databases, and supporting services need to operate as one portable environment.
  • Cloudflare — when you want an edge-native agent architecture and are comfortable building around Workers and Durable Objects.
  • Railway — when you want conventional long-running services and standard databases with minimal platform-specific architecture.
  • Fly.io — when you need full Linux environments, unusual networking, or regional placement and don't mind more infrastructure work.
  • Vercel — when you're building primarily around serverless/JavaScript and want durable workflows for jobs that outlive a function invocation.

FAQ

What transport does MCP use, and which of these hosts support it?

The current MCP spec defines two standard transports: stdio for a client-launched subprocess and Streamable HTTP for remote servers. On Streamable HTTP, each message is an HTTP POST and replies arrive as either a JSON object or a request-scoped Server-Sent Events stream. Stdio is a local process convention, so remote MCP servers use Streamable HTTP. All five platforms can host an MCP server over HTTP.

Do I need GPUs to host an agent?

Not for API-driven agents. If the LLM runs at Anthropic, OpenAI, or a provider fronted by OpenRouter, the agent host only needs enough CPU and memory to run the agent's code and its MCP tools. GPU-first infrastructure becomes relevant when you're serving the model yourself.

How do agent workloads differ from typical web workloads?

Agents can hold sessions open for minutes rather than milliseconds, stream large responses, keep per-session state between requests, call several external APIs during a response, and use several services together. The differences between hosting platforms therefore show up in streaming support, state management, recovery, and how easily those services can be connected.

Have more questions? Join our Discord.

Deploy anything.

Available now.