How Expanso Cloud Runs: Kubernetes, Cell-Based Architecture, and Practicing What We Preach

How we re-architected Expanso Cloud with cell-based Kubernetes clusters, custom operators, and sub-10-second orchestrator provisioning - and why we run on the same infrastructure we recommend to customers.

Earlier this week we described our private control-plane design-partner preview. This post explains the managed Expanso Cloud architecture that serves as the preview’s technical starting point. The private package is not generally available or self-service.

We tell customers to process data where it lives, on edge devices, behind firewalls, across distributed infrastructure. Compute should be portable, resilient, and close to the source. That’s the whole pitch. So when we re-architected Expanso Cloud, we held ourselves to the same standard. It now runs on Kubernetes with a cell-based architecture that spins up fully functional orchestrators in under ten seconds.

Quick refresher: what Expanso Cloud actually does

If you already know the product, skip ahead. Otherwise, the way we built Expanso Cloud follows directly from how the product works, so it’s worth a quick rundown.

The foundation is Expanso Edge, a lightweight agent that runs on customer-controlled infrastructure. It can transform, filter, and route pipeline payloads locally. Data is sent only to the outputs a customer configures; Expanso Cloud still receives the operational metadata required to coordinate and monitor the fleet.

Each Edge node connects to an Expanso Orchestrator, the control plane that coordinates what runs where. You deploy a logging pipeline or a data transformation job, and the orchestrator persists that intent and distributes it to matching edge nodes. It handles versioning, gradual rollouts, rollbacks, and monitoring. Pipeline payloads do not traverse the orchestrator, while fleet and execution metadata do.

The orchestrator talks to edge nodes over NATS, a lightweight messaging system built for high-throughput bidirectional communication. Edge nodes connect inbound; the orchestrator pushes configuration outbound. That’s the entire data flow through the control plane.

Expanso Cloud is the fully managed experience on top of all this. You create an account, set up an org, provision orchestrators through the web UI. It’s the front door to the platform, and the piece we just re-architected.

From minutes to seconds: the architecture behind fast provisioning

Spinning up a new orchestrator, one that’s ready to accept edge node connections and deploy pipelines, used to take minutes. Now it takes under ten seconds. Getting there meant rethinking how we provision and manage orchestrators at the infrastructure level.

Cells: independent, isolated Kubernetes clusters

Expanso Cloud uses a cell-based architecture. Each cell is a fully independent, isolated Kubernetes cluster. When you create a new network (that’s our term for a managed orchestrator instance), the platform picks the right cell and provisions your orchestrator there.

Cells aren’t single-tenant. Each one hosts orchestrators for multiple customers, but with strict isolation at the Kubernetes level, separate namespaces, dedicated service accounts, independent secrets, isolated network routing. If something goes wrong in one cell, it doesn’t cascade to others. Blast radius is contained by design.

For us, cells give us the operational boundaries we need to scale, update, and maintain infrastructure without coordinated global deployments. They’re also the foundation for geographic distribution: we can place orchestrators closer to the infrastructure they’re coordinating.

The provisioning itself is Kubernetes-native. We built a custom operator around a CRD called ExpansoCloudNetworkInstance. When the platform creates this resource in a cell, the operator’s reconciliation loop kicks in and stands up everything the orchestrator needs:

  • A Kubernetes Deployment running the orchestrator with an embedded NATS broker
  • A ClusterIP Service
  • Traefik IngressRoutes for HTTP API access and raw TCP for NATS connections
  • Secrets for credentials and platform configuration
  • Persistent storage for orchestrator state

The operator handles the full lifecycle: creation, updates, health monitoring, and teardown. It’s level-triggered, so it continuously reconciles toward the desired state. Pod crashes, it comes back. Secret rotates, the deployment rolls forward.

Provisioning is fast because creating a network is just creating a Kubernetes custom resource and letting the operator converge.

The controller: orchestrating the orchestrators

Between the Expanso Cloud web app and the cells sits the controller. It’s the piece of this re-architecture we’re most proud of, and it solves problems that don’t seem like problems until you’re managing hundreds of orchestrators.

The controller is an async service that manages orchestrator lifecycles across all cells. When you click “Create Network” in the UI, we don’t call a cell directly. Instead, we record the intent: your network config, capacity requirements, and target version. The controller picks it up from there. It selects the right cell, talks to the cell’s provisioning API, and monitors the result. If something fails, it retries with backoff.

This decoupling is what makes provisioning reliable, not just fast. The web app declares what should exist; the controller makes sure it does.

But the real value shows up beyond initial provisioning. The controller continuously monitors every orchestrator across every cell, health status, node counts, endpoint availability. And it handles the thing that gets genuinely painful at scale: rolling updates across hundreds of orchestrators.

When we ship a new orchestrator version, we don’t need to coordinate a synchronized deployment across every cell. We tell the controller: update from version X to version Y. It handles the rollout declaratively, same philosophy as Kubernetes Deployments, rolling through networks at a controlled pace, checking health at each step. Same thing when you update your orchestrator config through the platform. Whether it’s our ops team pushing a version or you changing a setting, the controller reconciles the desired state.

This is the same reconciliation pattern that makes Kubernetes operators work: declare desired state, let the system converge, and handle failures gracefully. Here it applies at the platform level, not just infra.

Direct-to-cell: no global routing overhead

Once your network is provisioned, all traffic goes directly to the cell. Your orchestrator’s HTTP API endpoint and NATS endpoint point straight to the cell’s ingress, no global proxy, no centralized routing layer, no extra hop.

When you interact with your orchestrator, through the CLI, the API, or the Cloud UI, traffic goes directly to the cell. When your edge nodes connect over NATS, they connect directly to the cell. The global control plane is only involved during provisioning and lifecycle management.

The upshot: our global control plane can go down for maintenance and your running orchestrators don’t notice. Edge nodes keep processing, pipelines keep running, API calls keep working. Each cell is self-sufficient once provisioned.

Preparing a private control-plane preview

Building on Kubernetes and cells gives us a starting point for a customer-operated control plane. We are validating that deployment only through qualified enterprise design-partner engagements today; it is not generally available or self-service.

The orchestrator does not process pipeline payloads; it coordinates the fleet and receives operational metadata. Many users choose the managed service on that basis. Some organizations have requirements that also place the control plane inside their perimeter, which is the need the design-partner preview is validating.

Each managed cell is a self-contained Kubernetes deployment with an operator, CRDs, and Helm charts. Those components are the technical base for the preview, but customer-operated packaging has additional installation, upgrade, secret-management, security, and support requirements. We qualify those requirements per design partner rather than claiming the managed deployment can be installed unchanged in any customer cluster.

When we tell customers to deploy on portable, resilient infrastructure, our managed platform runs on those same Kubernetes primitives. Preview environments use a narrow, reviewed support envelope.

Every developer on the team runs a full Expanso Cloud environment locally, not a simplified mock, not a Docker Compose approximation, but the actual cell-based architecture with the operator, provisioning, and monitoring. We test against the same topology that runs in production.

What this means for you

If you’re already running pipelines on Expanso: faster provisioning, better reliability through cell isolation, and orchestrator updates that don’t touch your running workloads. Direct-to-cell means your day-to-day doesn’t depend on global infrastructure.

If you’re evaluating us: Managed Expanso Cloud is generally available. If a private control plane is mandatory, book a demo to discuss whether your environment qualifies for the enterprise design-partner preview.

We will announce broader availability only after the design-partner installation, upgrade, security, and support gates pass. You can try managed Expanso Cloud today.