Docs / Survive a VM outage

Keep your app and database running when one VM dies

Many apps run on a single VM: the web app and PostgreSQL side by side. It's simple, but when that VM goes down, your customers stop too. This guide moves such an app onto two VMs that take over for each other, without changing the app's code.

Test results

Tested on 30 September 2026 with a simple point-of-sale (POS) app writing one transaction per second to PostgreSQL 18 in zero-data-loss mode. The main VM held the web app and the primary database, and was powered off hard while transactions were running:

EventResult
VM A powered off hardVM B served transactions again after about 21 seconds (the database moved to B and visitors were sent to B, automatically).
Transactions already saved0 lost. The running transaction count continued without a gap.
VM A powered on againA rejoined as the database standby by itself and visitors went back to A. One request was slow; no data was lost.

During those ~21 seconds the cashier sees an error and has to retry. A transaction that was reported as successful is never lost.

Layout: 2 VMs + 1 small witness

VMRunsSize
AWeb app + PostgreSQL (primary)Like your current VM
BCopy of the web app + PostgreSQL (standby, always in sync)Same as A
C (witness)No data. Only votes on which server is primary, so there are never two primaries.Small (1 vCPU, 1 GB)

Steps

1. Connect the VMs to Saka

Prepare VMs A and B (and C if you're not reusing the old VM), then install the Saka agent on each, including the old VM. See Getting started. Each VM needs a public IP and UDP port 51871 open between the VMs (Saka opens it itself when it manages the firewall).

2. Create a PostgreSQL cluster

  1. Open Databases, press + Create database cluster, choose PostgreSQL.
  2. Data servers: A and B. Witness: C.
  3. Mode: Zero data loss (synchronous). A must for money transactions.
  4. Version: the same as or newer than the PostgreSQL you use now (default 18).

The cluster is usually ready in 1 to 5 minutes; you're told on Telegram.

3. Move the old database

  1. Stop the old app so no new transactions happen during the copy (e.g. after closing time).
  2. On the cluster page, card Move an existing database into the cluster: choose the old VM, paste the old database address as in the app's config, e.g. postgresql://kasir:[email protected]:5432/kasir, and press Copy to cluster.
  3. Saka copies everything (tables, indexes, views, sequences, extensions) in a single transaction, then compares the row count of every table. Test result: 200,500 rows in 3 seconds, counts identical.

4. Deploy the app on A

  1. Open Sites, Create site, App from Git, choose server A. A repo with a Dockerfile is used as is; see Apps from Git.
  2. On the cluster page, card Connect app servers: install the proxy on A. It always leads to the primary database, whichever VM that is.
  3. Set the app's DATABASE_URL variable to the proxy address for containers, e.g. postgresql://postgres:[email protected]:15432/kasir?sslmode=require (the port is shown on the card; press Show password for the password).
  4. Deploy, open the app, and check that the old data is there.

The app must reconnect to the database when its connection drops (nearly every framework does: Rails, Laravel, Django, Prisma, node-postgres with a pool). Without that, the app keeps failing after the database moves until it is restarted.

5. Add the second location on B

  1. On the app page, card Two locations, press Add a second location and choose B. B builds the same commit with the same variables; the database proxy on B is installed automatically on the same port.
  2. Change the app domain's DNS to a CNAME pointing at the load balancer address shown on the card (e.g. lbxxxx.lb.saka.work). A bare domain (e.g. shop.com) needs a DNS provider with CNAME flattening, such as Cloudflare; or use app.shop.com.
  3. HTTPS is set up on both VMs. Every deploy, rollback, and variable change on A is applied to B too.

6. Test it yourself (recommended)

Outside busy hours, power off VM A from your cloud provider's panel. Within about 30 seconds the app opens again from B with all data. Power A back on; it rejoins by itself. You're told on Telegram at each step.

Good to know

Technical details: Managed databases and Sites in two locations.