Upgrades
Back up the complete system and update coordinated image pins during maintenance.
Before an upgrade
Read the new upstream compatibility requirements and release notes. Keep the Sentry-family releases coordinated rather than upgrading services independently. Current versions and digests are in images.lock.json.
The current revision verifies restart and redeploy persistence. It does not claim a fresh backup-and-restore drill. Complete your own restore drill before relying on backups for production recovery. See the recovery procedure.
Coordinated backup
- Block SDK ingestion at Gateway and stop telemetry generators.
- Drain Sentry and Snuba consumers. Record consumer lag, a known error event ID, log marker and trace ID.
- Stop scheduler, taskworker and consumer groups after queues drain.
- Export PostgreSQL and preserve roles as needed. Keep exports encrypted and out of public repositories.
- Stop stateful services and take consistent volume backups for PostgreSQL, ClickHouse, Kafka, Valkey and Taskbroker. Record backup identifiers, time, image digests and configuration checksum.
- Copy every Nodestore and Filestore object with its metadata to protected backup storage. Preserve
relay/public.jsonand the generatedRELAY_KEY_SEEDsecurely. Railway has no native bucket snapshot/versioning in the recorded procedure. - Resume the services in dependency order. Query the saved error, log and trace before reopening ingestion.
A PostgreSQL-only backup misses raw event bodies, attachments, queryable telemetry and pending queues. Persistence alone does not prove recoverability.
Deploy coordinated pins
During maintenance, update compatible Sentry-family images and associated configuration together. Start stores, Relay, Snuba, Web, consumers/tasks and Gateway. Verify administrator login, cleanup, saved telemetry and fresh SDK error, log and trace data.
Consumer deployments deliberately omit Railway deployment healthchecks to avoid Kafka partition-sharing deadlock during replacement. Internal heartbeat supervision remains enabled. Do not add deployment healthchecks without testing overlapping consumers.
Rollback or fix forward
After a schema migration, reverting images does not reverse the schema. Stop ingestion and restore all compatible stores and objects from the same backup point before redeploying previous pins. Verify both saved and fresh telemetry before reopening traffic.
A compatible forward fix may preserve more data than restoring an older snapshot. Choose it only when compatibility is supported by evidence. Never delete production volumes or buckets to force an upgrade through.
Restore drill
Restore to an empty isolated environment with the same pins and compatible configuration. Restore volumes and buckets while applications are stopped. Preserve Relay identity and all object metadata. Update references to the restored buckets, then start services in dependency order.
Query the exact saved event ID, log marker and trace ID. Inspect their details and send fresh examples. A restorable disk alone does not prove the application can recover its data.
Source: full recovery procedure.