Architecture
How the private ingestion path, query services and persistent stores fit together.
Data flow
Your application → private Gateway → Relay → Kafka
├→ Sentry consumers → event storage
└→ Snuba consumers → ClickHouse
Browser → public Gateway → Sentry Web → Snuba API → ClickHouse
└→ PostgreSQL / Valkey / object buckets
Sentry tasks → Taskbroker / PostgreSQL / scheduled cleanupThe diagram is a simplified view of the template architecture. Service configuration and official image pins are in images.lock.json.
Service responsibilities
| Service | Responsibility |
|---|---|
| gateway | Public authenticated interface and separate private SDK ingress; bootstrap readiness checks |
| web | Sentry interface, API, initialization and relational migrations |
| relay | Validates and forwards SDK envelopes |
| kafka | Buffers telemetry for consumers |
| sentry-consumers | Processes Sentry event pipelines |
| snuba-consumers | Indexes queryable telemetry into ClickHouse |
| snuba-api | Queries ClickHouse for Sentry |
| sentry-tasks | Taskworker, scheduler and cleanup supervision |
| taskbroker | Persistent task queue |
| postgres | Relational application state |
| clickhouse | Queryable event, log and span data |
| valkey | Shared runtime state and queues |
| memcached | Cache |
| symbolicator | Symbolication service |
Persistent storage
PostgreSQL, ClickHouse, Kafka, Valkey and Taskbroker use persistent volumes. Separate private Nodestore and Filestore buckets hold event bodies and attachments or source maps. Nodestore also preserves Relay's public identity; the generated Relay seed must remain stable across redeploys.
A PostgreSQL-only backup is incomplete. Use coordinated backups across volumes and both object buckets, preserving object metadata.
Network and region
Only the Gateway interface is public. The private SDK listener must not receive a public domain or TCP proxy. All application telemetry producers must share the Railway environment.
The native template uses Railway's default service region, with object buckets fixed in AMS. The recorded validation used EU West/AMS. Native S3 uploads use public TLS and incur service egress. Treat a region move as a storage migration, not an in-place configuration edit.
Retention and health
Indexed retention defaults to seven days. Kafka queue retention defaults to three hours. These configuration values are separate policies, not recovery guarantees. Queue segment deletion and consumer lag determine which pending records survive an outage.
Deployment readiness is not continuous monitoring. Observe consumer progress, scheduler activity and cleanup completion after startup. A running process alone does not prove that work is being processed.
Sources: operations, recovery, Web variables, Kafka variables.