- Purpose
- Survive consumer downtime without losing events.
- Prerequisites
- Delivery history access
- Architecture
- Failed deliveries are retried with exponential backoff, then parked.
Configuration
| Schedule | 1m, 5m, 30m, 2h, 6h (illustrative) |
Implementation steps
- 01Monitor failure rate
- 02Replay parked deliveries after recovery
- 03Reconcile with a REST sweep for anything beyond the retry horizon
Testing procedure
- Return 500 for 10 minutes and confirm recovery without data loss
Troubleshooting
Duplicate processing after replay
Store processed event ids for at least the retry horizon.