Designing resilient .NET 8 background services
BackgroundService in .NET 8 makes it easy to spin up long-running workers,
but the default “while (true) with a delay” isn’t enough once you hit production scale.
Here’s how I usually structure background workers so they behave like good citizens.
1. Respect cancellation and shutdown
The runtime gives you a CancellationToken for a reason. Your worker should exit
quickly when the app is shutting down — no hanging 60 seconds while you “just finish this
batch”.
Key rules I follow:
- Check
token.IsCancellationRequestedbetween iterations. - Avoid long blocking calls without timeouts.
- Flush in-flight work, then exit.
2. Use a small “orchestrator” loop
Instead of putting everything in ExecuteAsync, I keep the loop thin and delegate
the real work into a separate service. It keeps the worker testable and easier to reason about.
// Pseudo-structure, not full code
protected override async Task ExecuteAsync(CancellationToken stoppingToken)
{
while (!stoppingToken.IsCancellationRequested)
{
await _processor.ProcessBatchAsync(stoppingToken);
await Task.Delay(_pollInterval, stoppingToken);
}
}
The _processor hides all the complexity: retrieving messages, saving to SQL,
emitting metrics,
etc. This separation also lets you reuse the same processor in integration tests.
3. Handle transient failures with retries (but cap them)
When you talk to external systems (queues, APIs, databases), failures are normal, not exceptional. I wrap those calls in a retry policy (Polly or a simple custom backoff).
- Retry a few times with exponential backoff.
- Emit metrics/logs after each failure, not just the last one.
- Give up and move on instead of blocking the entire worker forever.
4. Guard the loop from “poison” data
A classic issue: one bad message or one corrupt record throws an exception every loop, and your worker keeps crashing and restarting.
To avoid this:
- Catch exceptions per item, not only at the outer loop.
- Send poison items to a quarantine table/queue.
- Keep the main loop alive unless the process truly can’t continue.
5. Add basic observability from day one
Background services are invisible to users, so they’re easy to forget until they fail. I always add at least:
- A counter for processed items (success + failures).
- A gauge or log for queue depth / pending items.
- A health check endpoint that verifies the worker is “doing work recently”.
6. Summary
Resilient workers are less about clever code and more about boring discipline: handle shutdown, isolate the “work”, treat failures as normal, and make them observable.