The silent retry bug killing ASP.NET Core reliability (and how idempotency keys save you)
The Silent Retry Bug
If you've spent any time working with distributed systems in C# and ASP.NET Core, you know the drill: network calls fail, timeouts happen, and services become temporarily unavailable. The standard advice? Use Polly. Add a retry policy. Make it resilient. But here's the hard truth: retries without idempotency are dangerous. When you retry a GET request, it's usually safe-just reading data. But when you retry a POST or PUT request, you risk changing state. If the first request actually succeeded (but the client didn't receive the response due to a timeout), the retry will execute the logic a second time. In a payment system, that's a double charge. In an e-commerce platform, that's a duplicate order. This "duplicate-state problem" is the quiet killer of reliability.
The Problem with "At Least Once"
Most retry strategies (Polly, Hangfire, Azure Service Bus) operate on an "at least once" delivery guarantee. This means you will retry if something looks like it failed. However, unless your endpoint is idempotent, a successful first call followed by a failed second call (or a successful second call that thinks it's the first) leads to data inconsistency.
The Solution: Idempotency Keys
The fix isn't a more sophisticated retry policy. It's making your POSTs safely repeatable. The industry-standard pattern is the Idempotency Key.
How It Works
Client Side: The client generates a unique identifier (usually a GUID) for the intent of the operation (e.g., "Place Order #1234"). This key is attached to the HTTP request, often in the Idempotency-Key header.
Server Side: The ASP.NET Core endpoint checks if this key has been seen before. If new, it executes the business logic, saves the result, and stores the key-result pair (with a TTL). If existing, it skips the business logic and returns the stored result immediately.
Implementing in ASP.NET Core
While you can build this manually, it's best practice to encapsulate it in a middleware or a filter so every endpoint benefits from it. Here is a conceptual outline of how the server logic changes:
[HttpPost("orders")]
public async Task<IActionResult> CreateOrder(
[FromBody] OrderDto order,
string idempotencyKey)
{
// 1. Check cache/dal for existing result
var existingResult = await _idempotencyStore.GetByKeyAsync(idempotencyKey);
if (existingResult != null)
{
return Ok(existingResult); // Return the original outcome
}
// 2. Execute business logic
var result = await _orderService.ProcessOrderAsync(order);
// 3. Store the key with the result
await _idempotencyStore.StoreAsync(idempotencyKey, result);
return CreatedAtAction(nameof(GetOrder), new { id = result.Id }, result);
}
You'll want to use a fast, distributed cache like Redis for _idempotencyStore to ensure this works across multiple instances of your API.
Integration with Polly
When using Polly, you don't need to change your retry policy drastically. You just need to ensure that every retryable HTTP call includes the idempotency key. For HttpClient in .NET 5+, you can use HttpMessageInvoker or a custom DelegatingHandler to inject the header automatically for specific clients.
protected override async Task<HttpResponseMessage> SendAsync(
HttpRequestMessage request,
CancellationToken cancellationToken)
{
if (request.Headers.TryGetValues("Idempotency-Key", out var values) &&
!values.Any())
{
// Fallback if client didn't provide one (though client SHOULD provide it)
request.Headers.Add("Idempotency-Key", Guid.NewGuid().ToString());
}
return await base.SendAsync(request, cancellationToken);
}
Why This Matters for Background Jobs
If you're using Hangfire or Azure Functions, the same principle applies. If a job fails mid-execution and retries, it must be safe to run again. By ensuring that the underlying HTTP calls within the job are idempotent, you can confidently use automatic retries without corrupting your database.
Practical Takeaway
Resilience isn't just about not crashing. It's about maintaining data integrity under failure. Key actions include:
- Don't assume POSTs are safe. They aren't.
- Implement Idempotency Keys for all state-changing endpoints.
- Use a Distributed Cache to store these keys with a reasonable Time-To-Live (TTL).
- Test your retries. Simulate a timeout after a successful write and verify your system doesn't duplicate state.
By decoupling "retry logic" from "state mutation," you build a system that is truly resilient.
Comments
No comments yet. Start the discussion.