Design PatternsData ManagementObservability

Eventual Consistency in DDD: Measure the Window with Sagas

OL
Oscar van der Leij
••13 min read
Eventual Consistency in DDD: Measure the Window with Sagas

My most memorable lesson in eventual consistency came from a business process management application. Two people could work on the same data, one in our application and the other in a business partner's. One would make an update, the other would not see it yet, and they would happily override each other's changes. We found out the usual way, through a user complaint. The fix was optimistic concurrency: every record carried a version, and a write based on a stale version got rejected instead of silently winning.

What stuck with me was that nobody had ever agreed how long two parts of that landscape were allowed to disagree. It was implicit, as it has been on every project since, and the teams I worked with found drift the way most teams do: a reconciliation job that tells you in the morning what went wrong yesterday.

Every landscape of bounded contexts has an inconsistency window. Most have one nobody chose. This article argues for making that window a business number, and then builds a small .NET 10 booking flow with Wolverine where the window is enforced, compensated and measured.

Harbour Manifests

Picture a container port. Customs, the shipping line, the terminal warehouse and the haulier each keep their own manifest, and each manifest is that authority's truth. When something happens, a stamped copy of the paperwork goes to the next authority. A good terminal stamps the bill of lading in the same moment the crane lifts the container, so paper and steel never disagree about whether it moved. A good clerk recognises a duplicate manifest and files it once. A container sitting in the yard has a clock on it, and when the clock runs out demurrage starts. Once the container leaves through the gate it cannot be recalled, and a recall order sent after it may itself go astray.

Bounded contexts are the authorities, events are the stamped copies, and your inconsistency window is the age of the oldest manifest nobody has reconciled yet.

The analogy breaks in two places, and both matter. The container is one physical object, so when the manifests disagree someone can walk over and look. In software there is nothing to walk over to: the records are the thing. And a port has a harbourmaster, while bounded contexts have nobody in charge unless you build one.

Eventual Consistency: Many Truths, One Customer

In a domain-driven landscape each bounded context owns its datastore, its rules and its truth. Contexts talk through events, so between an event being published and being handled, they see different realities.

Take a hotel booking. One click from a customer touches Booking, Inventory, Payment, Notification, Loyalty and Analytics. Six manifests for one container.

Now let it go wrong. Booking accepted the reservation, Inventory reserved the room, Notification sent the confirmation email, and then Payment failed. Every context is internally correct, and one customer holds an email that says they're booked.

The saga pattern is the textbook answer: local transactions, with compensating transactions that undo earlier steps when a later one fails. Microsoft's saga guidance is honest about the catch: "Compensating transactions might not always succeed, which can leave the system in an inconsistent state." The same guidance names the pivot transaction, the point of no return. In the hotel case the confirmation email is a container through the gate. You cannot unsend it, you can only send an apology after it.

Each context you add brings more events, dependencies, failure scenarios, compensations and monitoring, and they multiply. The challenge grows faster than the number of services.

Eventual consistency remains a design choice worth making: it buys scalability, team autonomy and the freedom to change one context without a release train for the other five. The useful question is how to make it visible, manageable and acceptable for the business process.

Getting Your Hands Dirty

We'll build Booking, Inventory and Payment in one .NET 10 project, each context as its own namespace and handlers, with PostgreSQL for durable state and RabbitMQ between Booking and Payment. In production these are separate services with separate databases. Here they share a process so you can follow along with the .NET 10 SDK and Docker.

I chose orchestration for this flow, a saga acting as the harbourmaster, because the point is to have one place that holds the clock. For many business process management flows I think orchestration is overkill; for resource provisioning it is really useful.

Step 1: Bring Up the Port

Create an ASP.NET Core web project called HarbourBooking and add the packages:

dotnet new web -n HarbourBooking -f net10.0
cd HarbourBooking
dotnet add package WolverineFx --version 6.46.0
dotnet add package WolverineFx.RuntimeCompilation --version 6.46.0
dotnet add package WolverineFx.Postgresql --version 6.46.0
dotnet add package WolverineFx.RabbitMQ --version 6.46.0
dotnet add package OpenTelemetry.Extensions.Hosting --version 1.19.1
dotnet add package OpenTelemetry.Exporter.Console --version 1.19.1

Now save this as compose.yaml inside the HarbourBooking folder. The health checks matter, because RabbitMQ takes noticeably longer to accept connections than Postgres.

services:
  postgres:
    image: postgres:17
    environment:
      POSTGRES_PASSWORD: postgres   # local demo only
    ports: ["5432:5432"]
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U postgres"]
      interval: 2s
      retries: 30
  rabbitmq:
    image: rabbitmq:4-management
    ports: ["5672:5672", "15672:15672"]
    healthcheck:
      test: ["CMD", "rabbitmq-diagnostics", "-q", "ping"]
      interval: 5s
      retries: 30

Start both containers with docker compose up -d --wait, which blocks until both health checks pass. You should see Container harbourbooking-postgres-1 Healthy and the same for rabbitmq-1.

Step 2: Stamp the Paperwork

Every message carries OccurredAt, the moment the fact happened, because that is what you will measure against later. Add Messages.cs:

using Wolverine;

namespace HarbourBooking;

public record BookingRequested(Guid Id, decimal Amount, DateTimeOffset OccurredAt);
public record HoldRoom(Guid Id);
public record RoomHeld(Guid Id);
public record RequestPayment(Guid Id, decimal Amount, DateTimeOffset OccurredAt);
public record PaymentSucceeded(Guid Id);
public record PaymentFailed(Guid Id, string Reason);
public record ReleaseRoom(Guid Id);
public record SendConfirmation(Guid Id);

// The yard clock: delivered 30 seconds after the saga schedules it.
public record BookingTimeout(Guid Id) : TimeoutMessage(TimeSpan.FromSeconds(30));

Then replace Program.cs. The endpoint calls nothing downstream. It writes BookingRequested to a durable queue in Postgres and returns 202 Accepted. From there every handler's state changes and outgoing messages are persisted together, which is the transactional outbox: the bill of lading stamped as the crane lifts. In a real Booking service you would save your own entity through EF Core in that same transaction; here the saga row doubles as the booking record.

using HarbourBooking;
using JasperFx.Resources;
using OpenTelemetry.Metrics;
using Wolverine;
using Wolverine.Postgresql;
using Wolverine.RabbitMQ;
using Wolverine.RDBMS;

var builder = WebApplication.CreateBuilder(args);
const string pg = "Host=localhost;Port=5432;Database=postgres;Username=postgres;Password=postgres";

builder.Host.UseWolverine(opts =>
{
    opts.PersistMessagesWithPostgresql(pg, "wolverine"); // inbox, outbox, timeouts
    opts.AddSagaType<BookingSaga>();                     // saga state as a Postgres table
    opts.UseRabbitMq(rabbit => rabbit.HostName = "localhost").AutoProvision(); // default guest account
    opts.PublishMessage<RequestPayment>().ToRabbitQueue("payments");
    opts.ListenToRabbitQueue("payments");
    opts.Policies.UseDurableOutboxOnAllSendingEndpoints();
    opts.Policies.UseDurableLocalQueues();
});
builder.Services.AddResourceSetupOnStartup(); // creates tables and queues on first run

var app = builder.Build();

app.MapPost("/bookings", async (decimal amount, IMessageBus bus) =>
{
    var id = Guid.NewGuid();
    await bus.PublishAsync(new BookingRequested(id, amount, DateTimeOffset.UtcNow));
    return Results.Accepted($"/bookings/{id}", new { id });
});

app.Run();

Step 3: Give the Harbourmaster a Clock

The saga moves from Requested to RoomHeld to Paid to Confirmed, and starts the timeout the moment it begins. Note where SendConfirmation lives: only after PaymentSucceeded. Payment is the pivot, so nobody tells the customer anything before it. Add Booking/BookingSaga.cs:

using Wolverine;

namespace HarbourBooking;

public class BookingSaga : Saga
{
    public Guid Id { get; set; }
    public decimal Amount { get; set; }
    public DateTimeOffset RequestedAt { get; set; }
    public string State { get; set; } = "Requested";

    public static (BookingSaga, HoldRoom, BookingTimeout) Start(BookingRequested e) =>
        (new BookingSaga { Id = e.Id, Amount = e.Amount, RequestedAt = e.OccurredAt },
         new HoldRoom(e.Id), new BookingTimeout(e.Id));

    public RequestPayment Handle(RoomHeld e)
    {
        State = "RoomHeld";
        return new RequestPayment(Id, Amount, DateTimeOffset.UtcNow);
    }

    // The pivot. Only past this point may the customer hear "you're booked".
    public SendConfirmation Handle(PaymentSucceeded e) { End("confirmed"); return new SendConfirmation(Id); }

    public ReleaseRoom Handle(PaymentFailed e) { End("released"); return new ReleaseRoom(Id); }

    public ReleaseRoom Handle(BookingTimeout t) { End("timed_out"); return new ReleaseRoom(Id); }

    // Messages arriving after the saga closed: this is where reconciliation begins.
    public static void NotFound(PaymentSucceeded e, ILogger<BookingSaga> log) =>
        log.LogWarning("Payment for {Id} landed after the room was released: refund needed", e.Id);

    public static void NotFound(BookingTimeout t) { } // finished in time, nothing to do

    void End(string outcome)
    {
        State = outcome;
        MarkCompleted();
    }
}

Step 4: Make Payment Idempotent, and Let It Fail

The outbox guarantees at-least-once delivery, which means a relay may publish the same message twice. The clerk has to recognise the duplicate. A static dictionary stands in for a payments table with a unique key on the booking id. Any amount above 1000 is declined, so you can force the failure path. Add Contexts.cs:

using System.Collections.Concurrent;

namespace HarbourBooking.Inventory
{
    public class InventoryHandler
    {
        public static RoomHeld Handle(HoldRoom cmd, ILogger<InventoryHandler> log)
        {
            log.LogInformation("Inventory: room held for {Id}", cmd.Id);
            return new RoomHeld(cmd.Id);
        }

        public static void Handle(ReleaseRoom cmd, ILogger<InventoryHandler> log) =>
            log.LogInformation("Inventory: room released for {Id}", cmd.Id);
    }
}

namespace HarbourBooking.Payment
{
    public class PaymentHandler
    {
        static readonly ConcurrentDictionary<Guid, bool> Processed = new();

        public static object? Handle(RequestPayment cmd, ILogger<PaymentHandler> log)
        {
            if (!Processed.TryAdd(cmd.Id, true))
            {
                log.LogInformation("Payment: duplicate request for {Id} ignored", cmd.Id);
                return null;
            }
            if (cmd.Amount > 1000) return new PaymentFailed(cmd.Id, "card declined");
            log.LogInformation("Payment: charged {Id}", cmd.Id);
            return new PaymentSucceeded(cmd.Id);
        }
    }
}

namespace HarbourBooking.Notification
{
    public class NotificationHandler
    {
        public static void Handle(SendConfirmation cmd, ILogger<NotificationHandler> log) =>
            log.LogInformation("Notification: confirmation sent for {Id}", cmd.Id);
    }
}

Step 5: Measure the Window

Here is the part most landscapes skip. Record how long each booking spent between BookingRequested and a terminal state, tagged with the outcome, plus how old a message is when Payment consumes it. Add BookingMetrics.cs:

using System.Diagnostics.Metrics;

namespace HarbourBooking;

public static class BookingMetrics
{
    public const string MeterName = "HarbourBooking";
    static readonly Meter Meter = new(MeterName);

    static readonly Histogram<double> Window = Meter.CreateHistogram<double>(
        "booking.consistency_window", "s", "BookingRequested until confirmed or released");
    static readonly Histogram<double> EventAge = Meter.CreateHistogram<double>(
        "booking.event_age", "s", "Message age when consumed");

    public static void RecordWindow(DateTimeOffset since, string outcome) =>
        Window.Record((DateTimeOffset.UtcNow - since).TotalSeconds,
            KeyValuePair.Create<string, object?>("outcome", outcome));

    public static void RecordAge(DateTimeOffset occurredAt, string message) =>
        EventAge.Record((DateTimeOffset.UtcNow - occurredAt).TotalSeconds,
            KeyValuePair.Create<string, object?>("message", message));
}

Wire it up with three small edits: call BookingMetrics.RecordWindow(RequestedAt, outcome); at the top of End() in the saga, call BookingMetrics.RecordAge(cmd.OccurredAt, nameof(RequestPayment)); as the first line of the Payment handler, and register the exporter in Program.cs before builder.Build():

builder.Services.AddOpenTelemetry().WithMetrics(m => m
    .AddMeter(BookingMetrics.MeterName)
    .AddConsoleExporter((_, reader) =>
        reader.PeriodicExportingMetricReaderOptions.ExportIntervalMilliseconds = 10_000));

Now the window has a number, and the number needs an owner. A sentence like "a booking is either confirmed or released within 30 seconds" is the kind of thing you agree with the business owner and alert on. That 30 is an example threshold for this demo, the real one comes out of that conversation. In production you would swap the console exporter for Prometheus or OTLP. For traces, OpenTelemetry's messaging conventions connect producer and consumer through span links, so a long-running saga shows up as linked traces rather than one tidy waterfall.

Step 6: Break It on Purpose

Start the app and send one booking that succeeds and one that Payment declines.

dotnet run --urls http://localhost:5080
# in a second terminal:
curl -X POST "http://localhost:5080/bookings?amount=200"
curl -X POST "http://localhost:5080/bookings?amount=5000"

Each curl returns {"id":"<guid>"} immediately. In the app's log you should see the first booking held, charged and confirmed, and the second held then released:

Inventory: room held for <id-1>
Inventory: room held for <id-2>
Payment: charged <id-1>
Notification: confirmation sent for <id-1>
Inventory: room released for <id-2>

Now take a context away. Run docker compose stop rabbitmq, send another booking with amount=200, and wait. The room is held, the payment request sits in the outbox, and the log fills with Error trying to start a new Rabbit MQ channel, each with a full stack trace, every few seconds for as long as RabbitMQ is down. Shortly after the 30 seconds pass, the timeout fires and logs Inventory: room released. It is a little later than 30, because scheduled messages are picked up by polling. That one line scrolls past between the stack traces, so search the terminal output for released rather than waiting for the errors to stop. They won't, and that is the point: the sender keeps retrying while the saga has already moved on. Run docker compose start rabbitmq. Once Wolverine reconnects, the outbox delivers the payment request it was holding, Payment charges the card, and the saga is already gone: you get Payment for <id> landed after the room was released: refund needed. That warning is a compensation that did not quite compensate, exactly what the saga guidance warned about, and it is the reconciliation job's inbox.

Every ten seconds the console exporter prints booking.consistency_window with a histogram per outcome (confirmed, released, timed_out), and the Max of booking.event_age jumps to show how long the delayed payment request sat unreconciled. Your timings will differ from run to run, which is the point of measuring them.

Stop the app and tear it all down with docker compose down -v.

Who Owns the Inconsistency Window

The code is the easy half. The window only means something once three groups have answered their part, and each answer turns into something concrete.

Audience Question What the answer becomes
Business How long may inconsistency last? The SLO on booking.consistency_window and its alert
Business Which processes are critical? Which flows get a saga, a pivot and a timeout
Business When is compensation acceptable? Compensation handlers versus blocking the customer until the pivot
Architects Which contexts may decide on stale data? Where a context acts on its own copy and where it must wait for an event
Architects How do we detect inconsistent states? NotFound handlers, reconciliation jobs, staleness metrics
Dev teams How do we make event processing observable? OccurredAt on every message, event-age histograms, linked traces
Dev teams How do we test cross-domain scenarios? Forced-failure tests like Step 6, run on purpose instead of in production

Notice how few of those questions are technical. The business ones decide the numbers everything else enforces.

Eventual Consistency as a Business Number

  • The window exists whether you choose it or not. An implicit window gets discovered through a user complaint.
  • Put the pivot before the promise. Nothing tells the customer "you're booked" until the step that cannot be undone has succeeded.
  • Assume duplicates. At-least-once delivery means every consumer is a clerk who files each manifest once.
  • Expect compensation to fail sometimes. Design the place where late arrivals land, and give it an owner.
  • Measure from OccurredAt. The age of the oldest unreconciled fact is the number the business actually cares about.

As a DDD landscape grows, the challenge shifts from storing data to managing several truths that coexist for a while. That is a business and architecture question, and a port with no harbourmaster and no clock on the yard is just a pile of containers with opinions.

How long are your contexts allowed to disagree today, and who signed off on that number? And most importantly, would you find out it was exceeded from a metric, or from a customer?

Share this article

Enjoyed this article?

Subscribe to get more insights delivered to your inbox monthly

No spam, unsubscribe anytime.

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.