Demonstrate Acceptance and Recovery

Run the assembled order service through repeats, races, rollback, ownership and interrupted publication without introducing new mechanisms.

A customer submits A-1001 and the screen spins. The database may have committed the order; the response may have been lost. The customer presses Submit again. This is where an integration stops being a sequence of green components and becomes a promise — one intended order, a recoverable outcome and an answer the caller can trust. A-1001 is accepted only when the service can explain what committed and recover the work that remains, and the complete local implementation lets us test that claim now rather than adding a missing database protocol in the last chapter.

Checkpoint 30 combines the principal seam and ownership query from chapter 23, the transaction and recorded result from chapter 18, the outbox and dispatcher from chapter 19, and the honest transition log from chapter 26. Its core order behavior is the same implementation those lessons introduced.

Submit and repeat: One committed order: Conflict on changed key. Interrupt work: Before commit: After receiver effect. Inspect and recover: No partial transaction: Two sends; one effect. The capstone combines previously taught mechanisms.

Rehearse recovery from ACB

Open book/checkpoints/30-recovery-capstone as the active ACB project. Start the controlled dependency with python3 book/stubs/server.py from the companion root in a separate terminal, then use Run and Debug → Run Mule Application. Wait for deployment before sending requests. Stop any previous checkpoint first because the local listener port is shared.

Keep Flow List → accept-once and the dispatcher available while running the acceptance commands below. Use Debug Mule Application for one synthetic request when you need to inspect a variable at a boundary; do not interpret a paused debugger as a timeout or throughput test. Run python3 book/run.py verify 30 from the companion root for the repeatable scenario checks, and preserve failures before using a diagnostic reset.

Capstone Choice injects a failure inside the local transaction

In commit-order, the Choice on vars.failAfterOrder sits inside the owning Try, after the order Insert and before the outbox Insert. POST /lab/accept-failure from chapter 18 sets that flag; ordinary orders leave it false.

Prepare one isolated run

A recovery test is only as good as the state it starts from, so isolate the run first. Start the dependency/receiver service and the isolated Mule runtime. Run checkpoint 30 from ACB, wait for deployment, then run python3 book/run.py verify 30. The verification resets synthetic teaching tables and receiver rows. Do not point it at a shared environment or use that reset when performing the separate restart-persistence exercise.

The fixture is the server-priced command from chapter 14. A caller supplies order ID and product quantities; the catalogue determines prices. The recorded response contains tenant, priced lines, total 32 USD and state ACCEPTED. It does not claim shipment or payment.

Read the acceptance matrix

ScenarioRequired local observationBoundary established
New valid command201 and recorded orderLocal durable acceptance
Identical repeat200 and identical recorded bodyRequest identity and replay
Changed command under same key409Fingerprint conflict
Four simultaneous identical commandsOne 201, three 200; equal bodiesTested unique-key race
Failure before transaction commit500, then absent orderLocal rollback
Retry after that rollbackNew 201No abandoned successful claim
Other tenant reads A-1001404Application tenant ownership
Missing principal401 from local guardRequired application context

Each row names a particular observable result, and the observation is what counts: a handler can return any text it likes regardless of what the database did, so the evidence for the rollback row is the absence of the order afterward — not the handler’s reassuring response. A backend timeout belongs to the same discipline. It should stay an operational failure rather than being recast as invalid customer data, because the status a caller receives has to say whether correction, retry or support is the next step. The local principal is synthetic, so its tests do not establish JWT verification. The H2 transaction is real, but it does not establish PostgreSQL isolation or a remote warehouse rollback.

Read the publication matrix

Writing the order and then publishing directly would leave a gap, because the database can commit while publication fails; publishing first opens the opposite gap, where fulfillment learns about an order that never committed. The outbox from chapter 19 closes both by committing the intent with the order, and this matrix is where that claim gets tested. After reset, accepting one order must leave one request result, one order and one pending event. The receiver is then configured to commit the next event and return 503 — the ambiguous case the design exists for. Dispatch reports uncertainty, the sender retains pending intent and the receiver has one effect.

A second dispatch sends the same event. The receiver recognizes it, the sender marks it sent, and the receiver still has one effect after two attempts. Dispatching an empty outbox reports zero attempted records. This is the local recovery promise: a repeated delivery is visible, and its stable identity prevents a second row in the controlled receiver.

That receiver row is the entire demonstrated downstream effect. A real warehouse must offer a corresponding idempotency or reconciliation contract around its own effect, because an outbox is a durable publication plan — it does not create exactly-once execution across every system. Writing a marker before an unrelated irreversible operation would reopen the same atomicity gap on the receiving side.

Restart without erasing the evidence

A process that comes back with its files is a different experiment from a replacement that comes back without them, and only one of them is being run here. Leave a committed order and an unsent event. Stop the isolated runtime and restart it over the same orders-data directory. Retrieve the order and repeat the command without running /lab/setup. Then dispatch pending work and inspect receiver state.

This exercise distinguishes process restart with retained storage from a fresh deployment with a new filesystem. The former can preserve this teaching database; the latter needs the external persistence design from chapter 25 — so record which failure you actually introduced before calling the result durable.

The SQLite receiver also owns a file-backed ledger. If the receiver and sender both lose their files, the local experiment has lost both sources of evidence. Backups, retention and reconciliation across storage loss require a larger operating design than this process-restart check.

Carry the tests to a different environment

Replacing H2 with PostgreSQL, HTTP publication with MQ, or a local principal with JWT enforcement changes a tested boundary — keep the same business cases, but run them again through the new adapter and account for its failure model. Do not replace the labels on local output and call it a cloud result.

For a platform acceptance run, add direct-listener bypass, policy cold start, wrong audience, expired token, broker redelivery, acknowledgement loss, replica replacement and configuration-aware rollback. Retain artifact, configuration and policy identities with each result. The reference projects make those tasks concrete while preserving the distinction between supplied source and executed behavior. A release is ready when its remaining uncertainty is explicit and acceptable to the people who own it.

Try it

1. Trace one identity across recovery. Which IDs remain stable when the dispatcher retries an event?

Show answer

Tenant/order identify the accepted business object and eventId identifies its notification. A new execution can have a different correlation ID. The retry must retain event identity so the receiver recognizes the same effect.

2. Move the failure. Why are before-commit failure and after-receiver-commit failure different tests?

Show answer

The first should leave no transaction records. The second deliberately leaves a receiver effect and uncertain sender acknowledgement, requiring repeat or reconciliation. One rollback test cannot establish both boundaries.

3. State the completed promise. What can this local capstone claim, and what still needs target-specific evidence?

Show answer

It demonstrates the supplied local acceptance, conflict, race, rollback, ownership and repeated-publication behaviors on its pinned stack. Provider authentication, PostgreSQL, MQ, cloud replacement and gateway enforcement require their own acceptance runs.

The service has grown from a greeting into an acceptance protocol whose uncertain outcomes can be examined. When the customer presses Submit a second time, the useful answer comes from a durable record of the intended operation rather than a guess about which network hop failed. The habit is the same at every scale: identify the next consumer, make the state transition explicit, and inspect the result at the boundary where it becomes authoritative.

Next: Appendix A: Setup, Versions and Evidence

Comments