Repeat a Request Safely

Separate retry from business identity and implement concurrent, transactional command acceptance with recorded results.

A caller submits A-1001, waits, and sees a timeout. The order may have committed before the response was lost. Retrying can be the right action, but only if the service recognizes the request as the same business operation.

Two separate mechanisms answer that, and they are easy to confuse. A bounded retry repeats an operation after failure. Business idempotency records which command was accepted and returns that recorded result when the same command comes back — one cannot substitute for the other.

Checkpoint 18 changes the teaching database from in-memory H2 to a file under $MULE_HOME/orders-data/orders-recovery, because a claim about surviving a restart needs something that survives a restart. Use the isolated runtime. POST /lab/setup intentionally erases this checkpoint’s synthetic order/request records — it is a reset command, not a startup migration for a business service.

Repeated command: Tenant + order key: Stable fingerprint. Local transaction: Unique request claim: Order + response. Result: New: 201; repeat: 200: Different content: 409. The constraint resolves competing claims; recovery reads the winner.

Open this checkpoint in ACB

Stop the previous local application, then use File → Open Folder to open book/checkpoints/18-retry-and-idempotency. Open src/main/mule/app.xml and use Flow List to select the named flow for each example. Start the controlled dependency in a separate terminal from the companion root with python3 book/stubs/server.py. Choose Run and Debug → Run Mule Application and wait for deployment. Save canvas edits and use Save and Hot-deploy to Local Runtime before repeating a request.

Run python3 book/run.py verify 18 from the companion root to exercise this running checkpoint with its synthetic fixtures. The verifier supplies requests and checks results; it does not start the ACB application. Keep the editor on this checkpoint while reading a failure so an old deployment cannot supply a misleading answer.

First, count attempts

A dependency that fails once and succeeds on the next call is the easy half of this problem, and Until Successful covers it. The question worth answering is how many times the far end is actually called.

Example 046 — Retry a temporarily unavailable read

The complete source is in checkpoints/18-retry-and-idempotency/src/main/mule/app.xml.

Choose Flow List → retry-read. Select Until Successful: Max Retries is 2, and Millis Between Retries is 25. Expand it and select its single Request: connection Dependency_HTTP, Method GET, Path /flaky.

Keep the scope limited to that read while testing retries. A processor added inside the scope becomes part of every attempt.

Configure the controlled service with POST /control and {"failNext":2}, then call GET /lab/retry. The first two responses fail and the third succeeds. Max Retries 2 permits two retries after the initial attempt, so this scenario makes three calls. Millis Between Retries 25 keeps the small test quick — it is not a recommended production delay.

Until Successful retries its whole scope on failure, not only the operation that failed. Suppose the scope held a successful database insert followed by a failing HTTP call: the next attempt starts again at the insert, because the scope rewinds the Mule execution and not the committed row. Put two operations inside one retry scope only if replaying both is safe.

Each attempt also begins with the scope’s incoming variable state, so changes from a failed attempt are not a durable attempt counter — use the controlled service’s call record to establish how often it was invoked. If attempts exhaust, handle MULE:RETRY_EXHAUSTED at a boundary that can express the final failure.

Reconnection has a different purpose: it reestablishes connector connectivity according to the connector strategy. A successful reconnection does not decide whether replaying a business write is safe. Nor does repetition make every failure go away — a connection reset may recover, while an unknown SKU will not become valid through repetition. Nested retry mechanisms can also multiply attempts, so count the complete path, including the caller and any broker redelivery.

Until Successful retry count and delay around a single Request

Until Successful retry count and delay around a single Request.

Name the command before storing it

The command key is (tenant_id, orderId). Two retailers can each use A-1001 without owning the same order. For now the local service supplies the fixture tenant retailer-a; chapter 23 will replace that assumption with an explicit caller context.

The RAML response contract now declares 200 for a replay as well as 201 for a new acceptance. The request ledger stores this key, a fingerprint and the serialized accepted response. The fingerprint serializes the normalized command fields in a fixed object order. Item order remains significant in this contract. Rearranging JSON object keys does not create a new command; rearranging items does. Decide such equivalences explicitly before calling something a duplicate.

The primary key on (tenant_id, request_key) arbitrates concurrent claims. A preceding Select is useful for the ordinary repeat path, but two callers can both observe absence — putting a lookup next to an insert on the canvas does not make the pair one atomic claim. The unique constraint, inside the transaction, resolves that race.

Example 047 — Accept or replay a command

The complete source is in checkpoints/18-retry-and-idempotency/src/main/mule/app.xml.

Choose Flow List → accept-once. Its first Set Variable retains command with Value payload. The next creates fingerprint with this Value expression:

write({orderId: payload.orderId,
       items: payload.items map (item) -> {sku: item.sku, qty: item.qty}},
      'application/json', {indent: false})

Select Select, whose Target Variable is previous. It queries the recorded request result by tenant and command identity:

SELECT fingerprint AS "fingerprint", response AS "response"
FROM request_results WHERE tenant_id = :tenant AND request_key = :key

Input Parameters is {tenant: vars.tenant, key: vars.command.orderId}. Expand the following Choice. Its condition not isEmpty(vars.previous) selects the replay route. That route checks the recovered fingerprint, reads the stored response into result, and sets httpStatus to 200.

In Otherwise, Flow Reference calls price-order, Set Variable retains order, and Try calls commit-order. Only after that call succeeds does the normal route assign result = vars.order and httpStatus = 201. The Try’s On Error Continue matches DB:QUERY_EXECUTION and calls recover-recorded-result. The final Transform Message serializes vars.result as JSON.

Use the component cards to distinguish the outer recovery Try from the transactional Try inside commit-order. They do not own the same work.

An existing row is compared with the new fingerprint. A match returns its recorded response with 200. A mismatch raises APP:CONFLICT, mapped to 409 by the API error handler. A new request is priced, then passed to the transaction. Only a successful new commit returns 201.

The catalogue call happens before the transaction begins. It may therefore run twice for concurrent first attempts, but only one result can become authoritative for the key — losing callers read the committed winner, and do not return their own independently calculated price.

Example 048 — Commit the order and request result

The complete source is in checkpoints/18-retry-and-idempotency/src/main/mule/app.xml.

Choose Flow List → commit-order and select Try. Its Transactional Action is ALWAYS_BEGIN and Transaction Type is LOCAL. Both Insert components inside it use Orders_DB and ALWAYS_JOIN.

InsertTarget VariableSQL destinationInput Parameters expression
FirstclaimResultrequest_results(tenant_id,request_key,fingerprint,response){tenant: vars.tenant, key: vars.command.orderId, fingerprint: vars.fingerprint, response: write(vars.order,"application/json")}
SecondinsertResultaccepted_orders(tenant_id,order_id,response){tenant: vars.tenant, id: vars.order.orderId, response: write(vars.order,"application/json")}

The first SQL statement uses VALUES (:tenant,:key,:fingerprint,:response); the second uses VALUES (:tenant,:id,:response). Inspect the bound SQL and input expressions in their respective component panels.

After the inserts, Choice tests vars.failAfterOrder default false and raises APP:INJECTED when true. The Try’s Error handler contains On Error Propagate with Type ANY. Leave the failure injector inside the transaction while testing rollback.

Both inserts join the local transaction. The injected failure comes after the order insert but before commit. Propagate fails the transaction owner so both records roll back. A successful claim without its order, or an order without a retrievable recorded result, would leave the repeated-request contract incomplete.

Example 049 — Recover the committed winner

The complete source is in checkpoints/18-retry-and-idempotency/src/main/mule/app.xml.

Choose Flow List → recover-recorded-result. Its Select repeats the tenant/key query from Example 047 and stores the result in previous. Inspect the following Choice routes in order:

ConditionRaise Error type
isEmpty(vars.previous)APP:PERSISTENCE
(vars.previous[0].fingerprint as String) != vars.fingerprintAPP:CONFLICT

After Choice, Set Variable assigns result with read(vars.previous[0].response as String, "application/json"). The next Set Variable assigns numeric 200 to httpStatus. This subflow runs after the failed transactional call has unwound; do not move its read inside that failed transaction.

The outer handler invokes this subflow after a database query failure. It reads outside the failed transaction, compares the fingerprint and returns a winner only if one exists. An unrelated database failure with no recorded result raises APP:PERSISTENCE — the application must not turn every SQL error into a successful duplicate.

Observe the competing outcomes

The verification command submits a new command, repeats it, changes a quantity under the same key and reads back the order. It also submits four concurrent identical commands and requires exactly one 201 and three 200 responses with equal bodies. That establishes the tested local race, not arbitrary distributed database isolation.

POST /lab/accept-failure runs the transaction with the failure injection enabled. A subsequent GET must return 404. A normal retry with that identity must then succeed with 201. An error-shaped HTTP response alone cannot establish rollback — the independent read is what establishes it.

Do not reset the database before testing restart durability. Leave a committed order, stop and restart the same isolated runtime, then retrieve and repeat it. Replacing or deleting the storage directory is a different failure than restarting the same process over retained files. Cloud replica replacement requires an external persistence design.

Retention is part of the promise

Deleting a request result while the caller can still repeat its key reopens the operation. Choose the idempotency retention window with the business and caller contract, and archive enough evidence to reconcile older requests. A cache TTL chosen for memory use is not an adequate payment or order-retention rule.

This implementation accepts one version of an order. Supporting amendments needs an explicit version or operation identity, concurrency rule and price decision. Reusing A-1001 with changed content is deliberately a conflict here.

Try it

1. Count attempts before changing a retry. What happens if a caller retries three times and each request permits two internal retries?

Show answer

There can be four caller attempts, each making up to three dependency attempts: twelve in the fully failing case. A retry budget must include every layer.

2. Lose the response after commit. Which stored value should the repeat return if the catalogue price has since changed?

Show answer

The response recorded for that accepted command. Repricing the repeat would change the result of the same operation and could disagree with the stored order.

3. Race two different commands. Two requests use one tenant/key but different quantities. What protects the identity, and what should the loser receive?

Show answer

The database unique constraint resolves the claim. After the losing transaction rolls back, recovery compares the winner fingerprint and returns conflict for different content. A Select-before-Insert check alone cannot arbitrate the race.

The transaction now records a reliable acceptance decision. Sending that decision to another system introduces a second durable boundary, which the next chapter makes explicit.

Next: Publish What the Database Committed

Comments