Reliabilityoptimistic lockingetagif-matchversionconflict412

Optimistic Concurrency: Versions and If-Match

Let concurrent writers proceed without locks, but make every update state which version it read. A stale version gets a 409 or 412 instead of silently destroying someone else's write — and the contract must say who untangles the conflict.

▶ Run the labFollow the failure

Frame the contract

API design starts with a consumer, a design question and a guarantee — never with a URL.

Design question
When two clients update the same resource from the same starting point, how does the second one find out — and what is it supposed to do then?
Consumers
Any client that reads, lets a human or a process think, then writes: admin UIs editing records, mobile apps syncing offline changes, background jobs updating rows they fetched minutes ago, and multi-tab users racing themselves.
The promise
An update succeeds only against the state the client actually read. Concurrent modification is surfaced as an explicit, machine-readable conflict with a defined resolution path — never as a silent overwrite.
RequirementConsumersResource ModelStyleContractValidationAuthorizationErrorsIdempotencyPaginationVersioningObservabilityEvolutionTrade-offs

The mechanism: a version travels with the read

Every read carries the resource's current version — an ETag header, a version integer in the body, or both. Every write sends that version back: If-Match: "v7" as a precondition, or "version": 7 in the payload. The server compares atomically at write time: match → apply and bump; mismatch → reject with 412 Precondition Failed (the HTTP-native form) or 409 Conflict (the body-version form). Nothing locks, nothing waits — hence *optimistic*: writers proceed assuming no conflict and pay only when one actually happened.

The comparison must be atomic with the write — a WHERE version = 7 on the UPDATE, or a compare-and-set in the store — not an application-level read-check-write, which just narrows the race window without closing it. Databases give you this primitive cheaply (Concurrency Anomalies covers what happens underneath); the API's job is to surface it as a contract clause rather than absorb the anomaly silently.

Optimistic beats pessimistic (lock on read) for API use almost by forfeit: HTTP clients disappear without unlocking, hold locks across human think-time, and retry — every one of those behaviors poisons lock-based schemes. Pessimistic concurrency survives inside a transaction on one connection; it does not survive being stretched across stateless requests. If a workflow truly needs exclusive access for minutes, model the *claim* explicitly (a checkout/lease resource with an expiry) so the lock is visible, owned and reclaimable, instead of implied.

The stale write is refused, and the refusal carries what the client needs next
Request
PUT /articles/42 HTTP/1.1
If-Match: "v7"
Content-Type: application/json

{ "title": "Q3 Plan", "body": "…edited from v7…" }
Response
HTTP/1.1 412 Precondition Failed
ETag: "v9"
Content-Type: application/json

{
  "error": {
    "code": "version_conflict",
    "message": "Resource changed since your read (v7 → v9).",
    "current_version": "v9",
    "request_id": "req_3ab…"
  }
}

The contract decision: who resolves the conflict

Rejecting the stale write is the easy half. The client now holds changes it cannot save — and the contract must say what happens next, because "show an error" pushed onto an unprepared client becomes "user loses ten minutes of edits", which is barely better than the lost update you prevented.

There are three honest resolution assignments. The user resolves: the client re-fetches, shows a diff or a "this changed while you edited" screen, and the human merges — right for documents and rich edits, and it requires UI investment the API team must warn clients about up front. The client resolves automatically: re-fetch, re-apply the change to the fresh state, retry with the new version — right when the change is a pure function of intent ("set status to approved") rather than of the stale snapshot; this loop is so common the docs should spell it out as the recommended retry recipe. The server resolves: field-level merge of non-overlapping changes — powerful and dangerous, because "non-overlapping" is a semantic judgment (two fields can be logically coupled), so it should be opt-in per field group, not a default.

One structural lever shrinks the conflict surface more than any resolution policy: narrower writes. A full-document PUT conflicts with every concurrent change; a PATCH of two fields, or a command like POST /articles/42/publish, conflicts only with changes it semantically overlaps (PUT vs PATCH, Designing State Transitions). The less state a write claims to know, the fewer 412s anyone has to resolve.

Version check added, resolution unassigned
1PUT /articles/42 If-Match: "v7"
2412 Precondition Failed
3{ "error": "conflict" }
4
5# Client team, on discovering this in production:
6# - no current_version in the error
7# - no guidance: refetch? retry? merge?
8# - ships: on 412 → refetch → resend user's
9# payload with fresh If-Match
10# Result: an auto-overwrite loop. The 412 is now
11# a lost update with extra steps.
The conflict is a documented workflow, not just a status
1PUT /articles/42 If-Match: "v7"
2412 { code: "version_conflict",
3 current_version: "v9",
4 retry: "refetch, reapply intent, resend" }
5
6Docs, per endpoint:
7 status: auto-retry safechange is intent-based
8 ("publish"), reapply and resend
9 body: user-mediatedrefetch, present merge UI
10 tags: server-mergednon-overlapping PATCHes
11 are combined, overlap412

A 412 without a resolution path trains client teams to write blind refetch-and-resend loops, which reintroduce the exact overwrite the version check exists to prevent. Naming the resolution per field group makes the safe loop the easy one.

ETag vs version field, and choosing preconditions

The ETag/If-Match pair is the HTTP-native mechanism: it composes with Conditional Requests: ETags, 304 and 412 caching (If-None-Match reads and If-Match writes share the same token), intermediaries understand it, and 412 has one unambiguous meaning. Its friction is practical: clients must thread a header through layers that mostly handle bodies, and debugging tools show bodies more readily than headers. A version field in the representation is more visible and serializes naturally into client state; it costs you a custom 409 contract that you must document as carefully as HTTP documents 412. Many mature APIs ship both, backed by the same underlying counter.

Whichever token you pick, decide its granularity honestly. A per-resource version is simple and over-conflicts (any change bumps it, so unrelated edits collide). Per-field or per-section versions conflict precisely but multiply bookkeeping. Start per-resource; split only where measured conflict rates on genuinely independent fields justify it.

Also decide whether unconditioned writes remain legal. Accepting a bare PUT without If-Match keeps old clients working — and silently keeps last-write-wins for exactly the clients most likely to cause conflicts. Requiring the precondition (reject bare writes with 428 Precondition Required) is the safe endpoint's posture; the migration between the two is an API Migration: Running the Change End to End with telemetry on who still writes blind.

  • `ETag` + `If-Match` — standard, cache-coherent, proxy-friendly; token lives in headers.
  • Body `version` + `409` — visible, easy to persist in client state; semantics are yours to document.
  • `428 Precondition Required` — the endpoint's way of saying "blind writes are not a thing here".
  • Weak vs strong ETags — concurrency control needs strong ones; a weak ETag (W/"…") says "equivalent", not "identical", and preconditions on it are undefined ground.
  • Version tokens are opaque — clients that parse or arithmetic them (v7 + 1) break the day you switch to content hashes.

Key points

  • Optimistic concurrency = version travels with the read, write states its precondition, server compares atomically, mismatch is an explicit 412/409.
  • The comparison must be compare-and-set at the store (WHERE version = ?), not an application-level check — otherwise the race merely narrows.
  • Rejecting the stale write is half the design; the contract must assign resolution — user merge, client reapply-and-retry, or server field merge.
  • A blind refetch-and-resend loop on 412 reintroduces the lost update; document the safe retry recipe so clients don't invent the unsafe one.
  • Narrower writes (PATCH, commands) shrink the conflict surface more than any resolution policy.
  • Locks don't survive stateless HTTP; if exclusivity is truly needed, model the lease as an explicit, expiring resource.

Lost Update Lab

Change the contract and observe which guarantee moves.

Lost Update Lab
Two clients edit the same document. Read on both, save on both, and see who wins.
Client A
has not read yet
Client B
has not read yet
Server
v1 · "Q3 launch plan"
Exchange log

Without the version check, the second save silently erases the first — a lost update. With it, the stale writer gets 412 and must re-read; the contract turned a data-loss bug into a visible conflict the client can resolve.

Follow the failure

How the contract fails or gets misused, hop by hop — and what it costs when it completes.

  1. 1
    Team → API: adds version to responses and a check in the handler — as a read-then-write in application code.
  2. 2
    Two clients → API: read v7 simultaneously; both checks pass in the race window; the second write silently wins anyway.
  3. 3
    Team → API: fixes the atomicity, ships bare 412 conflict with no current version and no guidance.
  4. 4
    Client team → users: implements refetch-and-resend on 412 to "make the errors go away"; overwrites resume, now invisible to metrics.
  5. 5
    Support → both teams: "my changes disappeared" tickets continue; everyone believes the version check made it impossible.
What breaks
  • Users lose edits — either silently (races, blind-resend loops) or loudly (412s that discard their work because the client had no merge path).
  • Sync-style clients (offline mobile) corrupt data confidently: their queued writes carry stale versions, and mishandled conflicts multiply across the queue.
  • Trust in the mechanism itself: after one bad 412 experience, client teams route around it with force=true-style escapes, and the API ends up with documented last-write-wins.

Design, observe, evolve

A contract decision is incomplete until you know how you would notice it failing and how it changes later.

Design the contract
  • • Implement the version check as compare-and-set in the store; return `412`/`409` with the current version and a machine-readable `version_conflict` code.
  • • Assign conflict resolution per endpoint or field group in the docs — auto-retry-safe, user-mediated, or server-merged — and spell out the safe retry recipe.
  • • Require preconditions on conflict-prone endpoints (`428` for blind writes), migrating existing clients with telemetry rather than by surprise.
  • • Prefer intent-shaped writes (PATCH, commands) over full-document PUT to shrink what each write claims to know.
Observe in production
  • • Track 412/409 rates per endpoint and per client: a spike is contention (or a client with a broken version cache); a rate of exactly zero on a multi-writer resource means blind writes are still getting through.
  • • Log both versions on every conflict (sent vs current) — the gap size distinguishes slow humans from broken clients replaying ancient state.
  • • Watch for refetch-immediately-resend patterns in traces: that is the unsafe auto-overwrite loop announcing itself.
Evolve without breaking
  • • Ship versions in responses first (additive), let clients adopt `If-Match` voluntarily, then require preconditions per endpoint with notice — each step is compatible.
  • • Switching version representation (counter → content hash) survives only if clients treated tokens as opaque; that opacity clause must be in the contract from day one.
  • • Server-side merge can be introduced later as an opt-in per field group without disturbing clients that resolve manually.
What it costs
  • • Conflicts move from silent to visible — which means client teams must build handling they previously didn't know they needed; the API team pays in documentation and advocacy.
  • • High-contention resources pay a retry tax: under real concurrency, some writers loop refetch-reapply several times; a queue or a command log fits those hotspots better.
  • • Requiring preconditions raises the integration floor: quick scripts and one-off tools now need a read before every write.

Misconceptions

Claim
“Conflicts are rare, so optimistic concurrency is over-engineering.”
Reality
Rare per-request, near-certain per-year — and the failure is silent data loss, the worst detectability class. The mechanism costs one header and one WHERE clause; the incident costs a customer's data.
Claim
“A 412 means the client did something wrong.”
Reality
It means the world changed between read and write — normal operation under concurrency. Clients should handle it as a routine workflow (refetch, reapply, retry), not log it as an error and give up.
Claim
“Database transactions already prevent this.”
Reality
A transaction protects the write statement, not the client's read-think-write span. The stale read happened minutes before the transaction began; only a version carried through the API can connect the two. See The Lost Update, Step by Step.

Apply it