Your API has an error contract. It is written down. It has a problem+json media type, a Retry-After header, and a webhook signature scheme. And almost nobody has watched a real client hit it under load.

The first time a client sees a genuine 429 with a long Retry-After, or a response body that stops mid-JSON, is usually during an incident. That is a bad time to discover that the retry logic ignores the header, the parser throws an unhandled exception, or the webhook handler compares signatures with ==.

A game day fixes the timing. It is a scheduled, controlled exercise where you inject the failure modes your contract already promises, then watch how your own clients, docs, and runbooks behave. The point is not to prove you can survive chaos. The point is to find the gap between what you published and what you actually built.

Why a game day is a contract test, not a stress test

Stress tests ask whether the system stays up. Game days ask whether the contract stays true. Those are different questions, and the second one is usually neglected.

RFC 9457 defines a problem detail format for carrying machine-readable error details in HTTP responses, specifically so APIs do not need to invent their own error format. It notes that problem details most naturally fit 4xx and 5xx responses, though they can be used with any status code. If you have adopted that format, you have a concrete shape to assert against.

RFC 9110, the HTTP semantics specification, is equally useful here. It says recipients should parse received protocol elements defensively, with only marginal expectations that they conform to grammar or fit a reasonable buffer size. That is a direct instruction to your clients: do not assume the body is well-formed. A game day is how you check whether they listened.

The exercise is narrow by design. You pick faults that map to clauses you already publish. If your contract does not say what happens on a 429, injecting one teaches you less than fixing the contract first.

Pick faults that correspond to written contract clauses

Start with the failures your documentation already describes. If a fault is not covered by a contract clause, either the contract is incomplete or the fault is out of scope for this exercise.

429 with Retry-After

RFC 9457 says a problem type definition may specify the use of the Retry-After response header in appropriate circumstances. It does not mandate a retry policy, a backoff algorithm, or a specific rate-limit header set. That means the behavior is yours to define and yours to test.

Inject a 429 with a Retry-After value that is long enough to be inconvenient. Watch whether clients honor it or hammer the endpoint. Watch whether the error body is a valid problem detail or a bare string.

HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 120

{
  "type": "https://api.example.com/problems/rate-limit",
  "title": "Rate limit exceeded",
  "status": 429,
  "detail": "You have exceeded the request quota for this endpoint.",
  "instance": "/requests/abc123"
}

The key lines to assert on are the Content-Type, the Retry-After value, and the presence of a resolvable type URI. RFC 9457 recommends that a problem type URI resolve to HTML documentation explaining how to resolve the problem. If your type URI 404s, that is a doc bug you can file the same day.

Truncated response bodies

A truncated body is not a valid problem detail. It is a partial JSON object, a cut-off string, or an empty response with a 200 status. RFC 9110’s guidance to parse defensively is the contract clause here: clients should not assume the body is complete or well-formed.

Inject a response that stops mid-object. Watch whether the client throws a parse error, retries, or silently treats the partial data as valid. The last outcome is the dangerous one.

HTTP/1.1 200 OK
Content-Type: application/json
Content-Length: 512

{
  "id": "ord_123",
  "status": "shipped",
  "items": [
    {"sku": "ABC", "qty": 2},
    {"sku": "DEF", "qty": 1}
  ],
  "tracking": "1Z999AA10123456784"
  // connection closes here

The client should detect that the JSON is incomplete. If it does not, you have found a real bug without waiting for a network partition to expose it.

Malformed webhook signatures

Webhook verification is a contract with a specific, testable shape. GitHub’s documentation describes the pattern clearly: GitHub uses your secret token to create a hash signature sent in the X-Hub-Signature-256 header on each delivery. The signature is an HMAC hex digest computed from the secret and the payload contents.

GitHub also recommends comparing the computed and received signatures with a constant-time comparison rather than a plain equality operator. The docs explicitly say never to use a plain == operator, and to consider methods like secure_compare or crypto.timingSafeEqual to mitigate timing attacks.

A webhook replayer that mutates the signature header is a cheap, high-signal test. It exercises the verification path, the error path, and the alerting path in one shot.

POST /webhooks/github HTTP/1.1
Host: api.example.com
X-Hub-Signature-256: sha256=0000000000000000000000000000000000000000000000000000000000000000
Content-Type: application/json

{"action": "opened", "issue": {"number": 42}}

Send a near-miss signature, a missing header, and a signature computed with the wrong secret. Each should produce a distinct, logged rejection. If any of them is accepted, you have a security bug, not a resilience bug.

GitHub provides fixed test values so you can verify your implementation against a known-good result. The secret is It's a Secret to Everybody and the payload is Hello, World!. The expected signature is published in the docs. If your implementation does not produce that signature, the bug is in your code, not in the test.

One more detail from the same source: if a proxy or load balancer modifies the payload or headers before verification, signature verification can fail. That is worth testing explicitly. A game day that routes webhooks through your actual ingress path will catch it.

Build the injection harness

You do not need a chaos platform. You need a proxy or middleware that can rewrite responses and a webhook replayer that can mutate headers.

Response rewriting

A simple reverse proxy in front of your staging API can intercept responses and apply faults based on a header or a query parameter. For example, a request with X-GameDay: truncate gets a body cut at a fixed byte offset. A request with X-GameDay: rate-limit gets a 429 with a Retry-After header.

Keep the fault set small and named. If you cannot list the faults on one page, the exercise is too broad to debrief usefully.

Webhook replay

Capture a real webhook delivery from staging, then replay it with a mutated signature header. The replayer should be able to send:

  • A valid signature (control case)
  • A signature with one hex character changed
  • A missing X-Hub-Signature-256 header
  • A signature computed with a different secret

Each case should produce a logged rejection with enough detail to debug. If the logs say only “invalid signature,” that is a doc bug for the runbook.

What to capture

Before the exercise, decide what evidence you will collect. Client logs, retry counts, support tickets, and the time between fault injection and first alert are all useful. The goal is to compare observed behavior to the contract, not to produce a pass/fail score.

Run the debrief like a blameless postmortem

The most valuable output of a game day is not a pass/fail result. It is a list of doc bugs: places where the contract was ambiguous, the runbook was stale, or the client team’s mental model diverged from the spec.

Ask three questions for each injected fault:

  • Did the client behave the way the contract says it should?
  • Did the on-call runbook tell the responder what to do?
  • Did the documentation explain the fault clearly enough that a new engineer could predict the behavior?

If the answer to any of these is no, file a doc bug. If the answer to the first is no, file a client bug. If the answer to the second is no, update the runbook before the next game day.

Some faults should become permanent CI tests. A 429 with Retry-After is easy to assert in a contract test. A truncated body is harder but still testable with a fixture. A malformed webhook signature is a unit test. The game day is how you discover which ones matter enough to automate.

Scope it to staging with a named abort criterion

This is not a production exercise. The goal is to observe behavior, not to prove you can survive production chaos. Run it in staging with a rollback path and a named abort criterion.

An abort criterion is a specific condition that ends the exercise immediately. Examples: error rate exceeds a threshold, a participant cannot reach the staging environment, or a fault affects a shared dependency outside the scope of the test. Write it down before you start.

If you cannot describe the rollback in one sentence, the exercise is too risky for a first run.

The artifact: a one-page game day runbook

Adapt this template for your team. Keep it to one page. If it grows beyond that, split it into multiple exercises.

# API Game Day Runbook

## Scope
- Environment: staging
- Services: [list]
- Participants: [names and roles]

## Faults
1. 429 with Retry-After: 120
2. Truncated response body at 512 bytes
3. Webhook with mutated X-Hub-Signature-256

## Hypothesis
For each fault, what do we expect the client, on-call, and docs to do?

## Abort criteria
- Error rate exceeds [threshold]
- [Other condition]

## Rollback
- [One sentence]

## Evidence to capture
- Client logs
- Retry counts
- Time to first alert
- Support tickets

## Follow-ups
- Doc bugs filed
- Client bugs filed
- Faults promoted to CI tests

The runbook is the deliverable. The game day is just the excuse to write it.

FAQ

How often should we run a game day?

Often enough that the runbook stays current. If your API changes quarterly, once a quarter is reasonable. If it changes weekly, the runbook will rot faster than you can test it, and you should invest in automated contract tests instead.

What if we do not use RFC 9457?

The exercise still works. The point is to test the contract you actually publish, whatever its shape. RFC 9457 is useful because it gives you a concrete, machine-readable format to assert against, but it is not a prerequisite.

Can we run this in production?

Not for a first run. The value of the exercise is in observing behavior, and production adds risk without adding much signal. Start in staging, and only consider production if you have a mature fault-injection practice and a very specific question that staging cannot answer.

What if the client team is in a different time zone?

Record the exercise and share the evidence. The debrief can be asynchronous. The important thing is that the client team sees the observed behavior and has a chance to file bugs.

How do we know if the game day was successful?

You have a list of doc bugs, client bugs, and CI tests that did not exist before. If you have none of those, either your system is perfect or the exercise was too gentle.