Chaos Engineering & Resilience Testing: SDET Interview Questions 2026
Master chaos engineering and resilience testing for SDET interviews: failure injection, circuit breakers, bulkheads, steady-state hypothesis, and Game Days. Mitchell Agoma's 20-year perspective from government, defence, financial services, and national retailers.
Published 14 June 2026 • By Mitchell Agoma
It is 11:42pm. You have spent the last three evenings revising your test automation framework answers, memorising the differences between stubs and mocks, and practising your CI/CD pipeline explanation until you can recite it in your sleep. You feel ready — or ready enough — for the lead SDET role at a fintech company that handles millions of transactions per day. Then, as you scroll through the interview confirmation email one final time, you spot it: "The technical interview will include a 30-minute session on resilience testing and chaos engineering principles." Your stomach tightens. You open Google. You type "chaos engineering interview questions SDET." You find an article about Netflix's Chaos Monkey from 2017, a GitHub repository with a Kubernetes pod killer script, and a Medium post titled "What Is Chaos Engineering?" that reads like a conference talk abstract. Nothing that tells you what an interviewer will ask. Nothing that tells you how to describe failure injection testing without sounding like you have just read the Wikipedia page. Nothing that gives you the language to say "I have tested systems under controlled failure — here is how I did it, what I learned, and why it matters" — which is exactly what a lead SDET interviewer expects to hear. Chaos engineering is the fastest-growing senior SDET interview topic that almost nobody prepares for — because most SDETs think it is a DevOps discipline, not a testing discipline. They are wrong. And the candidates who can explain steady-state hypothesis design, describe how to automate a circuit-breaker resilience test, and articulate the difference between fault injection and chaos engineering walk into the interview visibly ahead of every other applicant.
Mitchell Agoma has spent 20 years in test engineering — across government departments, the defence sector, financial services, and national retailers — and at every organisation that operated at scale, the question was not whether systems would fail but how gracefully they would fail and how quickly they would recover. At a financial services organisation, Mitchell's team was responsible for testing a payment-processing system that handled £2.3 million in transactions per hour. The architecture had circuit breakers, retry policies, fallback mechanisms, and rate limiters — and every single one of them was untested. The team ran 1,200 integration tests every night. Not one of them injected a network delay. Not one of them killed a database connection pool. Not one of them verified that when the downstream fraud-checking service timed out, the payment was queued for retry rather than silently dropped. When the fraud-checking service went down during a Black Friday peak — a scenario that the architecture was designed to handle — the system did not handle it. It cascaded. The payment service timed out. The circuit breaker did not open because it had never been tested under load. The retry policy flooded the already-overloaded fraud service. The queue backed up. The outage cost an estimated £180,000 in lost transactions over 47 minutes — all because the resilience mechanisms worked on the whiteboard but not in production. The lesson Mitchell's team took from that day: resilience mechanisms that have never been tested are not resilience mechanisms — they are assumptions. And assumptions fail in production, spectacularly, at the worst possible moment.
At a national retailer (Asda), the e-commerce platform used bulkhead patterns to isolate the checkout service from the recommendation engine — so a surge in browsing traffic would not take down the ability to buy. The architecture diagram was beautiful. The implementation was untested. When a marketing campaign drove 10x normal traffic to the product recommendation service, the bulkhead did not hold — because the connection pool configuration had been copied from a staging environment with 1/50th the capacity. Mitchell's team built a chaos experiment that simulated the exact traffic pattern: ramp recommendation requests to 10x, verify the checkout service remained responsive, and assert that the bulkhead's thread pool rejected excess requests rather than exhausting shared resources. The experiment failed on the first run. It took three iterations — adjusting thread pool sizes, connection timeouts, and queue depths — before the bulkhead held. The experiment was then automated into the deployment pipeline. Every release since has verified bulkhead resilience before reaching production. The thread that connects Mitchell's experience across government, defence, financial services, and retail: resilience testing is not a nice-to-have for systems that handle money, personal data, or critical infrastructure. It is the testing discipline that answers the question your users will ask during an outage — "why didn't you test for this?" — before they have to ask it.
The SDET Interview Coach iOS app — with 800+ questions across 32 topics, Claude-graded mock interviews from Junior to Lead, and dedicated modules on resilience testing, chaos engineering, and non-functional testing strategies — gives you the structured practice to describe failure injection, circuit breaker testing, and Game Day scenarios with the specificity of someone who has done it, for £4.99 per month. Do not let a 30-minute resilience testing session be the reason you hear "we have decided to move forward with another candidate."
What Interviewers Are Actually Testing When They Ask About Chaos Engineering — It Is Never "Do You Know What Chaos Monkey Is?"
When an interviewer asks "have you done chaos engineering?", they are not checking whether you can name Netflix's tools. A candidate who has read the SRE book can list Chaos Monkey, Chaos Kong, Gremlin, and Litmus — that is a vocabulary test, and vocabulary tests do not get lead SDET offers. What the interviewer is actually testing is whether you understand that chaos engineering is a structured experimental discipline — not random destruction. The core of chaos engineering is the scientific method applied to distributed systems: you form a hypothesis about how the system will behave under a specific failure condition (the steady-state hypothesis), you inject that failure in a controlled experiment, you measure the system's behaviour against your hypothesis, and you either confirm the system is resilient or discover a weakness to fix. The measure of a senior SDET in this domain is not whether they have run chaos-monkey on a staging cluster. It is whether they can design a chaos experiment, define a meaningful steady-state hypothesis, automate the experiment into a pipeline, and articulate why testing resilience is fundamentally different from testing functionality. That is what the interviewer wants to hear — not "I know what Chaos Monkey is" but "I understand that resilience is a system property that must be tested, not assumed, and here is how I design experiments to test it."
Signal 1: You Understand the Scientific Method Behind Chaos Engineering
The strongest candidates immediately frame chaos engineering as experimental science, not reckless destruction. "Chaos engineering follows a four-step experimental cycle. Step 1 — Define the steady-state hypothesis: 'When the payment-authorisation service experiences 500ms of added latency on 30% of requests, 99.5% of checkout transactions will complete within 3 seconds without error.' The hypothesis is specific (which service, what failure, what blast radius), measurable (99.5%, 3 seconds), and falsifiable (the experiment can prove it wrong). Step 2 — Design the experiment: select the failure injection method (network delay via Toxiproxy, latency injection via Gremlin, or an application-level toggle), define the blast radius (30% of traffic, not 100%), and establish the abort conditions (if error rate exceeds 5%, stop the experiment immediately). Step 3 — Execute in a controlled environment: run the experiment in a staging environment that mirrors production topology, with monitoring dashboards visible, and with a rollback plan ready. Step 4 — Analyse the results: compare actual behaviour to the steady-state hypothesis. If the hypothesis holds, confidence increases. If it does not, a weakness has been discovered — and discovering weaknesses in a controlled experiment is infinitely better than discovering them during a production outage." This framing demonstrates that you approach chaos engineering as a disciplined testing practice, not as a tool you run occasionally. At a government department, Mitchell's team applied this four-step method to a citizen-facing portal with 12 microservices — running one chaos experiment per week, each targeting a different failure mode — and discovered 7 resilience weaknesses over 3 months, all of which were fixed before they could cause production incidents.
Signal 2: You Can Automate Resilience Testing Into CI/CD — Not Just Run It Ad Hoc
Junior candidates describe chaos experiments they ran once in a workshop. Senior candidates describe chaos experiments that run automatically on every deployment. "I integrate resilience testing into the deployment pipeline so that every release is verified against a baseline set of chaos experiments. The approach: maintain a suite of automated chaos tests — each a scripted experiment that injects a specific failure, monitors the system behaviour, and asserts against the steady-state hypothesis. In the CI/CD pipeline, after the deployment to staging, the chaos test suite runs: (1) inject 500ms latency on the payment service and verify checkout latency remains under 3 seconds for 99.5% of requests, (2) kill one instance of the inventory service and verify the load balancer routes traffic to healthy instances within 5 seconds, (3) saturate the database connection pool to 95% and verify connection timeouts are handled gracefully rather than crashing the application, (4) simulate a downstream service returning 503 errors for 60 seconds and verify the circuit breaker opens and the fallback path activates. Each chaos test produces a pass/fail result and a report of resilience metrics — recovery time, error rate during failure, throughput degradation. If any chaos test fails, the deployment is blocked — the same as a failing functional test. This turns resilience from an aspiration into a gate." At a financial services organisation, Mitchell's team integrated 12 automated chaos experiments into their staging pipeline. In the first month, 3 deployments were blocked by failing chaos tests — deployments that would have passed functional tests and reached production with latent resilience weaknesses. The cost of fixing those weaknesses in staging was measured in hours. The cost of discovering them in production would have been measured in thousands of pounds per minute.
Signal 3: You Understand the Hierarchy of Resilience Patterns — and How to Test Each One
The most technically impressive answer Mitchell hears in resilience-testing interviews is not about tools — it is about testing the resilience patterns themselves. "I test resilience at four levels, corresponding to the patterns that protect the system. Level 1 — Timeouts: does every outbound call have a timeout, and does the system handle a timeout gracefully (not by crashing, not by hanging indefinitely)? I test this by making a downstream service unresponsive — accepting the TCP connection but never responding — and verifying the caller times out within the configured duration and returns a controlled error or fallback response. Level 2 — Retries: does the retry policy work — with exponential backoff, jitter, and a maximum retry count — and does it avoid retry storms? I test this by making the downstream service fail for exactly N attempts and recover on attempt N+1, then verifying the caller retries the correct number of times with the correct backoff pattern and succeeds when the downstream recovers. I also test for the retry storm scenario: verify that when 100 concurrent requests all retry simultaneously, the backoff jitter distributes the retries across time rather than synchronising them. Level 3 — Circuit Breakers: does the circuit breaker open when the failure threshold is exceeded, transition to half-open after the recovery timeout, and close when the downstream recovers? I test this by failing a downstream service for 10 requests (exceeding the threshold), verifying the circuit opens (requests are rejected immediately without calling the downstream), waiting for the recovery timeout, sending a probe request, and verifying the circuit transitions through half-open to closed when the downstream responds successfully. Level 4 — Bulkheads: does the bulkhead isolate failures so that one exhausted resource pool does not cascade? I test this by saturating the thread pool for Service A — sending requests until every thread is busy — and verifying that Service B, which shares the same host but a different bulkhead, remains fully responsive. This layered approach proves that the resilience architecture works under failure, not just on the whiteboard." At BT, Mitchell's team applied this layered testing to a telecommunications provisioning system where a single misconfigured timeout had historically caused cascading failures across 4 downstream services. After implementing the four-level resilience test suite, the team discovered that the retry policy had no jitter — meaning 150 concurrent retries fired simultaneously — and that the circuit breaker's half-open probe was calling a deprecated endpoint. Both weaknesses were invisible to functional tests and would have surfaced in production under load.
Signal 4: You Know That Game Days Are Not Chaos Experiments — and You Can Explain Why
Many candidates conflate Game Days with chaos experiments. Senior candidates understand the distinction — and it is this distinction that impresses interview panels. "A chaos experiment is a single, controlled failure injection with a specific hypothesis — 'what happens when the payment service has 500ms latency on 30% of traffic?' A Game Day is a scheduled, cross-team exercise where multiple failure scenarios are injected simultaneously or sequentially to test the organisation's incident response — not just the system's resilience. The difference is scope: a chaos experiment tests the system; a Game Day tests the system, the monitoring, the runbooks, the escalation paths, and the humans. In a chaos experiment, the primary measurement is whether the system meets its SLOs during failure. In a Game Day, the primary measurements are time-to-detect (how long until the on-call engineer notices the failure), time-to-diagnose (how long until they identify the root cause), time-to-mitigate (how long until the system is recovering), and time-to-resolve (how long until normal operation is restored). The SDET's role in a Game Day is to design the failure scenarios, instrument the system for observability, and capture the data that drives the retrospective. I design Game Day scenarios that are realistic (inspired by past incidents), safe (with abort conditions and rollback plans), and educational (the goal is learning, not blaming)." At a government department, Mitchell facilitated quarterly Game Days where the team simulated a production incident — database failover, network partition, third-party API outage — and measured their response. The first Game Day revealed that the monitoring dashboard showed CPU utilisation but not error rates, so the team spent 14 minutes diagnosing a problem that should have been visible in 30 seconds. The monitoring was fixed before the next Game Day, and the time-to-diagnose dropped from 14 minutes to 2 minutes. The system itself had not changed — but the team's ability to operate it under failure had improved dramatically.
The one-sentence answer that anchors every strong chaos engineering interview response: "Resilience is a system property that must be tested, not assumed — and the measure of a senior SDET is not whether they can run a chaos tool, but whether they can design controlled experiments, automate resilience verification into CI/CD, and prove that the system degrades gracefully under failure rather than collapsing catastrophically."
The 7 Most Common Chaos Engineering and Resilience Testing Interview Questions — With Model Answers That Demonstrate Production Experience
Here are the questions that Mitchell has both asked in interviews and been asked — each with the model answer that distinguishes a candidate who has tested resilience from a candidate who has only read about Netflix's Chaos Monkey.
Q1: "What is chaos engineering — and how is it different from failure testing or negative testing?"
What the interviewer is testing: This is the foundational question. Can you define chaos engineering precisely and distinguish it from related disciplines? A candidate who says "it is breaking things on purpose" fails. A candidate who describes the scientific method and the distinction between known-unknowns and unknown-unknowns succeeds. Model answer: "Chaos engineering is the discipline of experimenting on a distributed system to build confidence in its ability to withstand turbulent conditions in production. It is distinct from failure testing and negative testing in three ways. First, scope: failure testing asks 'if I inject this specific failure, does the system handle it correctly?' — it tests known failure modes against known handling logic. Chaos engineering asks 'what happens when I introduce turbulence into the system?' — it explores unknown failure modes and emergent behaviour. Failure testing verifies the system handles failures it was designed for. Chaos engineering discovers failures the system was not designed for. Second, method: failure testing is assertion-based: inject failure, assert the error handler runs. Chaos engineering is hypothesis-driven: define a steady-state, inject turbulence, measure deviation from steady-state. Third, environment: failure testing can run in any environment because it tests code paths. Chaos engineering must run in an environment that approximates production because it tests system behaviour under realistic load, network topology, and resource constraints. The relationship is complementary: failure testing catches bugs in resilience code; chaos engineering catches gaps in resilience architecture. For example, a failure test verifies that the circuit breaker opens after 5 consecutive timeouts. A chaos experiment verifies that when the circuit breaker opens, the remaining services can handle the redirected traffic without themselves becoming overloaded — an emergent behaviour that no unit test or integration test can predict." For related guidance on systematic testing approaches, see our test case design techniques guide.
Q2: "How do you define a steady-state hypothesis — and what makes a good one?"
What the interviewer is testing: The steady-state hypothesis is the intellectual core of chaos engineering. The interviewer wants to know whether you can define a hypothesis that is specific, measurable, and falsifiable — or whether you will propose something vague like 'the system should still work.' Model answer: "A steady-state hypothesis is a statement of normal system behaviour — the baseline against which the chaos experiment's results are compared. A good hypothesis has four properties. Property 1 — Specific: it identifies the system, the metric, the threshold, and the duration. Bad: 'the system should handle failures.' Good: 'when 50% of inventory-service instances are terminated, the product-detail page continues to return HTTP 200 with inventory data for 99% of requests within 2 seconds, measured over a 5-minute window.' Property 2 — Measurable: the metric must be observable from outside the system — ideally from the same monitoring dashboards used in production (Prometheus, Datadog, Grafana). If you cannot measure it, you cannot verify it. Property 3 — Falsifiable: the hypothesis must be capable of being proven wrong. 'The system will not crash' is not falsifiable because the system could slow to a crawl without crashing. '99% of requests complete within 2 seconds' is falsifiable — if only 94% complete within 2 seconds, the hypothesis is disproved. Property 4 — Business-relevant: the metric should reflect the user experience or business outcome, not just technical health. 'CPU utilisation stays below 80%' is less meaningful than 'checkout completion rate stays above 99%.' The steady-state hypothesis is what transforms chaos engineering from 'let us break things and see what happens' into a rigorous experimental discipline. It is also what allows you to automate the experiment: if you can measure the steady-state programmatically (via an API query to your monitoring system), you can assert on it in an automated test. See our monitoring and observability guide for the instrumentation needed to make steady-state measurements reliable."
Q3: "Walk me through how you would test a circuit breaker — step by step."
What the interviewer is testing: Circuit breakers are the most commonly implemented resilience pattern — and the most commonly untested. The interviewer wants to see whether you can design a test that covers all three circuit states (closed, open, half-open) and the transitions between them. Model answer: "I test a circuit breaker through its full lifecycle across three states and five scenarios. Scenario 1 — Normal operation (closed state): the downstream service is healthy. Send N requests, verify all pass through the circuit breaker to the downstream, verify the success counter increments, and verify the circuit remains closed. Scenario 2 — Failure threshold exceeded (closed → open): configure the downstream to fail (HTTP 500) for M consecutive requests, where M equals the circuit breaker's failure threshold. Send M failing requests, then send one more request — and verify this request is rejected immediately by the circuit breaker without calling the downstream. The key assertion: the request fails fast (typically in under 5ms) rather than timing out after the downstream's configured timeout (typically 5-30 seconds). This fast-fail behaviour is the circuit breaker's primary value — it prevents resource exhaustion from waiting for a known-broken downstream. Scenario 3 — Recovery timeout (open → half-open): after the circuit opens, wait for the configured recovery timeout (e.g., 30 seconds). Then fix the downstream service (restore healthy responses). Send a probe request — this is the first request in the half-open state. Verify that (a) the request passes through to the downstream (the circuit allowed it), (b) the downstream returns a success response, and (c) the circuit transitions to closed. Scenario 4 — Half-open failure (half-open → open): similar to Scenario 3, but the downstream still fails. Verify that the single probe request fails and the circuit immediately returns to the open state — it does not require the full failure threshold again. Scenario 5 — Fallback verification: while the circuit is open, verify that the fallback behaviour works correctly. If the fallback is a cached response, verify the cache is returned. If the fallback is a degraded experience, verify the degradation is applied. If the fallback is an error message, verify the error is meaningful to the caller. The test should be automated — I use Toxiproxy (an open-source TCP proxy that simulates network conditions) to programmatically control the downstream's behaviour, or Spring Cloud Contract / WireMock to stub HTTP responses at the application layer. The test code looks something like:
@Test
void shouldOpenCircuitBreakerAfterFailureThresholdExceeded() {
// Arrange: make downstream fail consistently
toxiproxy.toxics()
.latency("downstream-latency", ToxicDirection.DOWNSTREAM, 5000);
// Act: send requests until threshold exceeded
for (int i = 0; i < circuitBreakerThreshold; i++) {
assertThrows(ServiceUnavailableException.class,
() -> paymentClient.authorise(testPayment()));
}
// Assert: circuit is now open — requests fail fast
long start = System.currentTimeMillis();
assertThrows(CircuitBreakerOpenException.class,
() -> paymentClient.authorise(testPayment()));
long duration = System.currentTimeMillis() - start;
assertThat(duration).isLessThan(50); // fast-fail, not timeout
}
The tooling is detailed in our Docker test automation guide — Toxiproxy and similar tools run in Docker containers, making them ideal for CI/CD pipeline integration.
Q4: "What is the difference between fault injection and chaos engineering — and when would you use each?"
What the interviewer is testing: This question separates practitioners from theorists. Fault injection and chaos engineering are often used interchangeably, but they serve different purposes. Model answer: "Fault injection is a technique — the deliberate introduction of faults into a system to test its error-handling behaviour. Chaos engineering is a discipline — the systematic experimentation on a system to discover its resilience weaknesses. Fault injection asks a specific question: 'does the system correctly handle a network timeout on this specific call?' Chaos engineering asks an open-ended question: 'what weaknesses exist in this system's resilience?' The practical difference is in scope and intent. Fault injection is targeted and verification-oriented: I inject a specific fault (a 5-second TCP connection delay) into a specific component (the payment service's connection to the fraud service) and verify a specific behaviour (the payment service times out after 3 seconds and queues the payment for retry). This is a functional test of resilience code — I know exactly what I am testing and what the expected outcome is. Chaos engineering is exploratory and discovery-oriented: I introduce turbulence into the system (terminate 30% of instances, add latency to inter-service calls, saturate a database connection pool) and observe what happens — without a complete expectation of the outcome. The goal is not to verify that the system handles a known failure; it is to discover unknown failure modes. In practice, I use both — fault injection for automated resilience verification in CI/CD (every deployment must pass the fault-injection test suite), and chaos engineering for periodic exploration in staging or production (discover weaknesses that the fault-injection tests do not cover). Fault injection is the safety net. Chaos engineering is the exploration that expands the safety net. At a financial services organisation, Mitchell's team ran fault injection tests on every deployment (testing known failure modes) and weekly chaos experiments in a staging environment (exploring unknown failure modes). The fault injection tests caught regressions. The chaos experiments caught the weaknesses the team had not thought to test for — including a database failover that succeeded but caused a 47-second read-replica lag that broke a reporting service. No fault injection test had been written for read-replica lag because nobody had considered it a risk. The chaos experiment discovered it in 20 minutes." For the broader context of non-functional verification, see our non-functional testing guide.
Q5: "How do you decide what chaos experiments to run — and what do you prioritise?"
What the interviewer is testing: This is the strategy question. You cannot run chaos experiments against every component and every failure mode — the combinatorial space is infinite. The interviewer wants to see that you can prioritise based on risk. Model answer: "I prioritise chaos experiments using a risk-based framework with three dimensions. Dimension 1 — Impact radius: if this component fails, how many downstream services and user journeys are affected? A payment-authorisation service that is called by every checkout transaction has a larger impact radius than a product-recommendation service that is called only on the homepage. I prioritise high-impact-radius components first. Dimension 2 — Failure history: has this component or this failure mode caused an incident in the past? I review the incident post-mortems from the last 12 months and design chaos experiments for every failure mode that caused a production incident. If a database connection-pool exhaustion caused an outage in March, there must be a chaos experiment that reproduces connection-pool exhaustion and verifies the fix. Dimension 3 — Architectural dependency graph: I map the service dependency graph and identify the critical paths — the minimum set of services that must be operational for the core business function to work. For an e-commerce platform, the critical path is: browser → CDN → web application → checkout service → payment service → payment gateway. Every service on this path gets a chaos experiment before any service off the path. The prioritisation output: I produce a chaos experiment backlog — a prioritised list of experiments, each with a hypothesis, a blast radius, an owner, and a target date. The top of the backlog is reserved for components that score high on all three dimensions: high impact radius, history of incidents, and on the critical path. The bottom of the backlog is for low-impact, no-incident-history, off-critical-path components — these get experiments when time permits but do not block releases. The framework prevents the most common chaos engineering mistake: starting with interesting but low-risk experiments (like killing a non-critical background worker) while the critical-path services remain untested under failure." For more on risk-based testing strategy, see our test strategy planning guide.
Q6: "What tools do you use for chaos engineering and resilience testing — and how do you choose between them?"
What the interviewer is testing: Can you name specific tools, explain their trade-offs, and demonstrate that you choose tools based on requirements rather than brand recognition? Model answer: "I categorise chaos engineering tools into three tiers based on maturity and blast-radius control. Tier 1 — Application-level proxies (lowest blast radius, safest for automation): Toxiproxy, WireMock, MockServer. These sit between the application and its dependencies, intercepting network calls and injecting failures at the protocol level — latency, connection resets, bandwidth throttling, HTTP error responses. Toxiproxy is my default choice for automated resilience tests in CI/CD because it is deterministic, programmable via a REST API, and runs in Docker. A test script can add a toxic (latency, timeout, bandwidth limit) before the test and remove it after — the blast radius is limited to the single service under test. Tier 2 — Infrastructure-level agents (medium blast radius, for staging environments): Gremlin, Chaos Mesh (Kubernetes), LitmusChaos. These operate at the infrastructure level — terminating pods, saturating CPU, filling disk, dropping network packets. Gremlin provides fine-grained blast-radius control (target specific containers, hosts, or percentages of traffic) and an abort mechanism (halt all experiments with a single command). I use these in staging environments where the blast radius is contained and the experiments test infrastructure resilience (does the pod autoscaler respond to a terminated instance within 30 seconds?) rather than just application resilience. Tier 3 — Production-safe experimentation platforms: AWS Fault Injection Simulator, Azure Chaos Studio, Gremlin with production safeguards. These are designed for running chaos experiments in production with strict safety controls — automatic experiment duration limits, automatic rollback on alert triggers, and audit trails. I use these only after Tier 1 and Tier 2 experiments have passed, and only for experiments that cannot be realistically simulated in staging (like testing the behaviour of a managed service — RDS failover, ElastiCache node replacement — which behaves differently in AWS's production environment than in a local simulation). The decision framework: start with Tier 1 (Toxiproxy) for CI/CD automation — every resilience pattern gets a Tier 1 test before merge. Progress to Tier 2 (Chaos Mesh/Gremlin) for staging-environment Game Days. Reserve Tier 3 (AWS FIS) for production experiments that are approved, time-limited, and monitored. No team should jump straight to Tier 3 without mastering Tier 1 and Tier 2 — the risk of an uncontrolled production experiment causing an outage is too high." For setting up the infrastructure to support these tools, see our Kubernetes test infrastructure guide.
Q7: "How do you measure the success of a chaos engineering programme — and how do you report it to stakeholders who think you are just breaking things?"
What the interviewer is testing: This is the leadership question. Chaos engineering can look like expensive vandalism to stakeholders who do not understand it. The interviewer wants to see that you can justify the investment in business terms. Model answer: "I measure chaos engineering success in four metrics that map directly to business value. Metric 1 — Weaknesses discovered and fixed: the number of resilience weaknesses identified by chaos experiments that were fixed before causing production incidents. Each weakness fixed is a potential incident prevented. Over a 6-month period at a financial services organisation, Mitchell's team ran 18 chaos experiments, discovered 7 weaknesses, fixed all 7, and prevented an estimated 3 major incidents — each of which would have caused 30-90 minutes of degraded service. Metric 2 — Mean time to recovery (MTTR) improvement: measured during Game Days, not during real incidents. Before the chaos engineering programme, the team's Game Day MTTR was 47 minutes. After 6 months of quarterly Game Days with targeted improvements to monitoring, runbooks, and automation, the MTTR dropped to 11 minutes. This is a 77% improvement in the team's ability to recover from failure — and it was measured without waiting for a real incident. Metric 3 — Resilience coverage: the percentage of critical-path services that have automated resilience tests. Starting from 0% (no resilience tests, only functional tests), the goal is to reach 100% coverage of critical-path services within 12 months. Each service added to the resilience test suite reduces the risk of an unhandled failure. Metric 4 — Deployment confidence: the number of deployments blocked by failing resilience tests — evidence that the tests are working as a safety net. Each blocked deployment is a potential production incident prevented. I report these metrics quarterly to stakeholders using language they understand: 'In Q1, our chaos engineering programme discovered 3 resilience weaknesses, all fixed before reaching production. We estimate these would have caused approximately 2 hours of degraded service, affecting an estimated 15,000 customers. The programme cost £12,000 in engineering time. The estimated cost of those 2 hours of degraded service — in lost transactions, support tickets, and engineering time for incident response — was approximately £65,000. Return on investment: 5.4x.' Numbers speak louder than principles. When stakeholders see the ROI, 'breaking things on purpose' becomes 'investing in reliability.' For the foundational metrics framework, see our QA metrics and KPIs guide."
Common Mistakes in Chaos Engineering — and How to Avoid Them in an Interview
The five most damaging mistakes Mitchell has observed across organisations — each one a trap that interviewers specifically listen for.
Mistake 1: Starting Without a Steady-State Hypothesis
Running chaos experiments without a hypothesis is like running functional tests without expected results — you cannot tell whether the outcome is good or bad. The fix: every experiment begins with a written hypothesis that defines the steady-state behaviour, the failure to inject, the metric to measure, and the threshold for success. Without this, the experiment is a learning opportunity only — and learning opportunities are useful, but they do not gate deployments.
Mistake 2: Injecting Failures Without Blast-Radius Control
Killing 50% of instances in production without an abort condition is not chaos engineering — it is a self-inflicted outage. The fix: every experiment has a defined blast radius (which services, what percentage of traffic, which users), an automated abort condition (if error rate exceeds X%, stop the experiment), and a rollback plan. In an interview, always specify the blast radius when describing a chaos experiment — it demonstrates operational maturity.
Mistake 3: Testing Only Infrastructure Failures, Not Application Failures
Killing a pod tests the orchestrator's self-healing — valuable, but it does not test whether the application correctly handles a slow downstream, a malformed response, or a dependency returning stale data. The fix: chaos experiments must include application-level failures — injected latency, HTTP error codes, malformed response bodies, and resource exhaustion (connection pool, thread pool, memory). Infrastructure resilience and application resilience are different concerns; testing one does not verify the other.
Mistake 4: Running Chaos Experiments Only in Staging
Staging environments lack production traffic patterns, production data volumes, and production infrastructure behaviour (managed services behave differently under load than local simulations). The fix: after establishing confidence in staging, graduate the most critical chaos experiments to production — with strict blast-radius controls and during low-traffic periods. Production is the only environment that reveals how the system truly behaves under failure at scale.
Mistake 5: Treating Chaos Engineering as a Project, Not a Practice
Running a single chaos experiment, fixing the discovered weakness, and declaring "we do chaos engineering" is the most common failure mode. The fix: chaos engineering is an ongoing practice, like testing itself. The system changes with every deployment — new services, new dependencies, new failure modes. A chaos experiment that passed last month may fail this month because a configuration change reduced a timeout from 5 seconds to 2 seconds. The chaos experiment suite must run continuously — in CI/CD on every deployment, and as scheduled experiments in production — to catch resilience regressions.
How to Practise Chaos Engineering Before Your Interview — A 2-Day Preparation Plan
Chaos engineering cannot be learned from reading alone. You must inject failures into a real system, observe the behaviour, and articulate what you learned. Here is a two-day plan:
Day 1: Build a Resilient System and Break It
Create a small project — a simple e-commerce checkout service that calls a payment-authorisation service — with these components: (1) two Spring Boot or Express services running in Docker, (2) a circuit breaker (Resilience4j for Java, opossum for Node.js), (3) a retry policy with exponential backoff, (4) Toxiproxy sitting between the services for failure injection, (5) Prometheus metrics for measuring latency and error rates. Write chaos experiments: inject 2-second latency on the payment service and verify the circuit breaker opens; kill the payment service and verify the checkout service returns a controlled fallback; inject HTTP 500 errors on 50% of payment requests and verify the retry policy handles them. By the end of Day 1, you should have a working resilient system with automated chaos experiments that you can explain in an interview. The code should be on GitHub — interviewers will ask to see evidence of practical experience.
Day 2: Run a Personal Game Day and Practise Interview Answers
Schedule 90 minutes. Set up a monitoring dashboard (Grafana with Prometheus). Design three chaos experiments: (1) gradual latency injection — increase latency on the payment service from 100ms to 2 seconds in 200ms increments over 10 minutes and observe at what point the user experience degrades, (2) instance termination — kill one of two payment-service instances and measure recovery time, (3) dependency cascade — inject latency on a database query and observe whether the latency propagates through the service chain. Run the experiments, record the results (what you expected, what happened, what you learned), and write a one-page retrospective. Then practise answering the seven questions from this guide — out loud, to a timer, as if you were in the interview. The combination of hands-on experimentation (Day 1) and deliberate interview practice (Day 2) gives you both the technical experience and the articulation fluency that interview panels look for. Use the SDET Interview Coach iOS app to simulate the resilience testing and chaos engineering module — the AI mock interviewer asks scenario-based questions, follows up on your answers, and scores your responses on technical accuracy, completeness, and communication clarity. The Claude-graded feedback helps you refine your answers until they sound like someone who has been doing this for years — not someone who read a blog post the night before. Available on the iOS App Store for £4.99 per month.
Chaos Engineering Is Not Optional for SDETs Targeting Senior Roles in 2026
Resilience testing and chaos engineering were once the exclusive domain of site reliability engineers at Netflix and Amazon. That era is over. In 2026, any system that handles money, personal data, or critical business operations — which is to say, nearly every system that hires SDETs — is expected to degrade gracefully under failure, recover quickly, and demonstrate that its resilience mechanisms have been tested, not assumed. The patterns described in this guide — steady-state hypothesis design, circuit breaker lifecycle testing, fault injection versus chaos engineering, Game Day facilitation, and resilience coverage metrics — are the patterns that Mitchell's teams have applied in production across government departments, the defence sector, financial services, and national retailers like Asda, Co-op, and BT. They are not academic exercises. They are the difference between a system that absorbs failure and a system that amplifies it. In 2026, chaos engineering and resilience testing are what performance testing was in 2018 and what API testing was in 2015 — a skill that separates the SDETs who are choosing between offers from the SDETs who are waiting for rejection emails. Prepare now, practise deliberately, and walk into your interview with the specific, production-informed answers that make panels say "this candidate has operated systems under failure — not just read about it." Because by 2028, resilience testing will not be a specialisation. It will be part of the baseline SDET job description — and the candidates who prepared early will be the ones leading the interviews, not sitting in them.
For structured, AI-graded interview preparation across chaos engineering, resilience testing, and the full range of SDET topics, the SDET Interview Coach iOS app offers 800+ questions across 32 topics with Claude-powered mock interviews from Junior to Lead. The resilience testing module includes scenario-based questions on circuit breakers, bulkheads, retry policies, and steady-state hypothesis design — with AI-graded feedback on your technical accuracy and communication. Available on the iOS App Store for £4.99 per month.
If you are building your foundational system testing knowledge, start with our SDET system design interview guide — understanding system architecture is the prerequisite for testing system resilience. If you want to understand how resilience patterns integrate with deployment pipelines, see our continuous testing and DevOps guide. And if you are preparing for the non-functional testing dimensions that chaos engineering complements, see our non-functional testing interview guide.
Ready to Transform Your Testing?
The AI Test Automation Playbook gives you everything you need: Playwright setup, Claude AI integration, MCP deep dive, 10+ ready-to-use prompts, CI/CD pipeline setup, and a 30-day implementation roadmap.
By Mitchell Agoma, Senior SDET & AI Testing Specialist with 8+ years of experience