Kubernetes for SDET Test Infrastructure Interview Questions 2026 — K8s Architecture Fundamentals for QA Engineers (Pods, Deployments, Services, Ingress, Persistent Volumes), Deploying Selenium Grid on Kubernetes with Production-Grade Helm Charts, Running Playwright at Scale on K8s Clusters (Job Patterns, Sidecar Containers, Parallel Execution Strategies), CI/CD Integration with Argo CD and GitOps for Test Infrastructure, Monitoring Test Pods with Prometheus, Grafana, and Loki, Infrastructure as Code for QA Environments, and the Kubernetes Interview Questions Senior SDET Panels Ask in 2026
The definitive Kubernetes for SDET interview guide for 2026 — covering every K8s concept that infrastructure-aware interview panels drill into at architecture depth. When the hiring manager asks 'walk me through your test infrastructure scaling strategy' and your answer stops at Docker Compose, you've signalled you operate at the single-machine level. The follow-up question — 'and what happens when you need to run 500 browser tests in parallel across three geographies?' — is where the senior SDET conversation begins. In 2026, every engineering organisation running tests at scale is migrating — or has already migrated — test infrastructure to Kubernetes. Interview panels at Google, Amazon, Netflix, and fast-growing scale-ups want to hear that you've deployed Selenium Grid with Helm on a multi-node cluster, not just run minikube locally. They want to hear about your Playwright test Jobs with resource limits and liveness probes, your Argo CD-powered GitOps pipeline that syncs test infrastructure from a Git repository, and your Prometheus/Grafana dashboards that surface test execution trends in real time. This guide covers every K8s topic enterprise SDET panels probe in 2026: Kubernetes architecture fundamentals for testers — translating QA concepts into K8s primitives (Pods as atomic test execution units, Deployments for long-running services like Selenium Grid hubs, Services for cross-pod networking, ConfigMaps and Secrets for test configuration, Persistent Volumes for test artifacts and reports), Selenium Grid on Kubernetes — deploying the hub-router-node architecture with Helm, configuring session queues, scaling Chrome/Firefox/Edge node pools independently, and handling browser version lifecycle, Playwright on Kubernetes — Job-based patterns for ephemeral parallel test runs, the browserless/chromium sidecar container pattern, StatefulSet considerations for persistent browser profiles, and sharding strategies for test suites with 1000+ spec files, Helm charts for test infrastructure — templating your test environment, values.yaml overrides per environment, chart versioning and rollback, and reusable library charts for multi-team organisations, CI/CD with Argo CD — GitOps workflow where your test infrastructure is declared in Git and Argo CD reconciles the cluster state, progressive delivery with canary deployments of test framework updates, and GitHub Actions triggering Argo CD syncs, Monitoring with Prometheus, Grafana, and Loki — scraping Selenium Grid metrics, building dashboards for test throughput and failure rates, alerting on infrastructure anomalies, and log aggregation for debugging flaky test pods. Every section includes production-grade YAML manifests and kubectl commands you might be asked to write, explain, or debug during a live infrastructure interview round. The SDET Interview Coach iOS app includes K8s infrastructure design challenges with AI-scored feedback — you describe your scaling architecture, resource management strategy, and CI/CD integration pattern, and the app evaluates your answer against real senior SDET interview rubrics used at top technology companies.
Published 28 May 2026 • By Mitchell Agoma
You've containerised your test suite. You've got Selenium Grid humming in Docker Compose, Playwright tests running in CI with browser images, and a neat little docker-compose.yml that everyone on the team copies. Life is good. Then your engineering director announces the Q3 goal: "We're scaling to 500 parallel end-to-end tests across three environments, running on every PR." Your Docker Compose setup hits a wall — it was never designed for multi-node orchestration, self-healing, or dynamic scaling. You need Kubernetes. But here is the trap most SDETs fall into: they learn enough K8s to run kubectl apply -f pod.yaml and call it a day. Then the interview panel asks: "Walk me through your Selenium Grid deployment on Kubernetes. How do you handle hub failover? How do you manage browser version upgrades without downtime? What does your Argo CD Application manifest look like? What Prometheus metrics are you scraping from your test pods, and what thresholds trigger alerts?" These are not trivia questions — they reveal whether you understand K8s as an operational platform or just as "Docker with more YAML."
In 2026, Kubernetes has crossed the chasm from "nice-to-have for platform teams" to "baseline expectation for test infrastructure at any company running more than 50 parallel browser tests." This guide covers every K8s topic senior SDET interview panels drill into — from the architecture fundamentals (Pods, Deployments, Services, Ingress) through to production deployment patterns (Helm, Argo CD, Prometheus). Pair this with our deep-dives on Docker Test Automation Interview Questions 2026 for the container fundamentals that precede K8s, our CI/CD Pipeline Testing Interview Questions for the pipelines that deploy to your K8s clusters, and our SDET System Design Interview Questions 2026 for the architectural thinking behind distributed test infrastructure. The SDET Interview Coach iOS app includes Kubernetes-specific infrastructure design challenges — you describe your K8s scaling strategy and get AI-scored feedback against rubrics used at Google, Amazon, and Stripe for senior SDET roles.
Kubernetes Architecture Fundamentals — The QA Engineer's Translation Layer
Kubernetes documentation is written for platform engineers — not testers. But every K8s concept maps cleanly to a test automation concern. Here is the translation layer that turns K8s primitives into QA infrastructure building blocks:
Pods — The Atomic Unit of Test Execution
The interview question: "You're running 200 Playwright tests as individual pods. A pod hangs — the browser process is alive but unresponsive. How does Kubernetes detect and recover from this?" The answer panels want: Kubernetes does not magically detect application-level hangs. It knows about process exits (restartPolicy) but not about a browser that's consuming CPU while doing nothing useful. You must configure a livenessProbe — either an HTTP GET against a /healthz endpoint in your test runner, or an exec probe that checks a heartbeat file updated every 30 seconds. When the probe fails failureThreshold times (default: 3), K8s kills the pod. If the pod is part of a Job, K8s creates a replacement. The SDET nuance: Your liveness probe must distinguish between "test is running slowly because it's waiting for a slow API" (which is fine) and "test is hung and will never complete" (which requires intervention). A common pattern: write a timestamp to /tmp/heartbeat every 30 seconds from your test runner's main loop; configure the liveness probe to check that this file is no older than 90 seconds. If your test takes longer than 90 seconds without updating the heartbeat, K8s will restart the pod — and your test framework should handle this gracefully with retry logic.
Pods cheat sheet for SDETs: Pods are ephemeral — they can die at any time, and your test framework must tolerate this (idempotent test data setup, retry mechanisms). Pods can contain multiple containers that share a network namespace — this is the sidecar pattern where your test container runs alongside a browserless/chromium container, and they communicate via localhost. Pod lifecycle: Pending → Running → Succeeded/Failed — understanding the difference between Failed (test assertion failure) and OOMKilled (pod killed by out-of-memory) is essential for debugging CI failures where the root cause is resource starvation, not a test bug.
Deployments — Managing Long-Running Test Services
The interview question: "Your Selenium Grid hub runs as a Deployment with 3 replicas. You push a new hub image. How do you ensure zero in-flight test sessions are dropped during the rollout?" The answer: Configure strategy.type: RollingUpdate with maxUnavailable: 0 and maxSurge: 1 — this guarantees at least 3 pods are always available during the rollout. But the deeper answer: for the Selenium Grid hub (which maintains in-memory session queues), a rolling Deployment alone is insufficient — you need a graceful shutdown. Set terminationGracePeriodSeconds to 120 seconds and add a preStop lifecycle hook that signals the hub to enter draining mode: stop accepting new session requests, finish all active sessions, then exit. The replacement pod starts in parallel (thanks to maxSurge: 1), registers with the Service, and begins accepting new sessions while the old pod drains. The SDET architecture insight: Deployments are ideal for stateless, horizontally scalable components — Selenium Grid hub (stateless router), Allure report server, test data API mocks, and WireMock instances. They provide declarative rollouts, automatic rollbacks (kubectl rollout undo deployment/selenium-hub), and self-healing. What trips up candidates: Deployments create ReplicaSets which create Pods — when debugging why a pod keeps restarting, check all three levels with kubectl describe to trace ownership and events.
Deployments pattern for browser nodes: Selenium Grid browser nodes (Chrome, Firefox, Edge) are typically deployed as separate Deployments — one per browser type and version. This allows independent scaling: if your test suite is 70% Chrome, you scale deployment/selenium-node-chrome to 10 replicas while keeping deployment/selenium-node-firefox at 3. Each browser node registers with the hub Service, and the hub's session queue distributes incoming test requests across available nodes.
Services and Ingress — Network Connectivity for Test Infrastructure
The interview question: "Your Playwright test pods need to reach three things: the Selenium Grid hub, the application-under-test, and an external mock API service. How do you configure this networking in Kubernetes without hardcoding IPs?" The answer: Kubernetes Services provide stable DNS names for pod-to-pod communication. Your test code references http://selenium-hub-service:4444, http://app-under-test-service:8080, and http://mock-api-service:3000. Services use label selectors to dynamically route traffic — when a pod dies and is replaced, the Service's endpoints controller updates automatically. For external access (CI pipeline hitting the Selenium Grid from outside the cluster), use an Ingress resource with a hostname like selenium-grid.qa.example.com, or a LoadBalancer Service if you're on a cloud provider. The panel's follow-up: "What happens to in-flight requests when a pod restarts?" — The Service continues routing to the remaining healthy pods. The client may see a brief connection error for in-flight requests to the restarted pod; handle this with retry logic. For critical services, configure readinessProbe (separate from livenessProbe) — the readiness probe controls whether a pod receives traffic. A pod that is alive but not ready (e.g., still loading browser binaries) won't receive requests until readiness passes. Services types for QA: ClusterIP for internal test services, NodePort for local development (minikube/kind), LoadBalancer for cloud-hosted Selenium Grid exposed to CI runners outside the cluster.
ConfigMaps, Secrets, and Persistent Volumes — Configuration and State Management
The interview question: "Your test framework needs environment URLs, browser capabilities, API keys, and a location to store test reports for post-run analysis. How do you manage these in Kubernetes without baking them into the container image?" The answer: Use ConfigMaps for non-sensitive configuration (environment URLs, browser dimensions, retry counts, parallel worker settings) and Secrets for sensitive data (API keys, database credentials, test user passwords). Both can be consumed as environment variables or mounted as files. For test reports and artifacts — screenshots, videos, trace files, Allure results — use PersistentVolumeClaims (PVCs) backed by your cloud provider's block storage (EBS on AWS, Persistent Disk on GCP). Mount the PVC to each test pod at /app/test-results, and run a separate sidecar or post-run Job that archives results to a central location like S3. The architectural insight panels love: Use one ConfigMap per environment (test-config-staging, test-config-production) — the container image remains identical across all environments, and environment-specific values are injected at deploy time. Update a ConfigMap without restarting pods by mounting it as a volume (K8s updates the file in-pod within ~60 seconds via the kubelet sync loop). For Secrets: base64 encoding is not encryption — for production, integrate with a KMS (AWS KMS, GCP KMS) using the External Secrets Operator or Sealed Secrets to encrypt secrets at rest in Git.
# Production Kubernetes Manifest for SDET Test Infrastructure
# Covers: Pods, Deployments, Services, ConfigMaps, PVCs — Real Interview Patterns
---
# ConfigMap: Environment-specific test configuration
apiVersion: v1
kind: ConfigMap
metadata:
name: test-config
namespace: qa
data:
app-base-url: "https://staging.bankscanai.com"
browser-viewport: "1920x1080"
default-timeout-seconds: "30"
max-retries: "3"
parallel-workers: "10"
selenium-grid-url: "http://selenium-hub-service.qa.svc.cluster.local:4444"
---
# Secret: API keys and credentials (use External Secrets Operator in production!)
apiVersion: v1
kind: Secret
metadata:
name: test-credentials
namespace: qa
type: Opaque
data:
api-key: WW91clNlY3JldEFQSUtleUhlcmU= # base64-encoded placeholder
test-user-password: VGVzdFBhc3N3b3JkMTIzIQ==
---
# Deployment: Selenium Grid Hub (stateless router — Deployment is the right choice)
apiVersion: apps/v1
kind: Deployment
metadata:
name: selenium-hub
namespace: qa
labels:
app: selenium-grid
component: hub
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0 # Never drop below desired replicas during rollouts
maxSurge: 1 # Allow one extra pod during rollouts
selector:
matchLabels:
app: selenium-grid
component: hub
template:
metadata:
labels:
app: selenium-grid
component: hub
spec:
terminationGracePeriodSeconds: 120 # Allow hub to drain sessions
containers:
- name: selenium-hub
image: selenium/hub:4.27
ports:
- containerPort: 4442 # Subscriber events
- containerPort: 4443 # HTTP (session queue)
- containerPort: 4444 # HTTPS / GraphQL
envFrom:
- configMapRef:
name: test-config
- secretRef:
name: test-credentials
resources:
requests:
memory: "512Mi"
cpu: "500m"
limits:
memory: "1Gi"
cpu: "1000m"
livenessProbe:
httpGet:
path: /status
port: 4444
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /status
port: 4444
initialDelaySeconds: 15
periodSeconds: 5
failureThreshold: 2
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "curl -X POST http://localhost:4444/se/grid/api/drain && sleep 90"]
---
# Service: Stable endpoint for Selenium Grid Hub
apiVersion: v1
kind: Service
metadata:
name: selenium-hub-service
namespace: qa
spec:
type: ClusterIP
selector:
app: selenium-grid
component: hub
ports:
- name: http
port: 4444
targetPort: 4444
- name: events
port: 4442
targetPort: 4442
---
# Deployment: Chrome Browser Nodes — independently scalable
apiVersion: apps/v1
kind: Deployment
metadata:
name: selenium-node-chrome
namespace: qa
spec:
replicas: 5 # Scale independently of Firefox/Edge node pools
selector:
matchLabels:
app: selenium-grid
component: node
browser: chrome
template:
metadata:
labels:
app: selenium-grid
component: node
browser: chrome
spec:
containers:
- name: selenium-node-chrome
image: selenium/node-chrome:4.27
env:
- name: SE_EVENT_BUS_HOST
value: "selenium-hub-service"
- name: SE_EVENT_BUS_PUBLISH_PORT
value: "4442"
- name: SE_EVENT_BUS_SUBSCRIBE_PORT
value: "4443"
- name: SE_NODE_MAX_SESSIONS
value: "3"
- name: SE_VNC_NO_PASSWORD
value: "1"
resources:
requests:
memory: "1Gi"
cpu: "500m"
limits:
memory: "2Gi"
cpu: "2000m"
volumeMounts:
- name: test-reports
mountPath: /tmp/test-reports
volumes:
- name: test-reports
persistentVolumeClaim:
claimName: test-reports-pvc
---
# PersistentVolumeClaim: Shared storage for test artifacts
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-reports-pvc
namespace: qa
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 50Gi
storageClassName: standard
---
# Job: Playwright test execution as an ephemeral Job
apiVersion: batch/v1
kind: Job
metadata:
name: playwright-tests-{{ .Release.Revision }}
namespace: qa
labels:
app: playwright-tests
spec:
parallelism: 10 # Run 10 test pods in parallel
completions: 10 # Wait for all 10 to complete
backoffLimit: 2 # Retry each pod up to 2 times on failure
ttlSecondsAfterFinished: 3600 # Auto-cleanup after 1 hour
template:
spec:
restartPolicy: Never
containers:
- name: playwright
image: mcr.microsoft.com/playwright:v1.52-jammy
command: ["npx", "playwright", "test", "--shard={{ .podIndex }}/{{ .numPods }}"]
envFrom:
- configMapRef:
name: test-config
env:
- name: POD_INDEX
valueFrom:
fieldRef:
fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']
resources:
requests:
memory: "1Gi"
cpu: "1000m"
limits:
memory: "2Gi"
cpu: "2000m"
livenessProbe:
exec:
command:
- /bin/sh
- -c
- "find /tmp -name heartbeat -mmin -2 | grep -q heartbeat"
initialDelaySeconds: 120
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: test-reports
mountPath: /app/test-results
volumes:
- name: test-reports
persistentVolumeClaim:
claimName: test-reports-pvc
Helm Charts — Templating Test Infrastructure for Reusability
Raw Kubernetes YAML gets you running. Helm gets you maintainable across teams, environments, and time. The SDET interview question: "You have three squads, each needing their own Selenium Grid deployment on the shared K8s cluster. How do you avoid copy-pasting 500 lines of YAML per squad?"
Helm Architecture for Test Infrastructure
Helm is the package manager for Kubernetes — it bundles related K8s manifests into a versioned, templated chart. For test infrastructure, a Helm chart typically includes: Deployment manifests for the Selenium Grid hub and browser nodes, a Service for hub-to-node communication, a ConfigMap for test configuration, a Secret for credentials, a PVC for test artifacts, and optionally a Job definition for test execution. The values.yaml file is the secret sauce: it exposes every configurable parameter — replica counts, image tags, resource limits, environment URLs — so teams can override them without touching the chart templates. The interview answer: "I built a reusable selenium-grid Helm chart with a values.yaml that exposes replica counts per browser type, node resource limits, session concurrency, and environment-specific URLs. Each squad creates a values-squad-a.yaml override file with their specific configuration and deploys via helm install selenium-grid-squad-a ./selenium-grid -f values-squad-a.yaml -n squad-a. The chart is versioned in our internal Helm repository, and squads pin to specific chart versions for reproducibility."
Helm Patterns That Interviewers Probe
Chart versioning: "How do you roll back a bad Helm release?" — helm rollback selenium-grid 2 reverts to revision 2 of the release. Helm maintains a release history (default: 10 revisions) in the cluster as Secrets. Library charts: "How do you share common test infrastructure patterns across 20 microservice teams?" — Create a library chart (type: library in Chart.yaml) that contains named templates (helpers) for common patterns: health check probes, resource defaults, pod anti-affinity rules. Application charts import and use these templates, ensuring consistency. Values hierarchy: "Where do you set defaults vs overrides?" — Defaults in values.yaml (committed to the chart), environment overrides in separate files (values-staging.yaml), secrets via --set or external secrets management. Hooks: "How do you run database migrations before your test suite?" — Use Helm hooks (helm.sh/hook: pre-install,pre-upgrade) with a Job that runs migrations, and helm.sh/hook-weight to order multiple hooks. The test suite Job uses helm.sh/hook: post-install,post-upgrade to run after infrastructure is ready.
# Helm Chart Structure for SDET Test Infrastructure
# Chart.yaml — Package metadata
apiVersion: v2
name: selenium-grid
description: A reusable Helm chart for Selenium Grid test infrastructure
type: application
version: 2.3.0
appVersion: "4.27.0"
# values.yaml — All configurable parameters with sensible defaults
replicaCount:
hub: 2
chrome: 5
firefox: 2
edge: 1
selenium:
hub:
image:
repository: selenium/hub
tag: "4.27"
resources:
requests:
memory: "512Mi"
cpu: "500m"
limits:
memory: "1Gi"
cpu: "1000m"
service:
type: ClusterIP
port: 4444
terminationGracePeriodSeconds: 120
nodes:
chrome:
maxSessions: 3
resources:
requests:
memory: "1Gi"
cpu: "500m"
limits:
memory: "2Gi"
cpu: "2000m"
firefox:
maxSessions: 3
resources:
requests:
memory: "1Gi"
cpu: "500m"
limits:
memory: "2Gi"
cpu: "2000m"
testConfig:
env: "staging"
baseUrl: "https://staging.myapp.com"
viewport: "1920x1080"
timeout: 30
retries: 2
parallelWorkers: 8
monitoring:
prometheus:
enabled: true
scrapeInterval: "30s"
grafana:
dashboardLabel: "selenium-grid"
# templates/deployment-hub.yaml — Templated Hub Deployment
# Uses Go templating: {{ .Values.selenium.hub.resources.requests.memory }}
# Command: helm install selenium-grid ./selenium-grid -f values-staging.yaml -n qa
CI/CD with Argo CD — GitOps for Test Infrastructure
Running helm install from your CI pipeline works. But GitOps with Argo CD is what senior SDET panels want to hear about — it represents the shift from "CI pushes infrastructure changes" to "infrastructure converges to what's declared in Git."
Argo CD Fundamentals for QA Infrastructure
Argo CD is a Kubernetes controller that continuously reconciles your cluster state with the desired state declared in a Git repository. For test infrastructure, your Git repo contains Helm charts, Kustomize overlays, or plain YAML manifests. Argo CD watches this repo and ensures the cluster matches — if someone manually scales a Deployment, Argo CD reverts it. If you push a change (e.g., bump the Chrome node image tag), Argo CD applies it. The interview answer: "We store our test infrastructure definitions in a Git repository — Helm charts for Selenium Grid, ConfigMaps for test configuration, and Argo CD Application manifests that point to this repo. Argo CD deploys and continuously reconciles. When we update a browser version, we update the image tag in the values file, push to Git, and Argo CD handles the rollout — with automated health checks and rollback if the sync fails. The CI pipeline triggers Argo CD syncs via the Argo CD API or by committing to Git, not by running kubectl apply directly."
Argo CD Patterns for Test Environments
App of Apps pattern: "How do you manage 50 test environments across 10 squads?" — Create a root Argo CD Application (the "App of Apps") that points to a directory containing one Application manifest per squad/environment combination. Adding a new squad means adding a new Application YAML to the directory — Argo CD picks it up and provisions the infrastructure. Progressive delivery: "How do you roll out a new test framework version safely?" — Use Argo Rollouts (the Argo CD sibling project) for canary deployments. Deploy the new test framework image to a subset of test pods, run a smoke suite, and progressively increase the rollout if metrics are healthy. PR-based ephemeral environments: "How do you create and destroy test environments per feature branch?" — Use Argo CD's ApplicationSet with a Pull Request generator. When a PR is opened, Argo CD creates a namespace with a full Selenium Grid and test configuration. When the PR is merged or closed, Argo CD garbage-collects the resources. This is the infrastructure pinnacle that FAANG-level panels probe for.
# Argo CD Application Manifests for SDET Test Infrastructure
---
# Root Application (App of Apps) — manages all squad environments
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: test-infrastructure-root
namespace: argocd
spec:
project: qa-infrastructure
source:
repoURL: https://github.com/myorg/test-infrastructure-gitops
targetRevision: main
path: apps/ # Directory containing per-squad Application YAMLs
destination:
server: https://kubernetes.default.svc
namespace: argocd
syncPolicy:
automated:
prune: true # Remove resources when Application is deleted
selfHeal: true # Revert manual changes to match Git state
allowEmpty: false
---
# Per-Squad Application: Selenium Grid for Squad Alpha
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: selenium-grid-squad-alpha
namespace: argocd
spec:
project: qa-infrastructure
source:
repoURL: https://github.com/myorg/test-infrastructure-gitops
targetRevision: main
path: helm/selenium-grid
helm:
valueFiles:
- ../../values/squad-alpha/staging.yaml # Squad-specific overrides
destination:
server: https://kubernetes.default.svc
namespace: squad-alpha-qa
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true # Auto-create the squad namespace
- PrunePropagationPolicy=foreground
retry:
limit: 3
backoff:
duration: "10s"
factor: 2
maxDuration: "3m"
---
# ApplicationSet: Ephemeral environments per feature branch (PR generator)
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: selenium-grid-pr-envs
namespace: argocd
spec:
generators:
- pullRequest:
github:
owner: myorg
repo: myapp
tokenRef:
secretName: github-token
key: token
requeueAfterSeconds: 300
template:
metadata:
name: 'selenium-grid-pr-{{ "{{" }}branch_slug{{ "}}" }}'
spec:
project: qa-infrastructure
source:
repoURL: https://github.com/myorg/test-infrastructure-gitops
targetRevision: main
path: helm/selenium-grid
helm:
valueFiles:
- ../../values/pr-env/default.yaml
destination:
server: https://kubernetes.default.svc
namespace: 'pr-{{ "{{" }}branch_slug{{ "}}" }}-qa'
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
Monitoring Test Pods — Prometheus, Grafana, and Loki
Deploying test infrastructure on Kubernetes without monitoring is like running tests without assertions — you have no idea if things are working until something breaks visibly. Senior SDET roles demand observability literacy.
Prometheus — Scraping Metrics from Test Infrastructure
Prometheus is the de facto metrics collection system in the Kubernetes ecosystem. It scrapes HTTP endpoints on your pods at regular intervals, stores time-series data, and powers alerting rules. For Selenium Grid, Prometheus can scrape the Grid's built-in metrics endpoint to track: active sessions, queued sessions, session duration percentiles, node capacity utilisation, and error rates. The interview question: "What Prometheus metrics would you alert on for a production Selenium Grid deployment?" The answer: "I'd configure four alert rules: (1) selenium_queue_size > 20 for more than 5 minutes — indicates node capacity is insufficient and tests are queuing; (2) selenium_active_sessions / selenium_max_sessions > 0.85 for 10 minutes — cluster approaching capacity; (3) rate(selenium_session_errors_total[5m]) > 5 — unusual error rate that might indicate browser version or infrastructure issues; (4) kube_pod_status_phase{phase="Failed"} > 0 — any test pods in a Failed state. For Playwright test pods, I expose custom metrics from the test runner via a Prometheus client library: test pass/fail counts, duration histograms, and flakiness indicators per spec file."
Grafana Dashboards and Loki for Log Aggregation
Grafana visualises Prometheus metrics in dashboards — test throughput over time, failure rate by browser type, node utilisation heatmaps, and session duration distributions. A well-designed Grafana dashboard lets you spot infrastructure issues (e.g., "Chrome node memory is spiking at 10 AM every day — coincides with the daily regression suite") before they become test failures. Loki is the log aggregation counterpart to Prometheus — it collects and indexes logs from all your test pods, making them searchable by label (namespace, app, browser type). "Why did pod playwright-tests-abc123 fail?" → query Loki with {app="playwright-tests", pod="playwright-tests-abc123"} and see the complete test output. The interview insight: "Prometheus tells you that something is wrong; Loki tells you why; Grafana tells you when and how often. Together they form the observability triad that turns reactive debugging into proactive infrastructure management."
# Prometheus ServiceMonitor for Selenium Grid metrics scraping
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: selenium-grid-monitor
namespace: qa
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: selenium-grid
component: hub
endpoints:
- port: http
path: /metrics
interval: 30s
scrapeTimeout: 10s
namespaceSelector:
matchNames:
- qa
---
# PrometheusRule: Alerting rules for Selenium Grid health
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: selenium-grid-alerts
namespace: qa
spec:
groups:
- name: selenium-grid
rules:
- alert: SeleniumQueueBacklog
expr: selenium_queue_size > 20
for: 5m
labels:
severity: warning
annotations:
summary: "Selenium Grid has {{ $value }} queued sessions"
description: "Queue size exceeds 20 for more than 5 minutes. Consider scaling node replicas."
- alert: SeleniumNodeCapacityLow
expr: (selenium_active_sessions / selenium_max_sessions) > 0.85
for: 10m
labels:
severity: warning
annotations:
summary: "Selenium Grid at {{ $value | humanizePercentage }} capacity"
- alert: SeleniumHighErrorRate
expr: rate(selenium_session_errors_total[5m]) > 5
for: 2m
labels:
severity: critical
annotations:
summary: "Elevated Selenium session error rate: {{ $value }}/s"
- alert: TestPodFailed
expr: kube_pod_status_phase{namespace="qa", phase="Failed"} > 0
for: 1m
labels:
severity: warning
annotations:
summary: "{{ $value }} test pods in Failed state in QA namespace"
---
# Grafana Dashboard JSON snippet — Test Execution Overview panel
# Query: histogram_quantile(0.95, rate(test_duration_seconds_bucket[5m]))
# Panel type: Time series with threshold annotation
Playwright on Kubernetes — Job Patterns, Sharding, and Sidecars
Playwright's architecture makes it a natural fit for Kubernetes — each test runs in its own browser context, parallelisation is built into the test runner, and the mcr.microsoft.com/playwright Docker image ships with all browser binaries. But scaling Playwright on K8s requires understanding the Job pattern, sharding strategy, and the sidecar vs bundled browser decision.
Kubernetes Job vs Deployment for Playwright Tests
The interview question: "When would you use a Kubernetes Job for Playwright tests vs a constantly-running Deployment?" The answer: A Job is the right choice for batch test execution — it creates pods that run to completion (success or failure) and then terminate. Configure parallelism to control how many pods run concurrently, completions for how many successful pods are needed, and backoffLimit for how many times a failing pod retries. A Deployment would be wrong here because it tries to keep pods running indefinitely — when a test completes, the Deployment controller would restart it, creating an infinite loop of test runs. The Job pattern: Use ttlSecondsAfterFinished to auto-cleanup completed Job pods. For CI-triggered test runs, each CI pipeline creates a uniquely-named Job (e.g., playwright-tests-pr-1234), waits for it to complete, collects results from the PVC, and the Job auto-deletes after the TTL. Parallelism and sharding: Set parallelism: 10 and completions: 10 — ten pods run simultaneously, each executing a different shard of your test suite. Playwright's built-in sharding (--shard=1/10) distributes spec files across pods. Each pod writes results to the shared PVC, and a post-run aggregator merges them into a unified report.
Sidecar vs Bundled Browser Pattern
The interview question: "Should your Playwright test container include the browser binaries, or should you use a separate browser sidecar container?" The bundled approach (browsers inside the test container) is simpler — one container, no inter-pod networking, works out of the box with mcr.microsoft.com/playwright. Use this when you're starting out or running tests at moderate scale (< 50 parallel pods). The sidecar approach uses browserless/chromium or similar as a separate container in the same pod — your test code connects to ws://localhost:3000 (WebSocket CDP endpoint). Benefits: (1) Browser version independent of test container — update browser images without rebuilding test containers; (2) Resource isolation — you can set separate CPU/memory limits for the browser process vs the test runner; (3) Reuse across frameworks — the same browser sidecar serves Playwright, Puppeteer, and Cypress tests. The panel wants to hear: "I start with the bundled pattern for simplicity. When I hit scale limits — browser memory pressure affecting test runner performance, or needing to update browser versions independently — I graduate to the sidecar pattern. This evolutionary approach signals I make infrastructure decisions based on real constraints, not premature optimisation."
Scaling Test Execution — HPA, Resource Limits, and Node Affinity
The whole point of Kubernetes is scaling — and SDET panels want to hear specific scaling strategies, not generic "we use autoscaling" hand-waves.
Horizontal Pod Autoscaler (HPA) for Selenium Grid Nodes
HPA automatically scales Deployment replicas based on observed metrics. For Selenium Grid: "Configure HPA on the Chrome node Deployment to scale between 3 and 20 replicas based on CPU utilisation at 70%." But CPU is a poor metric for browser automation — a Chrome node at 100% CPU might be rendering a complex WebGL page (legitimate work) or stuck in an infinite JavaScript loop (waste). The better approach: Use custom metrics via Prometheus Adapter: scale based on selenium_queue_size (the Grid hub's session queue depth). When the queue grows, HPA adds node replicas. When the queue drains, HPA scales down. This directly matches capacity to demand. The YAML: The metrics block references a Prometheus query — selenium_queue_size{namespace="qa"} — and HPA converts the metric value to desired replicas: desiredReplicas = ceil(currentReplicas * (currentMetricValue / desiredMetricValue)).
Resource Requests, Limits, and Node Affinity
Resource requests are the minimum CPU/memory a pod needs — the Kubernetes scheduler uses them to decide which node to place the pod on. Resource limits are the hard ceiling — exceeding CPU limit throttles the process; exceeding memory limit kills the pod (OOMKill). The SDET sizing rule: "For Chrome browser nodes, I request 500m CPU and 1Gi memory, limit 2000m CPU and 2Gi memory. One Chrome tab typically uses 150-300MB — multiply by SE_NODE_MAX_SESSIONS to estimate total memory. For Playwright test pods, request 1000m CPU and 1Gi memory per shard; limit 2000m CPU and 2Gi memory. Node affinity places pods on specific nodes — use nodeSelector or affinity to ensure browser-intensive pods land on nodes with SSDs and sufficient memory, while lightweight test runners can run anywhere. Taints and tolerations reserve nodes exclusively for specific workloads — taint a node with workload=browser:NoSchedule and add a matching toleration to your browser node pods, ensuring non-browser workloads don't steal resources."
Ready to Transform Your Testing?
The AI Test Automation Playbook gives you everything you need: Playwright setup, Claude AI integration, MCP deep dive, 10+ ready-to-use prompts, CI/CD pipeline setup, and a 30-day implementation roadmap.
By Mitchell Agoma, Senior SDET & AI Testing Specialist with 8+ years of experience