You've containerised your test suite. You've got Selenium Grid humming in Docker Compose, Playwright tests running in CI with browser images, and a neat little docker-compose.yml that everyone on the team copies. Life is good. Then your engineering director announces the Q3 goal: "We're scaling to 500 parallel end-to-end tests across three environments, running on every PR." Your Docker Compose setup hits a wall — it was never designed for multi-node orchestration, self-healing, or dynamic scaling. You need Kubernetes. But here is the trap most SDETs fall into: they learn enough K8s to run kubectl apply -f pod.yaml and call it a day. Then the interview panel asks: "Walk me through your Selenium Grid deployment on Kubernetes. How do you handle hub failover? How do you manage browser version upgrades without downtime? What does your Argo CD Application manifest look like? What Prometheus metrics are you scraping from your test pods, and what thresholds trigger alerts?" These are not trivia questions — they reveal whether you understand K8s as an operational platform or just as "Docker with more YAML."

In 2026, Kubernetes has crossed the chasm from "nice-to-have for platform teams" to "baseline expectation for test infrastructure at any company running more than 50 parallel browser tests." This guide covers every K8s topic senior SDET interview panels drill into — from the architecture fundamentals (Pods, Deployments, Services, Ingress) through to production deployment patterns (Helm, Argo CD, Prometheus). Pair this with our deep-dives on Docker Test Automation Interview Questions 2026 for the container fundamentals that precede K8s, our CI/CD Pipeline Testing Interview Questions for the pipelines that deploy to your K8s clusters, and our SDET System Design Interview Questions 2026 for the architectural thinking behind distributed test infrastructure. The SDET Interview Coach iOS app includes Kubernetes-specific infrastructure design challenges — you describe your K8s scaling strategy and get AI-scored feedback against rubrics used at Google, Amazon, and Stripe for senior SDET roles.

Kubernetes Architecture Fundamentals — The QA Engineer's Translation Layer

Kubernetes documentation is written for platform engineers — not testers. But every K8s concept maps cleanly to a test automation concern. Here is the translation layer that turns K8s primitives into QA infrastructure building blocks:

Pods — The Atomic Unit of Test Execution

The interview question: "You're running 200 Playwright tests as individual pods. A pod hangs — the browser process is alive but unresponsive. How does Kubernetes detect and recover from this?" The answer panels want: Kubernetes does not magically detect application-level hangs. It knows about process exits (restartPolicy) but not about a browser that's consuming CPU while doing nothing useful. You must configure a livenessProbe — either an HTTP GET against a /healthz endpoint in your test runner, or an exec probe that checks a heartbeat file updated every 30 seconds. When the probe fails failureThreshold times (default: 3), K8s kills the pod. If the pod is part of a Job, K8s creates a replacement. The SDET nuance: Your liveness probe must distinguish between "test is running slowly because it's waiting for a slow API" (which is fine) and "test is hung and will never complete" (which requires intervention). A common pattern: write a timestamp to /tmp/heartbeat every 30 seconds from your test runner's main loop; configure the liveness probe to check that this file is no older than 90 seconds. If your test takes longer than 90 seconds without updating the heartbeat, K8s will restart the pod — and your test framework should handle this gracefully with retry logic.

Pods cheat sheet for SDETs: Pods are ephemeral — they can die at any time, and your test framework must tolerate this (idempotent test data setup, retry mechanisms). Pods can contain multiple containers that share a network namespace — this is the sidecar pattern where your test container runs alongside a browserless/chromium container, and they communicate via localhost. Pod lifecycle: Pending → Running → Succeeded/Failed — understanding the difference between Failed (test assertion failure) and OOMKilled (pod killed by out-of-memory) is essential for debugging CI failures where the root cause is resource starvation, not a test bug.

Deployments — Managing Long-Running Test Services

The interview question: "Your Selenium Grid hub runs as a Deployment with 3 replicas. You push a new hub image. How do you ensure zero in-flight test sessions are dropped during the rollout?" The answer: Configure strategy.type: RollingUpdate with maxUnavailable: 0 and maxSurge: 1 — this guarantees at least 3 pods are always available during the rollout. But the deeper answer: for the Selenium Grid hub (which maintains in-memory session queues), a rolling Deployment alone is insufficient — you need a graceful shutdown. Set terminationGracePeriodSeconds to 120 seconds and add a preStop lifecycle hook that signals the hub to enter draining mode: stop accepting new session requests, finish all active sessions, then exit. The replacement pod starts in parallel (thanks to maxSurge: 1), registers with the Service, and begins accepting new sessions while the old pod drains. The SDET architecture insight: Deployments are ideal for stateless, horizontally scalable components — Selenium Grid hub (stateless router), Allure report server, test data API mocks, and WireMock instances. They provide declarative rollouts, automatic rollbacks (kubectl rollout undo deployment/selenium-hub), and self-healing. What trips up candidates: Deployments create ReplicaSets which create Pods — when debugging why a pod keeps restarting, check all three levels with kubectl describe to trace ownership and events.

Deployments pattern for browser nodes: Selenium Grid browser nodes (Chrome, Firefox, Edge) are typically deployed as separate Deployments — one per browser type and version. This allows independent scaling: if your test suite is 70% Chrome, you scale deployment/selenium-node-chrome to 10 replicas while keeping deployment/selenium-node-firefox at 3. Each browser node registers with the hub Service, and the hub's session queue distributes incoming test requests across available nodes.

Services and Ingress — Network Connectivity for Test Infrastructure

The interview question: "Your Playwright test pods need to reach three things: the Selenium Grid hub, the application-under-test, and an external mock API service. How do you configure this networking in Kubernetes without hardcoding IPs?" The answer: Kubernetes Services provide stable DNS names for pod-to-pod communication. Your test code references http://selenium-hub-service:4444, http://app-under-test-service:8080, and http://mock-api-service:3000. Services use label selectors to dynamically route traffic — when a pod dies and is replaced, the Service's endpoints controller updates automatically. For external access (CI pipeline hitting the Selenium Grid from outside the cluster), use an Ingress resource with a hostname like selenium-grid.qa.example.com, or a LoadBalancer Service if you're on a cloud provider. The panel's follow-up: "What happens to in-flight requests when a pod restarts?" — The Service continues routing to the remaining healthy pods. The client may see a brief connection error for in-flight requests to the restarted pod; handle this with retry logic. For critical services, configure readinessProbe (separate from livenessProbe) — the readiness probe controls whether a pod receives traffic. A pod that is alive but not ready (e.g., still loading browser binaries) won't receive requests until readiness passes. Services types for QA: ClusterIP for internal test services, NodePort for local development (minikube/kind), LoadBalancer for cloud-hosted Selenium Grid exposed to CI runners outside the cluster.

ConfigMaps, Secrets, and Persistent Volumes — Configuration and State Management

The interview question: "Your test framework needs environment URLs, browser capabilities, API keys, and a location to store test reports for post-run analysis. How do you manage these in Kubernetes without baking them into the container image?" The answer: Use ConfigMaps for non-sensitive configuration (environment URLs, browser dimensions, retry counts, parallel worker settings) and Secrets for sensitive data (API keys, database credentials, test user passwords). Both can be consumed as environment variables or mounted as files. For test reports and artifacts — screenshots, videos, trace files, Allure results — use PersistentVolumeClaims (PVCs) backed by your cloud provider's block storage (EBS on AWS, Persistent Disk on GCP). Mount the PVC to each test pod at /app/test-results, and run a separate sidecar or post-run Job that archives results to a central location like S3. The architectural insight panels love: Use one ConfigMap per environment (test-config-staging, test-config-production) — the container image remains identical across all environments, and environment-specific values are injected at deploy time. Update a ConfigMap without restarting pods by mounting it as a volume (K8s updates the file in-pod within ~60 seconds via the kubelet sync loop). For Secrets: base64 encoding is not encryption — for production, integrate with a KMS (AWS KMS, GCP KMS) using the External Secrets Operator or Sealed Secrets to encrypt secrets at rest in Git.

# Production Kubernetes Manifest for SDET Test Infrastructure
# Covers: Pods, Deployments, Services, ConfigMaps, PVCs — Real Interview Patterns

---
# ConfigMap: Environment-specific test configuration
apiVersion: v1
kind: ConfigMap
metadata:
  name: test-config
  namespace: qa
data:
  app-base-url: "https://staging.bankscanai.com"
  browser-viewport: "1920x1080"
  default-timeout-seconds: "30"
  max-retries: "3"
  parallel-workers: "10"
  selenium-grid-url: "http://selenium-hub-service.qa.svc.cluster.local:4444"

---
# Secret: API keys and credentials (use External Secrets Operator in production!)
apiVersion: v1
kind: Secret
metadata:
  name: test-credentials
  namespace: qa
type: Opaque
data:
  api-key: WW91clNlY3JldEFQSUtleUhlcmU=  # base64-encoded placeholder
  test-user-password: VGVzdFBhc3N3b3JkMTIzIQ==

---
# Deployment: Selenium Grid Hub (stateless router — Deployment is the right choice)
apiVersion: apps/v1
kind: Deployment
metadata:
  name: selenium-hub
  namespace: qa
  labels:
    app: selenium-grid
    component: hub
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0   # Never drop below desired replicas during rollouts
      maxSurge: 1         # Allow one extra pod during rollouts
  selector:
    matchLabels:
      app: selenium-grid
      component: hub
  template:
    metadata:
      labels:
        app: selenium-grid
        component: hub
    spec:
      terminationGracePeriodSeconds: 120  # Allow hub to drain sessions
      containers:
        - name: selenium-hub
          image: selenium/hub:4.27
          ports:
            - containerPort: 4442  # Subscriber events
            - containerPort: 4443  # HTTP (session queue)
            - containerPort: 4444  # HTTPS / GraphQL
          envFrom:
            - configMapRef:
                name: test-config
            - secretRef:
                name: test-credentials
          resources:
            requests:
              memory: "512Mi"
              cpu: "500m"
            limits:
              memory: "1Gi"
              cpu: "1000m"
          livenessProbe:
            httpGet:
              path: /status
              port: 4444
            initialDelaySeconds: 30
            periodSeconds: 10
            failureThreshold: 3
          readinessProbe:
            httpGet:
              path: /status
              port: 4444
            initialDelaySeconds: 15
            periodSeconds: 5
            failureThreshold: 2
          lifecycle:
            preStop:
              exec:
                command: ["/bin/sh", "-c", "curl -X POST http://localhost:4444/se/grid/api/drain && sleep 90"]

---
# Service: Stable endpoint for Selenium Grid Hub
apiVersion: v1
kind: Service
metadata:
  name: selenium-hub-service
  namespace: qa
spec:
  type: ClusterIP
  selector:
    app: selenium-grid
    component: hub
  ports:
    - name: http
      port: 4444
      targetPort: 4444
    - name: events
      port: 4442
      targetPort: 4442

---
# Deployment: Chrome Browser Nodes — independently scalable
apiVersion: apps/v1
kind: Deployment
metadata:
  name: selenium-node-chrome
  namespace: qa
spec:
  replicas: 5  # Scale independently of Firefox/Edge node pools
  selector:
    matchLabels:
      app: selenium-grid
      component: node
      browser: chrome
  template:
    metadata:
      labels:
        app: selenium-grid
        component: node
        browser: chrome
    spec:
      containers:
        - name: selenium-node-chrome
          image: selenium/node-chrome:4.27
          env:
            - name: SE_EVENT_BUS_HOST
              value: "selenium-hub-service"
            - name: SE_EVENT_BUS_PUBLISH_PORT
              value: "4442"
            - name: SE_EVENT_BUS_SUBSCRIBE_PORT
              value: "4443"
            - name: SE_NODE_MAX_SESSIONS
              value: "3"
            - name: SE_VNC_NO_PASSWORD
              value: "1"
          resources:
            requests:
              memory: "1Gi"
              cpu: "500m"
            limits:
              memory: "2Gi"
              cpu: "2000m"
          volumeMounts:
            - name: test-reports
              mountPath: /tmp/test-reports
      volumes:
        - name: test-reports
          persistentVolumeClaim:
            claimName: test-reports-pvc

---
# PersistentVolumeClaim: Shared storage for test artifacts
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: test-reports-pvc
  namespace: qa
spec:
  accessModes:
    - ReadWriteMany
  resources:
    requests:
      storage: 50Gi
  storageClassName: standard

---
# Job: Playwright test execution as an ephemeral Job
apiVersion: batch/v1
kind: Job
metadata:
  name: playwright-tests-{{ .Release.Revision }}
  namespace: qa
  labels:
    app: playwright-tests
spec:
  parallelism: 10        # Run 10 test pods in parallel
  completions: 10        # Wait for all 10 to complete
  backoffLimit: 2        # Retry each pod up to 2 times on failure
  ttlSecondsAfterFinished: 3600  # Auto-cleanup after 1 hour
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: playwright
          image: mcr.microsoft.com/playwright:v1.52-jammy
          command: ["npx", "playwright", "test", "--shard={{ .podIndex }}/{{ .numPods }}"]
          envFrom:
            - configMapRef:
                name: test-config
          env:
            - name: POD_INDEX
              valueFrom:
                fieldRef:
                  fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']
          resources:
            requests:
              memory: "1Gi"
              cpu: "1000m"
            limits:
              memory: "2Gi"
              cpu: "2000m"
          livenessProbe:
            exec:
              command:
                - /bin/sh
                - -c
                - "find /tmp -name heartbeat -mmin -2 | grep -q heartbeat"
            initialDelaySeconds: 120
            periodSeconds: 30
            failureThreshold: 3
          volumeMounts:
            - name: test-reports
              mountPath: /app/test-results
      volumes:
        - name: test-reports
          persistentVolumeClaim:
            claimName: test-reports-pvc

Helm Charts — Templating Test Infrastructure for Reusability

Raw Kubernetes YAML gets you running. Helm gets you maintainable across teams, environments, and time. The SDET interview question: "You have three squads, each needing their own Selenium Grid deployment on the shared K8s cluster. How do you avoid copy-pasting 500 lines of YAML per squad?"

Helm Architecture for Test Infrastructure

Helm is the package manager for Kubernetes — it bundles related K8s manifests into a versioned, templated chart. For test infrastructure, a Helm chart typically includes: Deployment manifests for the Selenium Grid hub and browser nodes, a Service for hub-to-node communication, a ConfigMap for test configuration, a Secret for credentials, a PVC for test artifacts, and optionally a Job definition for test execution. The values.yaml file is the secret sauce: it exposes every configurable parameter — replica counts, image tags, resource limits, environment URLs — so teams can override them without touching the chart templates. The interview answer: "I built a reusable selenium-grid Helm chart with a values.yaml that exposes replica counts per browser type, node resource limits, session concurrency, and environment-specific URLs. Each squad creates a values-squad-a.yaml override file with their specific configuration and deploys via helm install selenium-grid-squad-a ./selenium-grid -f values-squad-a.yaml -n squad-a. The chart is versioned in our internal Helm repository, and squads pin to specific chart versions for reproducibility."

Helm Patterns That Interviewers Probe

Chart versioning: "How do you roll back a bad Helm release?" — helm rollback selenium-grid 2 reverts to revision 2 of the release. Helm maintains a release history (default: 10 revisions) in the cluster as Secrets. Library charts: "How do you share common test infrastructure patterns across 20 microservice teams?" — Create a library chart (type: library in Chart.yaml) that contains named templates (helpers) for common patterns: health check probes, resource defaults, pod anti-affinity rules. Application charts import and use these templates, ensuring consistency. Values hierarchy: "Where do you set defaults vs overrides?" — Defaults in values.yaml (committed to the chart), environment overrides in separate files (values-staging.yaml), secrets via --set or external secrets management. Hooks: "How do you run database migrations before your test suite?" — Use Helm hooks (helm.sh/hook: pre-install,pre-upgrade) with a Job that runs migrations, and helm.sh/hook-weight to order multiple hooks. The test suite Job uses helm.sh/hook: post-install,post-upgrade to run after infrastructure is ready.

# Helm Chart Structure for SDET Test Infrastructure

# Chart.yaml — Package metadata
apiVersion: v2
name: selenium-grid
description: A reusable Helm chart for Selenium Grid test infrastructure
type: application
version: 2.3.0
appVersion: "4.27.0"

# values.yaml — All configurable parameters with sensible defaults
replicaCount:
  hub: 2
  chrome: 5
  firefox: 2
  edge: 1

selenium:
  hub:
    image:
      repository: selenium/hub
      tag: "4.27"
    resources:
      requests:
        memory: "512Mi"
        cpu: "500m"
      limits:
        memory: "1Gi"
        cpu: "1000m"
    service:
      type: ClusterIP
      port: 4444
    terminationGracePeriodSeconds: 120
  nodes:
    chrome:
      maxSessions: 3
      resources:
        requests:
          memory: "1Gi"
          cpu: "500m"
        limits:
          memory: "2Gi"
          cpu: "2000m"
    firefox:
      maxSessions: 3
      resources:
        requests:
          memory: "1Gi"
          cpu: "500m"
        limits:
          memory: "2Gi"
          cpu: "2000m"

testConfig:
  env: "staging"
  baseUrl: "https://staging.myapp.com"
  viewport: "1920x1080"
  timeout: 30
  retries: 2
  parallelWorkers: 8

monitoring:
  prometheus:
    enabled: true
    scrapeInterval: "30s"
  grafana:
    dashboardLabel: "selenium-grid"

# templates/deployment-hub.yaml — Templated Hub Deployment
# Uses Go templating: {{ .Values.selenium.hub.resources.requests.memory }}
# Command: helm install selenium-grid ./selenium-grid -f values-staging.yaml -n qa

CI/CD with Argo CD — GitOps for Test Infrastructure

Running helm install from your CI pipeline works. But GitOps with Argo CD is what senior SDET panels want to hear about — it represents the shift from "CI pushes infrastructure changes" to "infrastructure converges to what's declared in Git."

Argo CD Fundamentals for QA Infrastructure

Argo CD is a Kubernetes controller that continuously reconciles your cluster state with the desired state declared in a Git repository. For test infrastructure, your Git repo contains Helm charts, Kustomize overlays, or plain YAML manifests. Argo CD watches this repo and ensures the cluster matches — if someone manually scales a Deployment, Argo CD reverts it. If you push a change (e.g., bump the Chrome node image tag), Argo CD applies it. The interview answer: "We store our test infrastructure definitions in a Git repository — Helm charts for Selenium Grid, ConfigMaps for test configuration, and Argo CD Application manifests that point to this repo. Argo CD deploys and continuously reconciles. When we update a browser version, we update the image tag in the values file, push to Git, and Argo CD handles the rollout — with automated health checks and rollback if the sync fails. The CI pipeline triggers Argo CD syncs via the Argo CD API or by committing to Git, not by running kubectl apply directly."

Argo CD Patterns for Test Environments

App of Apps pattern: "How do you manage 50 test environments across 10 squads?" — Create a root Argo CD Application (the "App of Apps") that points to a directory containing one Application manifest per squad/environment combination. Adding a new squad means adding a new Application YAML to the directory — Argo CD picks it up and provisions the infrastructure. Progressive delivery: "How do you roll out a new test framework version safely?" — Use Argo Rollouts (the Argo CD sibling project) for canary deployments. Deploy the new test framework image to a subset of test pods, run a smoke suite, and progressively increase the rollout if metrics are healthy. PR-based ephemeral environments: "How do you create and destroy test environments per feature branch?" — Use Argo CD's ApplicationSet with a Pull Request generator. When a PR is opened, Argo CD creates a namespace with a full Selenium Grid and test configuration. When the PR is merged or closed, Argo CD garbage-collects the resources. This is the infrastructure pinnacle that FAANG-level panels probe for.

# Argo CD Application Manifests for SDET Test Infrastructure

---
# Root Application (App of Apps) — manages all squad environments
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: test-infrastructure-root
  namespace: argocd
spec:
  project: qa-infrastructure
  source:
    repoURL: https://github.com/myorg/test-infrastructure-gitops
    targetRevision: main
    path: apps/  # Directory containing per-squad Application YAMLs
  destination:
    server: https://kubernetes.default.svc
    namespace: argocd
  syncPolicy:
    automated:
      prune: true       # Remove resources when Application is deleted
      selfHeal: true    # Revert manual changes to match Git state
      allowEmpty: false

---
# Per-Squad Application: Selenium Grid for Squad Alpha
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: selenium-grid-squad-alpha
  namespace: argocd
spec:
  project: qa-infrastructure
  source:
    repoURL: https://github.com/myorg/test-infrastructure-gitops
    targetRevision: main
    path: helm/selenium-grid
    helm:
      valueFiles:
        - ../../values/squad-alpha/staging.yaml  # Squad-specific overrides
  destination:
    server: https://kubernetes.default.svc
    namespace: squad-alpha-qa
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
    syncOptions:
      - CreateNamespace=true     # Auto-create the squad namespace
      - PrunePropagationPolicy=foreground
    retry:
      limit: 3
      backoff:
        duration: "10s"
        factor: 2
        maxDuration: "3m"

---
# ApplicationSet: Ephemeral environments per feature branch (PR generator)
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: selenium-grid-pr-envs
  namespace: argocd
spec:
  generators:
    - pullRequest:
        github:
          owner: myorg
          repo: myapp
          tokenRef:
            secretName: github-token
            key: token
        requeueAfterSeconds: 300
  template:
    metadata:
      name: 'selenium-grid-pr-{{ "{{" }}branch_slug{{ "}}" }}'
    spec:
      project: qa-infrastructure
      source:
        repoURL: https://github.com/myorg/test-infrastructure-gitops
        targetRevision: main
        path: helm/selenium-grid
        helm:
          valueFiles:
            - ../../values/pr-env/default.yaml
      destination:
        server: https://kubernetes.default.svc
        namespace: 'pr-{{ "{{" }}branch_slug{{ "}}" }}-qa'
      syncPolicy:
        automated:
          prune: true
          selfHeal: true
        syncOptions:
          - CreateNamespace=true

Monitoring Test Pods — Prometheus, Grafana, and Loki

Deploying test infrastructure on Kubernetes without monitoring is like running tests without assertions — you have no idea if things are working until something breaks visibly. Senior SDET roles demand observability literacy.

Prometheus — Scraping Metrics from Test Infrastructure

Prometheus is the de facto metrics collection system in the Kubernetes ecosystem. It scrapes HTTP endpoints on your pods at regular intervals, stores time-series data, and powers alerting rules. For Selenium Grid, Prometheus can scrape the Grid's built-in metrics endpoint to track: active sessions, queued sessions, session duration percentiles, node capacity utilisation, and error rates. The interview question: "What Prometheus metrics would you alert on for a production Selenium Grid deployment?" The answer: "I'd configure four alert rules: (1) selenium_queue_size > 20 for more than 5 minutes — indicates node capacity is insufficient and tests are queuing; (2) selenium_active_sessions / selenium_max_sessions > 0.85 for 10 minutes — cluster approaching capacity; (3) rate(selenium_session_errors_total[5m]) > 5 — unusual error rate that might indicate browser version or infrastructure issues; (4) kube_pod_status_phase{phase="Failed"} > 0 — any test pods in a Failed state. For Playwright test pods, I expose custom metrics from the test runner via a Prometheus client library: test pass/fail counts, duration histograms, and flakiness indicators per spec file."

Grafana Dashboards and Loki for Log Aggregation

Grafana visualises Prometheus metrics in dashboards — test throughput over time, failure rate by browser type, node utilisation heatmaps, and session duration distributions. A well-designed Grafana dashboard lets you spot infrastructure issues (e.g., "Chrome node memory is spiking at 10 AM every day — coincides with the daily regression suite") before they become test failures. Loki is the log aggregation counterpart to Prometheus — it collects and indexes logs from all your test pods, making them searchable by label (namespace, app, browser type). "Why did pod playwright-tests-abc123 fail?" → query Loki with {app="playwright-tests", pod="playwright-tests-abc123"} and see the complete test output. The interview insight: "Prometheus tells you that something is wrong; Loki tells you why; Grafana tells you when and how often. Together they form the observability triad that turns reactive debugging into proactive infrastructure management."

# Prometheus ServiceMonitor for Selenium Grid metrics scraping

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: selenium-grid-monitor
  namespace: qa
  labels:
    release: prometheus-stack
spec:
  selector:
    matchLabels:
      app: selenium-grid
      component: hub
  endpoints:
    - port: http
      path: /metrics
      interval: 30s
      scrapeTimeout: 10s
  namespaceSelector:
    matchNames:
      - qa

---
# PrometheusRule: Alerting rules for Selenium Grid health
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: selenium-grid-alerts
  namespace: qa
spec:
  groups:
    - name: selenium-grid
      rules:
        - alert: SeleniumQueueBacklog
          expr: selenium_queue_size > 20
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: "Selenium Grid has {{ $value }} queued sessions"
            description: "Queue size exceeds 20 for more than 5 minutes. Consider scaling node replicas."
        - alert: SeleniumNodeCapacityLow
          expr: (selenium_active_sessions / selenium_max_sessions) > 0.85
          for: 10m
          labels:
            severity: warning
          annotations:
            summary: "Selenium Grid at {{ $value | humanizePercentage }} capacity"
        - alert: SeleniumHighErrorRate
          expr: rate(selenium_session_errors_total[5m]) > 5
          for: 2m
          labels:
            severity: critical
          annotations:
            summary: "Elevated Selenium session error rate: {{ $value }}/s"
        - alert: TestPodFailed
          expr: kube_pod_status_phase{namespace="qa", phase="Failed"} > 0
          for: 1m
          labels:
            severity: warning
          annotations:
            summary: "{{ $value }} test pods in Failed state in QA namespace"

---
# Grafana Dashboard JSON snippet — Test Execution Overview panel
# Query: histogram_quantile(0.95, rate(test_duration_seconds_bucket[5m]))
# Panel type: Time series with threshold annotation

Playwright on Kubernetes — Job Patterns, Sharding, and Sidecars

Playwright's architecture makes it a natural fit for Kubernetes — each test runs in its own browser context, parallelisation is built into the test runner, and the mcr.microsoft.com/playwright Docker image ships with all browser binaries. But scaling Playwright on K8s requires understanding the Job pattern, sharding strategy, and the sidecar vs bundled browser decision.

Kubernetes Job vs Deployment for Playwright Tests

The interview question: "When would you use a Kubernetes Job for Playwright tests vs a constantly-running Deployment?" The answer: A Job is the right choice for batch test execution — it creates pods that run to completion (success or failure) and then terminate. Configure parallelism to control how many pods run concurrently, completions for how many successful pods are needed, and backoffLimit for how many times a failing pod retries. A Deployment would be wrong here because it tries to keep pods running indefinitely — when a test completes, the Deployment controller would restart it, creating an infinite loop of test runs. The Job pattern: Use ttlSecondsAfterFinished to auto-cleanup completed Job pods. For CI-triggered test runs, each CI pipeline creates a uniquely-named Job (e.g., playwright-tests-pr-1234), waits for it to complete, collects results from the PVC, and the Job auto-deletes after the TTL. Parallelism and sharding: Set parallelism: 10 and completions: 10 — ten pods run simultaneously, each executing a different shard of your test suite. Playwright's built-in sharding (--shard=1/10) distributes spec files across pods. Each pod writes results to the shared PVC, and a post-run aggregator merges them into a unified report.

Sidecar vs Bundled Browser Pattern

The interview question: "Should your Playwright test container include the browser binaries, or should you use a separate browser sidecar container?" The bundled approach (browsers inside the test container) is simpler — one container, no inter-pod networking, works out of the box with mcr.microsoft.com/playwright. Use this when you're starting out or running tests at moderate scale (< 50 parallel pods). The sidecar approach uses browserless/chromium or similar as a separate container in the same pod — your test code connects to ws://localhost:3000 (WebSocket CDP endpoint). Benefits: (1) Browser version independent of test container — update browser images without rebuilding test containers; (2) Resource isolation — you can set separate CPU/memory limits for the browser process vs the test runner; (3) Reuse across frameworks — the same browser sidecar serves Playwright, Puppeteer, and Cypress tests. The panel wants to hear: "I start with the bundled pattern for simplicity. When I hit scale limits — browser memory pressure affecting test runner performance, or needing to update browser versions independently — I graduate to the sidecar pattern. This evolutionary approach signals I make infrastructure decisions based on real constraints, not premature optimisation."

Scaling Test Execution — HPA, Resource Limits, and Node Affinity

The whole point of Kubernetes is scaling — and SDET panels want to hear specific scaling strategies, not generic "we use autoscaling" hand-waves.

Horizontal Pod Autoscaler (HPA) for Selenium Grid Nodes

HPA automatically scales Deployment replicas based on observed metrics. For Selenium Grid: "Configure HPA on the Chrome node Deployment to scale between 3 and 20 replicas based on CPU utilisation at 70%." But CPU is a poor metric for browser automation — a Chrome node at 100% CPU might be rendering a complex WebGL page (legitimate work) or stuck in an infinite JavaScript loop (waste). The better approach: Use custom metrics via Prometheus Adapter: scale based on selenium_queue_size (the Grid hub's session queue depth). When the queue grows, HPA adds node replicas. When the queue drains, HPA scales down. This directly matches capacity to demand. The YAML: The metrics block references a Prometheus query — selenium_queue_size{namespace="qa"} — and HPA converts the metric value to desired replicas: desiredReplicas = ceil(currentReplicas * (currentMetricValue / desiredMetricValue)).

Resource Requests, Limits, and Node Affinity

Resource requests are the minimum CPU/memory a pod needs — the Kubernetes scheduler uses them to decide which node to place the pod on. Resource limits are the hard ceiling — exceeding CPU limit throttles the process; exceeding memory limit kills the pod (OOMKill). The SDET sizing rule: "For Chrome browser nodes, I request 500m CPU and 1Gi memory, limit 2000m CPU and 2Gi memory. One Chrome tab typically uses 150-300MB — multiply by SE_NODE_MAX_SESSIONS to estimate total memory. For Playwright test pods, request 1000m CPU and 1Gi memory per shard; limit 2000m CPU and 2Gi memory. Node affinity places pods on specific nodes — use nodeSelector or affinity to ensure browser-intensive pods land on nodes with SSDs and sufficient memory, while lightweight test runners can run anywhere. Taints and tolerations reserve nodes exclusively for specific workloads — taint a node with workload=browser:NoSchedule and add a matching toleration to your browser node pods, ensuring non-browser workloads don't steal resources."

Ready to Transform Your Testing?

The AI Test Automation Playbook gives you everything you need: Playwright setup, Claude AI integration, MCP deep dive, 10+ ready-to-use prompts, CI/CD pipeline setup, and a 30-day implementation roadmap.

✅ Playwright + TypeScript✅ Claude AI Prompts✅ MCP Deep Dive✅ CI/CD with GitHub Actions✅ 30-Day Roadmap✅ Page Object Patterns
Get the AI Test Automation Playbook — $49.99

By Mitchell Agoma, Senior SDET & AI Testing Specialist with 8+ years of experience