← All projects
Single-Touch Edge AI Platform - cover art

Case study

Single-Touch Edge AI Platform

Turned a high-level edge-AI design into a single-touch deployment running on Kubernetes at the store edge.

Infrastructure / DevOps Engineer · Woolworths·2025 - Present

  • Kubernetes
  • Edge
  • NVIDIA GPU
  • CD pipelines
  • Helm
  • Python
Design → single-touch pipeline → readiness-gated GPU inference at the edge
Design → single-touch pipeline → readiness-gated GPU inference at the edge

Problem

Edge AI at retail scale lives or dies on repeatability. A computer-vision workload that runs perfectly in a lab has to come up the same way in a store with no on-site engineer, flaky connectivity, and a GPU that may not be ready the instant Kubernetes wants to schedule against it. The starting point was a high-level design and a pile of manual steps - exactly the gap between "it works" and "it ships."

Constraints

  • No hands at the edge. Deployment has to be hands-off and idempotent - a single press.
  • GPU timing. The kubelet can report Ready before the NVIDIA device plugin has advertised nvidia.com/gpu. An inference container that starts in that window crash-loops and poisons the rollout.
  • Heterogeneous stores. Per-site variables (network, hardware, identity) without forking the platform for every location.

Design

I took the high-level designs and turned them into low-level, problem-solving deployments driven by CD pipelines. The application is packaged as containers and shipped to a store-edge Kubernetes cluster via Helm with end-state manifests. Per-store configuration is injected from a single source of truth, so one pipeline produces a correct deployment for any site.

Three separate controls do three separate jobs. A nvidia.com/gpu resource request decides where the pod lands. An init container blocks until nvidia-smi enumerates a device, so the inference container cannot start before the hardware answers - and it fails loudly, with a bound, rather than waiting forever. Runtime probes and watchdogs then track health after start.

Security & reliability decisions

  • Init-gated GPU startup - the single biggest reliability win; the app container no longer races the device plugin at boot.
  • Single source of truth for config, rendered through templates - the class of divergence where two stores disagree on the same value fails at render time rather than in production.
  • Spec-driven, documented-as-code - the deployment is the documentation.

Outcome

A high-level idea becomes a real, repeatable deployment on a single press. New edge sites come up consistently, GPUs come online reliably, and the manual runbook is gone - replaced by a pipeline anyone on the team can trigger.

Future improvements

Push more of the per-store delta into declarative policy, and extend the readiness model to cover the full inference dependency chain (model artifacts, egress, downstream sinks) as a single health gate.