Skip to content
View JoshDanielWalker's full-sized avatar
  • JLR
  • England

Organizations

@odyssey-dev

Block or report JoshDanielWalker

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
JoshDanielWalker/README.md

Joshua Daniel Walker

Principal Site Reliability Engineer

Global Multi-Cloud Kubernetes Infrastructure

Manchester, UK · Open to London · Remote & Hybrid Roles

LinkedIn


11+ yrs in tech 5+ yrs enterprise SRE 100× cluster scale 99.9% SLO

Principal SRE and Kubernetes platform engineer operating at global enterprise scale, founding and leading platform teams, scaling infrastructure from 3 nodes to over 1000, and building the reliability foundations other engineers depend on. Currently Principal Site Reliability Engineer at Jaguar Land Rover, a global enterprise automotive group operating across international markets.


// SERVICE.HISTORY

Jaguar Land Rover · Global Enterprise · Sep 2021 — Present · 5 yrs+

Role Period
Principal Site Reliability Engineer Sep 2023 — Present
Senior Site Reliability Engineer Mar 2021 — Sep 2023

Kubernetes Platform — 3 → 9 Clusters · 3 → ~100 Nodes Each · 140+ Services

Technology Context Scale
Kubernetes (GKE · EKS · IBM K8s) Multi-cloud cluster fleet, custom CRDs & operators in Go 3 → 9 clusters · ~100 nodes each
Terraform + Helm + Kustomize Full IaC cluster lifecycle, GitOps-driven provisioning 9 production clusters, zero manual config
ArgoCD + Argo Rollouts GitOps delivery, canary and blue-green rollout strategies 140+ services across all clusters
Istio + Envoy Service mesh, mTLS, traffic management, circuit breaking Cluster-wide, all inter-service traffic
Cilium / eBPF CNI, NetworkPolicy enforcement, low-level observability Node-level across full cluster fleet
GCP · AWS · IBM Cloud Multi-cloud platform engineering, VPC design, interconnect 3 cloud providers, enterprise IAM

Observability — Enterprise-wide · 140+ Services

Technology Context Scale
OpenTelemetry + Tempo Distributed tracing, unified across all engineering teams 140+ services end-to-end
Prometheus + Grafana Metrics, alerting, dashboards, SLO tracking Full cluster fleet + business services
Dynatrace (Certified) AI-powered anomaly detection, org-wide observability rollout Organisation-wide transformation
PagerDuty (Certified) On-call leadership, escalation policies, incident workflows 24/7 coverage, cross-team ownership

SRE Leadership

Founding member of JLR's DDC SRE function Grew team from 0 to multi-engineer capability
SLO programme Established error budget culture and SLO/SLI frameworks across critical systems
Incident management Blameless retrospectives, automated remediation, escalation process design
Internal Developer Platform Built IDP on Backstage — self-service cloud deployments for remote engineering teams
IaC transition Led GitLab infrastructure migration to full Infrastructure as Code

Sykes Holiday Cottages · Mid-market e-commerce · Jul 2019 — Feb 2021

DevOps Engineer · AWS ECS/Fargate Terraform Docker Bitbucket CI/CD SQS

Containerised core platform, migrated from EC2 → ECS, modernised CI/CD from Bamboo to Bitbucket Cloud.


Earlier · Web Development · 2015 — 2018

Web Developer & Designer · WordPress Shopify PHP Laravel JavaScript


// CAP.MATRIX

CAP_01  Kubernetes Platform & Cluster Scaling
CAP_02  Observability & Reliability Engineering
CAP_03  Multi-cloud Infrastructure & Networking
CAP_04  Platform Tooling & Developer Experience

// TECH.STACK

Kubernetes & Platform GKE EKS IBM K8s CRDs / Operators (Go) ArgoCD Argo Rollouts Istio Envoy Cilium eBPF

IaC & Automation Terraform Helm Kustomize Argo Workflows Vault K6 Chaos Engineering

Cloud & Networking GCP AWS IBM Cloud VPC Design Cloud Interconnect FinOps Enterprise IAM RBAC Pod Security

Observability Prometheus Grafana OpenTelemetry Tempo Dynatrace PagerDuty SLOs / Error Budgets

Languages Go Python Bash Terraform HCL JavaScript / Node.js


// ACCREDITATION

  • BSc Computer Science — 2:1 · University of Chester · 2017
  • Dynatrace Certified
  • PagerDuty Certified
  • AWS & GCP Management Essentials

// SIG.INTEL

"First member of the DDC SRE team. His remarkable drive and determination, and his ability to plan and execute complex tasks with precision, has consistently moved the platform forward."

— Chapter Lead SRE, Jaguar Land Rover · 2023

"Josh's tenacity and diligence is the main reason we implemented distributed tracing. He actively seeks feedback and makes thoughtful, well-considered contributions to technical direction."

— Senior Engineer, Jaguar Land Rover · 2023


Pinned Loading

  1. The-MilkyWay-Project/MilkyWay The-MilkyWay-Project/MilkyWay Public

    The MilkyWay, Microservices Infrastructure Project

    1

  2. The-MilkyWay-Project/Planets The-MilkyWay-Project/Planets Public

    Part of the MilkyWay, Microservices Infrastructure Project

  3. The-MilkyWay-Project/Asteroids The-MilkyWay-Project/Asteroids Public

    Part of the MilkyWay, Microservices Infrastructure Project