Current RCDiags infrastructure uses Celery for task orchestration, which has significant limitations:
- No durable execution: Worker failures cause task loss and require manual intervention
- Limited visibility: Difficult to track task state and diagnose failures
- No built-in retry: Manual retry logic required for transient failures
- Complex error handling: No standardized approach to workflow failures
- No branching logic: Cannot dynamically adjust execution based on intermediate results
This PoC demonstrates how Temporal addresses these limitations by providing:
- Parallel Execution: Multiple devices orchestrated concurrently
- Device-Level Orchestration: Hierarchical workflows (Rack → Device → Tasks)
- Dynamic Branching: Conditional execution based on health check results
- Kubernetes Job Execution: Activities create and manage Kubernetes jobs
- Failure Recovery: Workflows survive worker crashes and resume from last known state
┌─────────────────────────────────────────────────────────────┐
│ Temporal Server │
│ (localhost:8088) │
└──────────────────────┬──────────────────────────────────────┘
│
│ Task Queue: test-queue
│
┌──────────────────────▼──────────────────────────────────────┐
│ Worker │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ RackWorkflow (Parent) │ │
│ │ ┌──────────────────────────────────────────────┐ │ │
│ │ │ DeviceWorkflow (Child) │ │ │
│ │ │ ┌────────────────────────────────────────┐ │ │ │
│ │ │ │ k8s_job_activity (health check) │ │ │ │
│ │ │ │ ┌────────────────────────────────────┐ │ │ │ │
│ │ │ │ │ IF success → firmware → config │ │ │ │ │
│ │ │ │ │ IF failed → diagnostic → END │ │ │ │ │
│ │ │ │ └────────────────────────────────────┘ │ │ │ │
│ │ │ └────────────────────────────────────────┘ │ │ │
│ │ └──────────────────────────────────────────────┘ │ │
│ │ (Parallel execution) │ │
│ └──────────────────────────────────────────────────────┘ │
└──────────────────────┬──────────────────────────────────────┘
│
│ Kubernetes API
│
┌──────────────────────▼──────────────────────────────────────┐
│ Kubernetes Cluster (kind) │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ health │ │ firmware │ │ diagnostic │ │
│ │ job │ │ job │ │ job │ │
│ └─────────────┘ └─────────────┘ └─────────────┘ │
└──────────────────────────────────────────────────────────────┘
- Temporal Server: Manages workflow state and history
- Worker: Polls for workflow tasks and executes activities
- Activities: Create Kubernetes jobs, poll for completion, clean up resources
- Workflows: Orchestrate activities with branching logic and parallel execution
- Kubernetes: Executes jobs as containers in pods
Temporal UI showing parallel execution of multiple device workflows
RackWorkflow with DeviceWorkflow children executing in parallel
Device workflows taking different paths based on health check results
Workflow resuming after worker crash without losing state
- Trigger RackWorkflow with 2 devices
- Both devices execute health checks in parallel
- Device A takes diagnostic path (health fails)
- Device B takes normal path (health succeeds)
- Observe parallel execution in Temporal UI
- Health check determines execution path
- Success path: health → firmware → config
- Failure path: health → diagnostic → END
- Branching decision preserved in workflow output
- Start workflow execution
- Kill worker process mid-execution
- Restart worker
- Workflow resumes from last known state
- No duplicate jobs created
See QUICK_START.md for a 10-minute setup guide.
For detailed phase-by-phase implementation and validation, see:
- Phase 0 Report - Initial setup and simple workflow
- Phase 1 Report - Sequential task execution
- Phase 2 Report - Parallel execution
- Phase 3 Report - Child workflows
- Phase 4 Report - Dynamic branching
- Phase 5 Report - Failure recovery
- Temporal: Workflow orchestration engine
- Python: Workflow and activity implementation
- Kubernetes (kind): Container orchestration for jobs
- Docker: Container runtime
Internal PoC - Dell Technologies