Legacy CI/CD pipelines relied on push-based deployment scripts that executed kubectl apply commands from external CI runners directly into production clusters. This anti-pattern exposed cluster credentials to CI runners and frequently caused environment drift between Git source code and live cluster states.

**GitOps** replaces push-based deployments with a pull-based model where Git acts as the Single Source of Truth for infrastructure and application code. In this guide, we explore GitOps reconciliation loops (ArgoCD/Flux), progressive delivery strategies (Canary vs Blue-Green), and production GitHub Actions workflows.

1. The Four Core Principles of GitOps

  1. Declarative System Descriptions: The entire environment (Kubernetes manifests, Helm charts, Kustomize overlays) is described declaratively in Git.
  2. Versioned and Immutable State: Desired state is stored in Git with a complete audit log of commits, branches, and tags.
  3. Automated Pull-Based Sync: In-cluster agents (ArgoCD or Flux) continuously pull changes from Git and apply them to the cluster.
  4. Self-Healing Continuous Reconciliation: If someone manually modifies a production pod via kubectl edit, the GitOps operator detects state drift and immediately overwrites the cluster back to the Git target state.

2. Progressive Delivery: Blue-Green vs Canary Deployments

Deployment Strategy Traffic Routing Mechanics Rollback Speed & Risk Profile
Blue-Green Deployment Spins up a 100% complete new environment (Green). Ingress router switches traffic from Blue to Green instantly. Instant Rollback: Switch router back to Blue. Requires $2\times$ hardware resource capacity during deployment window.
Canary Deployment Routes a small fraction of real user traffic (e.g., 5% -> 25% -> 50% -> 100%) to new version while monitoring error metrics. Lowest Risk: Automatically aborts rollout if HTTP 5xx error rates or P99 latencies spike in Prometheus. Minimal resource overhead.

3. Automated Canary Rollouts with Argo Rollouts & Prometheus

Tools like **Argo Rollouts** automate progressive canary deployments by querying Prometheus metrics during each step of the rollout:

# Argo Rollout Canary Specification with Prometheus Metric Analysis apiVersion: argoproj.io/v1alpha1 kind: Rollout metadata: name: api-service-rollout spec: replicas: 10 strategy: canary: analysis: templates: - templateName: success-rate-check args: - name: service-name value: api-service steps: - setWeight: 10 # Route 10% traffic to Canary - pause: { duration: 10m } # Pause 10 mins for metric analysis - setWeight: 50 # Route 50% traffic if metrics pass - pause: { duration: 30m } - setWeight: 100 # Full rollout

4. Production GitHub Actions CI Workflow

Below is a production-grade GitHub Actions workflow that runs automated unit tests, builds a multi-arch Docker image, scans for security vulnerabilities using Trivy, and updates the GitOps manifest repository:

name: Production Build & GitOps Sync Pipeline on: push: branches: [ "main" ] jobs: build-and-test: runs-on: ubuntu-latest steps: - name: Checkout Application Code uses: actions/checkout@v4 - name: Set up Node.js Environment uses: actions/setup-node@v4 with: node-version: 20 cache: 'npm' - name: Run Test Suite run: | npm ci npm test - name: Set up Docker Buildx uses: docker/setup-buildx-action@v3 - name: Log in to Container Registry uses: docker/login-action@v3 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - name: Build & Push Production Image uses: docker/build-push-action@v5 with: context: . push: true tags: | ghcr.io/mashaewilliams/api-service:${{ github.sha }} ghcr.io/mashaewilliams/api-service:latest cache-from: type=gha cache-to: type=gha,mode=max - name: Vulnerability Scan via Trivy uses: aquasecurity/trivy-action@master with: image-ref: 'ghcr.io/mashaewilliams/api-service:${{ github.sha }}' format: 'table' exit-code: '1' severity: 'CRITICAL' - name: Update GitOps Manifest Repository env: GITOPS_TOKEN: ${{ secrets.GITOPS_PAT }} run: | git config --global user.email "ci@mashaewilliamsllc.com" git config --global user.name "GitOps Bot" git clone https://x-access-token:${GITOPS_TOKEN}@github.com/mashaewilliams/k8s-gitops-config.git cd k8s-gitops-config/environments/production sed -i 's|image: ghcr.io/mashaewilliams/api-service:.*|image: ghcr.io/mashaewilliams/api-service:'"${GITHUB_SHA}"'|g' deployment.yaml git commit -am "chore(deps): bump api-service to ${GITHUB_SHA}" git push origin main

5. Key Takeaways for High-Velocity Engineering Teams

  • Separate Source Code & Infrastructure Repositories: Keep application source code in an application repo, and Kubernetes manifests in a dedicated GitOps repository to prevent infinite CI trigger loops.
  • Enforce Automated Rollbacks: Ensure canary metric analysis automatically triggers a rollback if error rates exceed 0.5% during the first 10 minutes of a release.