Skip to main content
Opik Documentation

Search documentation

Type to search this documentation.

On this pageOverview

Advanced clickhouse backup

This guide covers the two backup options available for ClickHouse in Opik's Kubernetes deployment:

  1. SQL-based Backup - Uses ClickHouse's native BACKUP command with S3
  2. ClickHouse Backup Tool - Uses the dedicated clickhouse-backup tool

ClickHouse backup is essential for data protection and disaster recovery. Opik provides two different approaches to handle backups, each with its own advantages:

  • SQL-based Backup: Simple, uses ClickHouse's built-in backup functionality
  • ClickHouse Backup Tool: More advanced, provides additional features like compression and incremental backups

This is the default backup method that uses ClickHouse's native BACKUP command to create backups directly to S3-compatible storage.

  • Uses ClickHouse's built-in BACKUP ALL EXCEPT DATABASE system command
  • Direct S3 upload with timestamped backup names
  • Configurable schedule via CronJob
  • Supports both AWS S3 and S3-compatible storage (like MinIO)

Create a Kubernetes secret with your S3 credentials:

Bash
kubectl create secret generic clickhouse-backup-secret \
  --from-literal=access_key_id=YOUR_ACCESS_KEY \
  --from-literal=access_key_secret=YOUR_SECRET_KEY

Then configure the backup:

YAML
clickhouse:
  backup:
    enabled: true
    bucketURL: "https://your-bucket.s3.region.amazonaws.com"
    secretName: "clickhouse-backup-secret"
    schedule: "0 0 * * *"

For AWS EKS clusters, you can use IAM roles instead of access keys:

YAML
clickhouse:
  serviceAccount:
    create: true
    name: "opik-clickhouse"
    annotations:
      eks.amazonaws.com/role-arn: "arn:aws:iam::ACCOUNT:role/clickhouse-backup-role"
  backup:
    enabled: true
    bucketURL: "https://your-bucket.s3.region.amazonaws.com"
    schedule: "0 0 * * *"

Required IAM Policy:

JSON
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "s3:*",
      "Resource": ["arn:aws:s3:::your-bucket", "arn:aws:s3:::your-bucket/*"]
    }
  ]
}

Trust Relationship Policy:

JSON
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": {
        "Federated": "arn:aws:iam::ACCOUNT:oidc-provider/oidc.eks.REGION.amazonaws.com/id/OIDCPROVIDERID"
      },
      "Action": "sts:AssumeRoleWithWebIdentity",
      "Condition": {
        "StringEquals": {
          "oidc.eks.REGION.amazonaws.com/id/OIDCPROVIDERID:sub": "system:serviceaccount:YOUR_NAMESPACE:opik-clickhouse",
          "oidc.eks.REGION.amazonaws.com/id/OIDCPROVIDERID:aud": "sts.amazonaws.com"
        }
      }
    }
  ]
}

You can customize the backup command if needed:

YAML
clickhouse:
  backup:
    enabled: true
    bucketURL: "https://your-bucket.s3.region.amazonaws.com"
    command:
      - /bin/bash
      - "-cx"
      - |-
        export backupname=backup$(date +'%Y%m%d%H%M')
        echo "BACKUP ALL EXCEPT DATABASE system TO S3('${CLICKHOUSE_BACKUP_BUCKET}/${backupname}/', '$ACCESS_KEY', '$SECRET_KEY');" > /tmp/backQuery.sql
        clickhouse-client -h clickhouse-opik-clickhouse --send_timeout 600000 --receive_timeout 600000 --port 9000 --queries-file=/tmp/backQuery.sql

The SQL-based backup:

  1. Creates a timestamped backup name (format: backupYYYYMMDDHHMM)
  2. Executes BACKUP ALL EXCEPT DATABASE system TO S3(...) command
  3. Uploads all databases except the system database to S3
  4. Uses ClickHouse's native backup format

To restore from a SQL-based backup:

Bash
# Connect to ClickHouse
kubectl exec -it deployment/clickhouse-opik-clickhouse -- clickhouse-client

# Restore from S3 backup
RESTORE ALL FROM S3('https://your-bucket.s3.region.amazonaws.com/backup202401011200/', 'ACCESS_KEY', 'SECRET_KEY');

The ClickHouse Backup Tool provides more advanced backup features including compression, incremental backups, and better restore capabilities.

  • Advanced backup management with compression
  • Incremental backup support
  • REST API for backup operations
  • Better restore capabilities
  • Backup metadata and validation
YAML
clickhouse:
  backupServer:
    enabled: true
    image: "altinity/clickhouse-backup:2.6.23"
    port: 7171
    env:
      LOG_LEVEL: "info"
      ALLOW_EMPTY_BACKUPS: true
      API_LISTEN: "0.0.0.0:7171"
      API_CREATE_INTEGRATION_TABLES: true

Set up S3 configuration for the backup tool:

YAML
clickhouse:
  backupServer:
    enabled: true
    env:
      S3_BUCKET: "your-backup-bucket"
      S3_ACCESS_KEY: "your-access-key" # can be ignored when use role
      S3_SECRET_KEY: "your-secret-key"
      S3_REGION: "us-west-2"
      S3_ENDPOINT: "https://s3.us-west-2.amazonaws.com" # Optional: for S3-compatible storage

Use Kubernetes secrets for sensitive data:

(can be ignored when using IAM roles)

Bash
kubectl create secret generic clickhouse-backup-tool-secret \
  --from-literal=S3_ACCESS_KEY=YOUR_ACCESS_KEY \
  --from-literal=S3_SECRET_KEY=YOUR_SECRET_KEY
YAML
clickhouse:
  backupServer:
    enabled: true
    env:
      S3_BUCKET: "your-backup-bucket"
      S3_REGION: "us-west-2"
    envFrom:
      - secretRef:
          name: "clickhouse-backup-tool-secret"
Bash
# Port-forward to access the backup server
kubectl port-forward svc/chi-opik-clickhouse-cluster-0-0 7171:7171

# Create a backup
curl -X POST "http://localhost:7171/backup/create?name=backup-$(date +%Y%m%d-%H%M%S)"

# List available backups
curl "http://localhost:7171/backup/list"
Bash
# Upload backup to S3
curl -X POST "http://localhost:7171/backup/upload/backup-20240101-120000"
Bash
# Download backup from S3
curl -X POST "http://localhost:7171/backup/download/backup-20240101-120000"

# Restore backup
curl -X POST "http://localhost:7171/backup/restore/backup-20240101-120000"

You can create a custom CronJob to automate the backup tool:

YAML
apiVersion: batch/v1
kind: CronJob
metadata:
  name: clickhouse-backup-tool-job
spec:
  schedule: "0 2 * * *" # Daily at 2 AM
  jobTemplate:
    spec:
      template:
        spec:
          containers:
            - name: backup-tool
              image: altinity/clickhouse-backup:2.6.23
              command:
                - /bin/bash
                - -c
                - |
                  BACKUP_NAME="backup-$(date +%Y%m%d-%H%M%S)"
                  curl -X POST "http://clickhouse-opik-clickhouse:7171/backup/create?name=$BACKUP_NAME"
                  sleep 30
                  curl -X POST "http://clickhouse-opik-clickhouse:7171/backup/upload/$BACKUP_NAME"
          restartPolicy: OnFailure

The Opik helm chart ships a Kubernetes Job that runs a complete restore: it picks restore or restore_remote depending on whether the backup is already on local disk, starts it through the backup server API, and polls until it finishes.

  1. Find the backup name

    backupName is required, and must match a name the backup server knows:

    Bash
    kubectl port-forward -n <namespace> svc/chi-opik-clickhouse-cluster-0-0 7171:7171
    
    # Backups in S3
    curl -s "http://localhost:7171/backup/list/remote"
    
    # Backups already on local disk
    curl -s "http://localhost:7171/backup/list/local"

    Use the name field, not the S3 prefix. With S3_PATH: shard-{shard}, a backup stored at s3://your-bucket/shard-0/2026-07-28/ has the name 2026-07-28.

    The service, pod and port used throughout this section are the chart defaults, matching the RESTORE_SERVICE the Job computes. If you set nameOverride, a different shard/replica layout, or clickhouse.backupServer.service.name / .port, substitute your own — list them with kubectl get pods,svc -n <namespace> -l clickhouse.altinity.com/chi. The port is clickhouse.backupServer.service.port when set, and otherwise clickhouse.backupServer.port (7171 by default).

  2. Render the job manifest

    The backup server has to be running in the release already. The command below renders only the Job, so nothing under backupServer reaches the cluster through it — if the server is not enabled yet, roll it out with helm upgrade first:

    YAML
    # part of your release values
    clickhouse:
      backupServer:
        enabled: true
        # S3 settings the backup server needs for restore_remote (downloads from S3)
        env:
          REMOTE_STORAGE: s3
          S3_BUCKET: YOURBUCKET
          S3_PATH: YOURPATH
          RESTORE_SCHEMA_ON_CLUSTER: cluster  # so the schema is restored on every replica

    Then put the restore settings, which are only needed at render time, in their own values file:

    YAML
    # restore-values.yaml
    clickhouse:
      backup:
        restore:
          createJob: true
          backupName: "2026-07-28"          # from step 1
          activeDeadlineSeconds: 604800     # 7 days; default is 24h
          image: "amazon/aws-cli:2.27.49"   # any image with bash and curl

    Render only the restore job. Pass your existing values file first, so the job inherits the same service account, node selector and tolerations as your ClickHouse pods:

    Bash
    helm template opik opik/opik \
      -f your-values.yaml -f restore-values.yaml \
      --show-only templates/clickhouse_restore_job.yaml > clickhouse-restore-job.yaml

    helm template does not write a namespace into the manifest, so either add namespace: to the job's metadata or pass -n when you apply it.

  3. Create the job

    Bash
    kubectl apply -f clickhouse-restore-job.yaml -n <namespace>

    The job is named <opik.name>-clickhouse-restoreopik-clickhouse-restore unless you set nameOverride, and always readable as metadata.name in the manifest you just rendered. Use that name in the commands here and in the next step. It is fixed per release, so delete the previous job before restoring again:

    Bash
    kubectl delete job opik-clickhouse-restore -n <namespace>

    Re-applying is safe: if this backup was already restored the Job exits without restoring it again, and if a restore is still running — for example after its pod was rescheduled — it follows that one instead of starting a second.

  4. Watch the restore

    Bash
    kubectl get job opik-clickhouse-restore -n <namespace>
    
    kubectl logs -f job/opik-clickhouse-restore -n <namespace>

    The job polls every 15 minutes, so the log stays quiet between checks — a large restore_remote runs for hours. It prints the final status and fails the job if the restore failed.

The job's own log only reports in progress. For actual progress, use these — in increasing cost order.

Per-table progress, from the backup server's log:

Bash
kubectl logs chi-opik-clickhouse-cluster-0-0-0 -c clickhouse-backup -n <namespace> \
  | grep download_data | tail -5

Each line is one finished table: progress=9/44 size=457.38GiB table=opik_prod.spans. Note that 9/44 is the table's position in the list, not a count of finished tables — count the distinct positions you have seen, or the numerator will look stuck while most tables are done.

Current status and start time:

Bash
kubectl exec chi-opik-clickhouse-cluster-0-0-0 -c clickhouse-backup -n <namespace> \
  -- curl -s localhost:7171/backup/actions | tail -3

Bytes landed on disk, against the backup's own size:

Bash
# how much is on the data volume now
kubectl exec chi-opik-clickhouse-cluster-0-0-0 -c clickhouse-backup -n <namespace> \
  -- df -h /var/lib/clickhouse

# the size the finished restore should reach (data_size)
kubectl exec chi-opik-clickhouse-cluster-0-0-0 -c clickhouse-backup -n <namespace> \
  -- curl -s localhost:7171/backup/list/remote

Two df readings a few minutes apart give a rough throughput and ETA, but treat that as a coarse capacity check rather than restore progress: df reports the whole volume, so merges, system tables and any other writes land in the same delta, and the total can pass the backup's data_size before the restore is done. The per-table download_data lines above are the authoritative signal. Prefer df over du -sb on the backup directory either way: du walks every file in a multi-terabyte tree and adds significant I/O to the volume the restore is already saturating.

Feature SQL-based Backup ClickHouse Backup Tool
Setup Complexity Simple Moderate
Compression No Yes
Incremental Backups No Yes
Backup Validation Basic Advanced
REST API No Yes
Restore Flexibility Basic Advanced
Resource Usage Low Moderate
S3 Compatibility Native Native
  1. Test Restores: Regularly test backup restoration procedures
  2. Monitor Backup Jobs: Set up monitoring for backup job failures
  3. Retention Policy: Implement backup retention policies
  4. Cross-Region: Consider cross-region backup replication for disaster recovery
  1. Access Control: Use IAM roles when possible instead of access keys
  2. Encryption: Enable S3 server-side encryption for backup storage
  3. Network Security: Use VPC endpoints for S3 access when available
  1. Schedule: Run backups during low-traffic periods
  2. Resource Limits: Set appropriate resource limits for backup jobs
  3. Storage Class: Use appropriate S3 storage classes for cost optimization
Bash
# Check backup job logs
kubectl logs -l app=clickhouse-backup

# Check CronJob status
kubectl get cronjobs
kubectl describe cronjob clickhouse-backup
Bash
# Test S3 connectivity
kubectl exec -it deployment/clickhouse-opik-clickhouse -- \
  clickhouse-client --query "SELECT * FROM system.disks WHERE name='s3'"
Bash
# Check backup server logs
kubectl logs -l app=clickhouse-backup-server

# Test API connectivity
kubectl port-forward svc/clickhouse-opik-clickhouse 7171:7171
curl "http://localhost:7171/backup/list"

Set up monitoring for backup operations:

YAML
# Example Prometheus alerts
- alert: ClickHouseBackupFailed
  expr: increase(kube_job_status_failed{job_name=~".*clickhouse-backup.*"}[5m]) > 0
  for: 0m
  labels:
    severity: warning
  annotations:
    summary: "ClickHouse backup job failed"
    description: "ClickHouse backup job {{ $labels.job_name }} has failed"
  1. Enable the backup server:

    YAML
    clickhouse:
      backupServer:
        enabled: true
  2. Create initial backup with the tool

  3. Disable SQL-based backup:

    YAML
    clickhouse:
      backup:
        enabled: false
  1. Disable backup server:

    YAML
    clickhouse:
      backupServer:
        enabled: false
  2. Enable SQL-based backup:

    YAML
    clickhouse:
      backup:
        enabled: true

For additional help with ClickHouse backups:

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu