Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions ChangeLog.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ NEW FEATURES
- `delete local|remote <backup_name>` and `POST /backup/delete/{where}/{name}` now refuse to delete a backup which other backups require via `required_backup` and report the dependent backup names, instead of silently breaking the incremental backups chain (the breakage surfaced only later, when a descendant was downloaded or restored, and for object disks the descendant `required` parts blobs were deleted together with the parent); pass `--force` (`force=1` for the API) to get the old behavior, or set `general.rebase_during_delete: true` (env `REBASE_DURING_DELETE`, default `false`) to rebase every dependent increment first (same as the `rebase` command) so the chain stays restorable — rebase copies the deleted backup parts into its dependents, so deletion time grows with the copied data size and a rebase failure aborts the delete. `backups_to_keep_local`/`backups_to_keep_remote` retention is not affected, fix [#1493](https://github.com/Altinity/clickhouse-backup/issues/1493)

IMPROVEMENTS
- fail fast instead of burning the whole `retries_on_failure` x `retries_duration` backoff budget per file when a remote object is permanently missing (S3 `NoSuchKey`/404, GCS 404, Azure `BlobNotFound`, FTP/SFTP not-found) during `download`, `restore_remote` and other retried remote operations; add `general.allow_missing_files_on_download` (env `ALLOW_MISSING_FILES_ON_DOWNLOAD`, default `false`), `--allow-missing-files` CLI flag for `download`/`restore_remote` and the `allow_missing_files` query argument for `POST /backup/download` and `POST /backup/restore_remote` — salvage mode which skips data parts missing on remote storage with an `error`-level log and a final summary, drops them from local table metadata and lets the intact tables/parts of a partially corrupted backup be restored, metadata files are never skipped, fix [#1456](https://github.com/Altinity/clickhouse-backup/issues/1456)
- add `s3.delete_batch_fallback_to_single` (env `S3_DELETE_BATCH_FALLBACK_TO_SINGLE`, default `true`) and `s3.delete_batch_min_size` (env `S3_DELETE_BATCH_MIN_SIZE`, default `0`) — when a whole `DeleteObjects` batch fails (some S3-compatible gateways such as DigitalOcean Spaces / Ceph RGW reset the response stream when the batch contains large objects, so retrying the same batch never succeeds and blocks retention and the `watch` loop), the batch is split in halves down to `delete_batch_min_size` and finally its objects are deleted one by one with `DeleteObject`, `s3.delete_concurrency` in parallel; `general.delete_batch_size` is now validated (1..1000 for `s3`), fix [#1532](https://github.com/Altinity/clickhouse-backup/issues/1532)
- `download --hardlink-exists-files` no longer does per-part filesystem and ClickHouse lookups, which dominated the download time on servers holding many local backups or many parts. Two changes: the shadow directories of the local backups are now indexed once per `download` run from their table metadata, so finding a hardlink candidate for a backup carrying legacy CRC64 `checksums` costs one map lookup plus one `stat` instead of a `filepath.Glob` over `<disk>/backup/*/shadow/...` (which ran twice per part, once for the free space check and once for the download itself); and the `hash_of_all_files` lookup in `system.parts` is now read once per table in chunks of 1000 hashes instead of one `SELECT` per part. Both paths keep their previous results: an indexed candidate is still verified by the CRC64 of its `checksums.txt` before being hardlinked, a local backup which can't be indexed (broken backup, unreadable metadata) marks the index incomplete and restores the old glob, and a `system.parts` candidate which a merge removed after the snapshot was taken is detected before hardlinking and re-resolved by a single live query for that part, fix [#1457](https://github.com/Altinity/clickhouse-backup/issues/1457)

Expand Down
5 changes: 5 additions & 0 deletions ReadMe.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,6 +132,7 @@ general:
# When `clickhouse->use_embedded_backup_restore: true`, throttling is delegated to the ClickHouse server via the `max_backup_bandwidth` query setting passed in the BACKUP/RESTORE SETTINGS clause (requires ClickHouse 25.1+); upload_max_bytes_per_second applies to BACKUP, download_max_bytes_per_second to RESTORE. On older ClickHouse versions embedded transfers are not throttled.
download_max_bytes_per_second: 0 # DOWNLOAD_MAX_BYTES_PER_SECOND, 0 means no throttling
upload_max_bytes_per_second: 0 # UPLOAD_MAX_BYTES_PER_SECOND, 0 means no throttling
allow_missing_files_on_download: false # ALLOW_MISSING_FILES_ON_DOWNLOAD, salvage mode for partially corrupted remote backups: `download` and `restore_remote` skip data parts whose files are missing on remote storage (404/NoSuchKey/BlobNotFound) with an `error` log and drop them from local table metadata instead of failing, metadata files are never skipped, requires `upload_by_part: true`, `--allow-missing-files` CLI argument or `allow_missing_files` API parameter overrides it per command, see https://github.com/Altinity/clickhouse-backup/issues/1456
download_disk_limit: 0 # DOWNLOAD_DISK_LIMIT, refuse `download` and `restore_remote` when usage of any local disk would exceed this percent (1..100) after download, 0 means no limit, `--disk-limit` CLI argument or `disk_limit` API parameter overrides it per command, see https://github.com/Altinity/clickhouse-backup/issues/1458
# MAX_BROKEN_PART_RATIO, maximum allowed fraction (0..1) of broken data parts (e.g. caused by S3-disk or filesystem failures) that still produces a successful but partial backup during backup creation (`create`, and the create stage of `create_remote`).
# 0 (default) preserves legacy behavior where any broken part stops the backup completely. When >0 and the broken/total part ratio stays at or below this value, creation skips the broken parts, logs a warning, and the backup is marked successful.
Expand Down Expand Up @@ -681,6 +682,7 @@ Download backup from remote storage: `curl -s localhost:7171/backup/download/<BA
only).
- Optional boolean query argument `resumable` works the same as the `--resumable` CLI argument (save intermediate download state and resume download if it already exists on local storage).
- Optional integer query argument `disk_limit` or `disk-limit` works the same as the `--disk-limit` CLI argument (refuse download when usage of any local disk would exceed this percent after download, 1-100).
- Optional boolean query argument `allow_missing_files` or `allow-missing-files` works the same as the `--allow-missing-files` CLI argument (skip data parts missing on remote storage instead of failing, overrides `general.allow_missing_files_on_download` for this request).
- Optional string query argument `callback` allow pass callback URL which will call with POST with `application/json` with payload `{"status":"error|success|cancel","error":"not empty when error happens", "operation_id" : "<random_uuid>", "command":"<full command line>", "duration":"<elapsed>"}`. When omitted or empty, falls back to `general.callback_url` if configured.

Note: this operation is asynchronous, so the API will return once the operation has started. The response includes an `operation_id` field that can be used to track the operation status via `/backup/status?operationid=<operation_id>`.
Expand Down Expand Up @@ -751,6 +753,7 @@ Download and restore data from remote backup: `curl -s localhost:7171/backup/res
- Optional boolean query argument `resume` works the same as the `--resume` CLI argument (resume download for object disk data).
- Optional boolean query argument `hardlink_exists_files` or `hardlink-exists-files` works the same as the `--hardlink-exists-files` CLI argument (Create hardlinks for existing files instead of downloading).
- Optional integer query argument `disk_limit` or `disk-limit` works the same as the `--disk-limit` CLI argument (refuse download when usage of any local disk would exceed this percent after download, 1-100).
- Optional boolean query argument `allow_missing_files` or `allow-missing-files` works the same as the `--allow-missing-files` CLI argument (skip data parts missing on remote storage instead of failing, overrides `general.allow_missing_files_on_download` for this request).
- Optional boolean query argument `streaming` works the same as the `--streaming` CLI argument (restore each table right after its download and delete its local copy, see [Streaming mode](#streaming-mode)).
- Optional boolean query argument `skip_empty_tables` or `skip-empty-tables` works the same as the `--skip-empty-tables` CLI argument (skip restoring tables that have no data).
- Optional boolean query argument `rebind_replica_path_if_exists` or `rebind-replica-path-if-exists` works the same as the `--rebind-replica-path-if-exists` CLI argument (overrides `clickhouse.rebind_replica_path_if_exists` for this request, rebind a restored ReplicatedMergeTree to `default_replica_path` when the original ZK path still has leftover state but our replica entry is absent). WARNING: never set during a concurrent HA multi-replica restore.
Expand Down Expand Up @@ -999,6 +1002,7 @@ OPTIONS:
--resume, --resumable Save intermediate download state and resume download if backup exists on local storage, ignored with 'remote_storage: custom' or 'use_embedded_backup_restore: true'
--hardlink-exists-files Create hardlinks for existing files instead of downloading
--disk-limit int Refuse download when usage of any local disk would exceed this percent (1-100) after download, overrides general->download_disk_limit, 0 means use config value, https://github.com/Altinity/clickhouse-backup/issues/1458 (default: 0)
--allow-missing-files Skip data part files which are missing on remote storage (404/NoSuchKey) with an error log and drop them from local table metadata instead of failing, salvage mode for partially corrupted backups, overrides general->allow_missing_files_on_download, https://github.com/Altinity/clickhouse-backup/issues/1456
--dry-run Show tables count and data size which would be downloaded, without downloading
--help, -h show help

Expand Down Expand Up @@ -1114,6 +1118,7 @@ OPTIONS:
--restore-schema-as-attach Use DETACH/ATTACH instead of DROP/CREATE for schema restoration
--hardlink-exists-files Create hardlinks for existing files instead of downloading
--disk-limit int Refuse download when usage of any local disk would exceed this percent (1-100) after download, overrides general->download_disk_limit, 0 means use config value, https://github.com/Altinity/clickhouse-backup/issues/1458 (default: 0)
--allow-missing-files Skip data part files which are missing on remote storage (404/NoSuchKey) with an error log and drop them from local table metadata instead of failing, salvage mode for partially corrupted backups, overrides general->allow_missing_files_on_download, https://github.com/Altinity/clickhouse-backup/issues/1456
--skip-empty-tables Skip restoring tables that have no data (empty tables with only schema)
--streaming Restore each table right after its download and delete its local copy, keeps only a small local footprint, https://github.com/Altinity/clickhouse-backup/issues/780
--rebind-replica-path-if-exists Override clickhouse.rebind_replica_path_if_exists, rebind a restored ReplicatedMergeTree to default_replica_path when the original ZK path still has leftover state but our replica entry is absent
Expand Down
10 changes: 10 additions & 0 deletions cmd/clickhouse-backup/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -527,6 +527,11 @@ func newRootCommand() *cli.Command {
Hidden: false,
Usage: "Refuse download when usage of any local disk would exceed this percent (1-100) after download, overrides general->download_disk_limit, 0 means use config value, https://github.com/Altinity/clickhouse-backup/issues/1458",
},
&cli.BoolFlag{
Name: "allow-missing-files",
Hidden: false,
Usage: "Skip data part files which are missing on remote storage (404/NoSuchKey) with an error log and drop them from local table metadata instead of failing, salvage mode for partially corrupted backups, overrides general->allow_missing_files_on_download, https://github.com/Altinity/clickhouse-backup/issues/1456",
},
&cli.BoolFlag{
Name: "dry-run",
Usage: "Show tables count and data size which would be downloaded, without downloading",
Expand Down Expand Up @@ -828,6 +833,11 @@ func newRootCommand() *cli.Command {
Hidden: false,
Usage: "Refuse download when usage of any local disk would exceed this percent (1-100) after download, overrides general->download_disk_limit, 0 means use config value, https://github.com/Altinity/clickhouse-backup/issues/1458",
},
&cli.BoolFlag{
Name: "allow-missing-files",
Hidden: false,
Usage: "Skip data part files which are missing on remote storage (404/NoSuchKey) with an error log and drop them from local table metadata instead of failing, salvage mode for partially corrupted backups, overrides general->allow_missing_files_on_download, https://github.com/Altinity/clickhouse-backup/issues/1456",
},
&cli.BoolFlag{
Name: "skip-empty-tables",
Hidden: false,
Expand Down
8 changes: 8 additions & 0 deletions pkg/backup/backuper.go
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ import (
"regexp"
"strings"
"sync"
"sync/atomic"

"github.com/Altinity/clickhouse-backup/v2/pkg/common"
"github.com/Altinity/clickhouse-backup/v2/pkg/metadata"
Expand Down Expand Up @@ -62,6 +63,8 @@ type Backuper struct {
// localPartIndex - read-only after build, maps parts of local backups to their shadow directories
// so `download --hardlink-exists-files` doesn't glob all local backups per part, see issues/1457
localPartIndex *localPartIndex
// skippedMissingParts - data parts skipped by allow_missing_files_on_download during the current download, see issues/1456
skippedMissingParts atomic.Uint64
}

func NewBackuper(cfg *config.Config, opts ...BackuperOpt) *Backuper {
Expand All @@ -87,6 +90,11 @@ func (b *Backuper) Classify(err error) retrier.Action {
if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) {
return retrier.Fail
}
// a missing remote object (404/NoSuchKey/BlobNotFound) will never heal, don't burn the retry budget on it,
// see https://github.com/Altinity/clickhouse-backup/issues/1456
if storage.IsNotFoundErr(err) {
return retrier.Fail
}
log.Warn().Err(err).Msgf("Will wait near %s and retry", common.AddRandomJitter(b.cfg.General.RetriesDuration, b.cfg.General.RetriesJitter))
return retrier.Retry
}
Expand Down
4 changes: 4 additions & 0 deletions pkg/backup/backuper_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,11 @@ func TestClassify(t *testing.T) {
{context.Canceled, retrier.Fail},
{context.DeadlineExceeded, retrier.Fail},
{fmt.Errorf("object_disk.CopyObject: %w", context.Canceled), retrier.Fail},
// https://github.com/Altinity/clickhouse-backup/issues/1456
{fmt.Errorf("DownloadCompressedStream StatFile: %w", storage.NewErrNotFound("shadow/default/t/default_all_1_1_0.tar")), retrier.Fail},
{&smithy.GenericAPIError{Code: "NoSuchKey", Message: "The specified key does not exist"}, retrier.Fail},
{errors.New("transient network error"), retrier.Retry},
{&smithy.GenericAPIError{Code: "SlowDown", Message: "Please reduce your request rate"}, retrier.Retry},
}
for _, tc := range testcases {
if got := b.Classify(tc.err); got != tc.expect {
Expand Down
Loading