zrepl

mirror of https://github.com/zrepl/zrepl.git synced 2024-11-22 00:13:52 +01:00

Author	SHA1	Message	Date
Christian Schwarz	def510abfd	chore: require go 1.22/1.23, upgrade protobuf, upgrade all deps Go upgrade: - Go 1.23 is current => use that for release builds - Go 1.22 is less than one year old, it's desirable to support it. - The [`Go Toolchains`](https://go.dev/doc/toolchain) stuff is available in both of these (would also be in Go 1.21). That is quite nice stuff, but required some changes to how we versions we use in CircleCI and the `release-docker` Makefile target. Protobuf upgrade: - Go to protobuf GH release website - Download latest locally - run `sha256sum` - replace existing pinned hashes - `make generate` Deps upgrade: - `go get -t -u all` - repository moves aren't handled well automatically, fix manually - repeat until no changes	2024-09-08 20:49:09 +00:00
Tercio Filho	2b3daaf9f1	zrepl status: hide progress bar once all filesystems reach terminal state (#674 ) * Added `IsTerminal` method * Made rendering of progress bar conditional based on IsTerminal	2023-05-02 19:28:56 +02:00
Christian Schwarz	a91fb873e4	fix incorrect use of sort.StringSlice A newer version of staticheck found these: > SA4029: sort.StringSlice is a type, not a function, and > sort.StringSlice(variants) doesn't sort your values; consider using > sort.Strings instead (staticcheck)	2022-10-24 22:22:41 +02:00
Christian Schwarz	c743c7b03f	refactor snapper & support cron-based snapshotting fixes https://github.com/zrepl/zrepl/issues/554 refs https://github.com/zrepl/zrepl/discussions/547#discussioncomment-1936126	2022-09-25 19:23:44 +02:00
Christian Schwarz	a9c61b4b0b	zrepl status UI: include `w` shortcut to wrap lines in help bar	2022-09-25 19:23:44 +02:00
Christian Schwarz	2d8c3692ec	rework resume token validation to allow resuming from raw sends of unencrypted datasets Before this change, resuming from an unencrypted dataset with send.raw=true specified wouldn't work with zrepl due to overly restrictive resume token checking. An initial PR to fix this was made in https://github.com/zrepl/zrepl/pull/503 but it didn't address the core of the problem. The core of the problem was that zrepl assumed that if a resume token contained `rawok=true, compressok=true`, the resulting send would be encrypted. But if the sender dataset was unencrypted, such a resume would actually result in an unencrypted send. Which could be totally legitimate but zrepl failed to recognize that. BACKGROUND ========== The following snippets of OpenZFS code are insightful regarding how the various ${X}ok values in the resume token are handled: - `6c3c5fcfbe/module/zfs/dmu_send.c (L1947-L2012)` - `6c3c5fcfbe/module/zfs/dmu_recv.c (L877-L891)` - https://github.com/openzfs/zfs/blob/6c3c5fc/lib/libzfs/libzfs_sendrecv.c#L1663-L1672 Basically, some zfs send flags make the DMU send code set some DMU send stream featureflags, although it's not a pure mapping, i.e, which DMU send stream flags are used depends somewhat on the dataset (e.g., is it encrypted or not, or, does it use zstd or not). Then, the receiver looks at some (but not all) feature flags and maps them to ${X}ok dataset zap attributes. These are funnelled back to the sender 1:1 through the resume_token. And the sender turns them into lzc flags. As an example, let's look at zfs send --raw. if the sender requests a raw send on an unencrypted dataset, the send stream (and hence the resume token) will not have the raw stream featureflag set, and hence the resume token will not have the rawok field set. Instead, it will have compressok, embedok, and depending on whether large blocks are present in the dataset, largeblockok set. WHAT'S ZREPL'S ROLE IN THIS? ============================ zrepl provides a virtual encrypted sendflag that is like `raw`, but further ensures that we only send encrypted datasets. For any other resume token stuff, it shoudn't do any checking, because it's a futile effort to keep up with ZFS send/recv features that are orthogonal to encryption. CHANGES MADE IN THIS COMMIT =========================== - Rip out a bunch of needless checking that zrepl would do during planning. These checks were there to give better error messages, but actually, the error messages created by the endpoint.Sender.Send RPC upon send args validation failure are good enough. - Add platformtests to validate all combinations of (Unencrypted/Encrypted FS) x (send.encrypted = true \| false) x (send.raw = true \| false) for cases both non-resuming and resuming send. Additional manual testing done: 1. With zrepl 0.5, setup with unencrypted dataset, send.raw=true specified, no send.encrypted specified. 2. Observe that regular non-resuming send works, but resuming doesn't work. 3. Upgrade zrepl to this change. 4. Observe that both regular and resuming send works. closes https://github.com/zrepl/zrepl/pull/613	2022-09-25 17:32:02 +02:00
Christian Schwarz	193abbe6b1	fix active child tasks panic with endpoint.ListAbstractionsStreamed The goroutine that does endTask() for "list-abstractions-streamed-producer" can be preempted after it has closed the out and outErrs channel, but before it calls endTask(). If the parent ("handler") then gets scheduled and and ends itself, it will observe an active child task "list-abstractions-streamed-producer". This is easy to demo by injecting a sleep here: --- a/endpoint/endpoint_zfs_abstraction.go +++ b/endpoint/endpoint_zfs_abstraction.go @@ -575,6 +576,7 @@ func ListAbstractionsStreamed(ctx context.Context, query ListZFSHoldsAndBookmark ctx, endTask := trace.WithTask(ctx, "list-abstractions-streamed-producer") go func() { defer endTask() + defer time.Sleep(10 * time.Second) defer close(out) defer close(outErrs) fixes https://github.com/zrepl/zrepl/issues/607	2022-07-17 21:44:03 +02:00
Cole Helbling	1df0f8912a	Add `--skip-cert-check` flag to `zrepl configcheck` to prevent checking cert files It may be desirable to check that a config is valid without checking for the existence of certificate files (e.g. when validating a config inside a sandbox without access to the cert files). This will be very useful for NixOS so that we can check the config file at nix-build time (e.g. potentially without proper permissions to read cert files for a TLS connection). fixes https://github.com/zrepl/zrepl/issues/467 closes https://github.com/zrepl/zrepl/pull/587	2022-07-08 20:18:41 +02:00
Christian Schwarz	e0c7ceedd5	prevent transient zrepl status error: Post "http://unix/status ": EOF See the comment added to client.go in this commit. fixes https://github.com/zrepl/zrepl/issues/483 fixes https://github.com/zrepl/zrepl/issues/262 fixes https://github.com/zrepl/zrepl/issues/379 fixes https://github.com/zrepl/zrepl/issues/379	2022-06-26 14:39:35 +02:00
Christian Schwarz	ce6701fb33	status: fix over-counted step when status != stepping This is a fixup of commit `b00b61e967` Author: Christian Schwarz <me@cschwarz.com> Date: Sun Nov 21 15:15:23 2021 +0100 status: user-visible replication step number should start at 1 fixes https://github.com/zrepl/zrepl/issues/589 refs https://github.com/zrepl/zrepl/issues/538	2022-04-24 15:24:39 +02:00
Christian Schwarz	c3f0041efd	zrepl test placeholder: fix panic if dataset does not exist fixes https://github.com/zrepl/zrepl/issues/406	2021-12-18 15:14:33 +01:00
Christian Schwarz	b00b61e967	status: user-visible replication step number should start at 1 fixes https://github.com/zrepl/zrepl/issues/538	2021-11-21 15:32:18 +01:00
Christian Schwarz	ac147b5a6f	replication: report a filesystem is active vs. blocked on something - `BlockedOn` prop in JSON report - Bring back the `*` in front of the filesystem report as an activity indicator. fixes https://github.com/zrepl/zrepl/issues/505	2021-11-14 17:34:32 +01:00
Christian Schwarz	4f9b63aa09	rework size estimation & dry sends - use control connection (gRPC) - use uint64 everywhere => fixes https://github.com/zrepl/zrepl/issues/463 - [BREAK] bump protocol version closes https://github.com/zrepl/zrepl/pull/518 fixes https://github.com/zrepl/zrepl/issues/463	2021-10-09 15:43:27 +02:00
Christian Schwarz	ad80bb3735	status: byteprogresshistory: disable averaging as workaround for #497 refs #497	2021-09-12 20:08:44 +02:00
Matthias Freund	bf1276f767	status: port status-v1 ETA calculation patch Must have forgotten to integrate it into the status-v2 branch at the time. refs https://github.com/zrepl/zrepl/issues/98#issuecomment-872154091 cc @dcdamien	2021-07-08 15:00:26 +02:00
InsanePrawn	ac4b109872	status/interactive: Revert to simple wakeup/reset signalling Signed-off-by: InsanePrawn <insane.prawny@gmail.com> closes #452	2021-03-25 22:26:17 +01:00
InsanePrawn	b2c6e51a43	client/signal: Revert "add signal 'snapshot', rename existing signal 'wakeup' to 'replication'" This was merged to master prematurely as the job components are not decoupled well enough for these signals to be useful yet. This reverts commit `2c8c2cfa14`. closes #452	2021-03-25 22:26:17 +01:00
Cole Helbling	1e85b1cb5f	client/status: allow raw mode without a tty A fairly common pattern when presented with JSON is to pipe it to `jq`. This also allows one to `zrepl status --mode raw > log` and operate on the JSON there. closes #442	2021-03-22 23:58:32 +01:00
Calistoc	ab7abd9686	status: show the current replication attempt's runtime	2021-03-14 18:24:28 +01:00
Christian Schwarz	a58ce74ed0	implement new 'zrepl status' Primary goals: - Scrollable output ( fixes #245 ) - Sending job signals from status view - Filtering of output by filesystem Implementation: - original TUI framework: github.com/rivo/tview - but: tview is quasi-unmaintained, didn't support some features - => use fork https://gitlab.com/tslocum/cview - however, don't buy into either too much to avoid lock-in - instead: port over the existing status UI drawing code and adjust it to produce strings instead of directly drawing into the termbox buffer Co-authored-by: Calistoc <calistoc@protonmail.com> Co-authored-by: InsanePrawn <insane.prawny@gmail.com> fixes #245 fixes #220	2021-03-14 18:24:25 +01:00
Calistoc	2c8c2cfa14	add signal 'snapshot', rename existing signal 'wakeup' to 'replication'	2021-03-14 18:16:23 +01:00
Christian Schwarz	61acc7494a	[#388 ] endpoint: fix incorrect use of trace.WithTaskGroup in ListAbstractionsStreamed Capturing behavior was broken.	2021-01-24 13:58:23 +01:00
Christian Schwarz	0a2dea05a9	[#385 ] status + replication: warning if replication succeeeded without any filesystem being replicated refs #385 refs #384	2020-11-01 13:51:28 +01:00
Christian Schwarz	4e702eedc9	cmd: zfs-abstraction list --json: fix panic (was panicking because `abstractions` is in fact a channel	2020-07-26 20:32:35 +02:00
Christian Schwarz	10a14a8c50	[#307 ] add package trace, integrate it with logging, and adopt it throughout zrepl package trace: - introduce the concept of tasks and spans, tracked as linked list within ctx - see package-level docs for an overview of the concepts - main feature 1: unique stack of task and span IDs - makes it easy to follow a series of log entries in concurrent code - main feature 2: ability to produce a chrome://tracing-compatible trace file - either via an env variable or a `zrepl pprof` subcommand - this is not a CPU profile, we already have go pprof for that - but it is very useful to visually inspect where the replication / snapshotter / pruner spends its time ( fixes #307 ) usage in package daemon/logging: - goal: every log entry should have a trace field with the ID stack from package trace - make `logging.GetLogger(ctx, Subsys)` the authoritative `logger.Logger` factory function - the context carries a linked list of injected fields which `logging.GetLogger` adds to the logger it returns - `logging.GetLogger` also uses package `trace` to get the task-and-span-stack and injects it into the returned logger's fields	2020-05-19 11:30:02 +02:00
Christian Schwarz	70f9c6482f	zfs: context propagation to ZFSListFilesystemVersions fixup of `9568e46f05`	2020-04-21 14:10:53 +02:00
Christian Schwarz	e0b5bd75f8	endpoint: refactor, fix stale holds on initial replication failure, zfs-abstractions subcmd, more efficient ZFS queries The motivation for this recatoring are based on two independent issues: - @JMoVS found that the changes merged as part of #259 slowed his OS X based installation down significantly. Analysis of the zfs command logging introduced in #296 showed that `zfs holds` took most of the execution time, and they pointed out that not all of those `zfs holds` invocations were actually necessary. I.e.: zrepl was inefficient about retrieving information from ZFS. - @InsanePrawn found that failures on initial replication would lead to step holds accumulating on the sending side, i.e. they would never be cleaned up in the HintMostRecentCommonAncestor RPC handler. That was because we only sent that RPC if there was a most recent common ancestor detected during replication planning. @InsanePrawn prototyped an implementation of a `zrepl zfs-abstractions release` command to mitigate the situation. As part of that development work and back-and-forth with @problame, it became evident that the abstractions that #259 built on top of zfs in package endpoint (step holds, replication cursor, last-received-hold), were not well-represented for re-use in the `zrepl zfs-abstractions release` subocommand prototype. This commit refactors package endpoint to address both of these issues: - endpoint abstractions now share an interface `Abstraction` that, among other things, provides a uniform `Destroy()` method. However, that method should not be destroyed directly but instead the package-level `BatchDestroy` function should be used in order to allow for a migration to zfs channel programs in the future. - endpoint now has a query facitilty (`ListAbstractions`) which is used to find on-disk - step holds and bookmarks - replication cursors (v1, v2) - last-received-holds By describing the query in a struct, we can centralized the retrieval of information via the ZFS CLI and only have to be clever once. We are "clever" in the following ways: - When asking for hold-based abstractions, we only run `zfs holds` on snapshot that have `userrefs` > 0 - To support this functionality, add field `UserRefs` to zfs.FilesystemVersion and retrieve it anywhere we retrieve zfs.FilesystemVersion from ZFS. - When asking only for bookmark-based abstractions, we only run `zfs list -t bookmark`, not with snapshots. - Currently unused (except for CLI) per-filesystem concurrent lookup - Option to only include abstractions with CreateTXG in a specified range - refactor `endpoint`'s various ZFS info retrieval methods to use `ListAbstractions` - rename the `zrepl holds list` command to `zrepl zfs-abstractions list` - make `zrepl zfs-abstractions list` consume endpoint.ListAbstractions - Add a `ListStale` method which, given a query template, lists stale holds and bookmarks. - it uses replication cursor has different modes - the new `zrepl zfs-abstractions release-{all,stale}` commands can be used to remove abstractions of package endpoint - Adjust HintMostRecentCommonAncestor RPC for stale-holds cleanup: - send it also if no most recent common ancestor exists between sender and receiver - have the sender clean up its abstractions when it receives the RPC with no most recent common ancestor, using `ListStale` - Due to changed semantics, bump the protocol version. - Adjust HintMostRecentCommonAncestor RPC for performance problems encountered by @JMoVS - by default, per (job,fs)-combination, only consider cleaning step holds in the createtxg range `[last replication cursor,conservatively-estimated-receive-side-version)` - this behavior ensures resumability at cost proportional to the time that replication was donw - however, as explained in a comment, we might leak holds if the zrepl daemon stops running - that trade-off is acceptable because in the presumably rare this might happen the user has two tools at their hand: - Tool 1: run `zrepl zfs-abstractions release-stale` - Tool 2: use env var `ZREPL_ENDPOINT_SENDER_HINT_MOST_RECENT_STEP_HOLD_CLEANUP_MODE` to adjust the lower bound of the createtxg range (search for it in the code). The env var can also be used to disable hold-cleanup on the send-side entirely. supersedes closes #293 supersedes closes #282 fixes #280 fixes #278 Additionaly, we fixed a couple of bugs: - zfs: fix half-nil error reporting of dataset-does-not-exist for ZFSListChan and ZFSBookmark - endpoint: Sender's `HintMostRecentCommonAncestor` handler would not check whether access to the specified filesystem was allowed.	2020-04-18 12:26:03 +02:00
Christian Schwarz	1336c91865	zfs: introduce pkg zfs/zfscmd for command logging, status, prometheus metrics refs #196	2020-04-05 20:47:25 +02:00
InsanePrawn	9568e46f05	zfs: use exec.CommandContext everywhere Co-authored-by: InsanePrawn <insane.prawny@gmail.com>	2020-03-27 13:08:43 +01:00
InsanePrawn	44bd354eae	Spellcheck all files Signed-off-by: InsanePrawn <insane.prawny@gmail.com>	2020-02-24 16:06:09 +01:00
Christian Schwarz	a3842155c5	zrepl test filesystems: support snap job type	2020-02-17 18:02:04 +01:00
Christian Schwarz	58c08c855f	new features: {resumable,encrypted,hold-protected} send-recv, last-received-hold - Resumable Send & Recv Support No knobs required, automatically used where supported. - Hold-Protected Send & Recv Automatic ZFS holds to ensure that we can always resume a replication step. - Encrypted Send & Recv Support for OpenZFS native encryption. Configurable at the job level, i.e., for all filesystems a job is responsible for. - Receive-side hold on last received dataset The counterpart to the replication cursor bookmark on the send-side. Ensures that incremental replication will always be possible between a sender and receiver. Design Doc ---------- `replication/design.md` doc describes how we use ZFS holds and bookmarks to ensure that a single replication step is always resumable. The replication algorithm described in the design doc introduces the notion of job IDs (please read the details on this design doc). We reuse the job names for job IDs and use `JobID` type to ensure that a job name can be embedded into hold tags, bookmark names, etc. This might BREAK CONFIG on upgrade. Protocol Version Bump --------------------- This commit makes backwards-incompatible changes to the replication/pdu protobufs. Thus, bump the version number used in the protocol handshake. Replication Cursor Format Change -------------------------------- The new replication cursor bookmark format is: `#zrepl_CURSOR_G_${this.GUID}_J_${jobid}` Including the GUID enables transaction-safe moving-forward of the cursor. Including the job id enables that multiple sending jobs can send the same filesystem without interfering. The `zrepl migrate replication-cursor:v1-v2` subcommand can be used to safely destroy old-format cursors once zrepl has created new-format cursors. Changes in This Commit ---------------------- - package zfs - infrastructure for holds - infrastructure for resume token decoding - implement a variant of OpenZFS's `entity_namecheck` and use it for validation in new code - ZFSSendArgs to specify a ZFS send operation - validation code protects against malicious resume tokens by checking that the token encodes the same send parameters that the send-side would use if no resume token were available (i.e. same filesystem, `fromguid`, `toguid`) - RecvOptions support for `recv -s` flag - convert a bunch of ZFS operations to be idempotent - achieved through more differentiated error message scraping / additional pre-/post-checks - package replication/pdu - add field for encryption to send request messages - add fields for resume handling to send & recv request messages - receive requests now contain `FilesystemVersion To` in addition to the filesystem into which the stream should be `recv`d into - can use `zfs recv $root_fs/$client_id/path/to/dataset@${To.Name}`, which enables additional validation after recv (i.e. whether `To.Guid` matched what we received in the stream) - used to set `last-received-hold` - package replication/logic - introduce `PlannerPolicy` struct, currently only used to configure whether encrypted sends should be requested from the sender - integrate encryption and resume token support into `Step` struct - package endpoint - move the concepts that endpoint builds on top of ZFS to a single file `endpoint/endpoint_zfs.go` - step-holds + step-bookmarks - last-received-hold - new replication cursor + old replication cursor compat code - adjust `endpoint/endpoint.go` handlers for - encryption - resumability - new replication cursor - last-received-hold - client subcommand `zrepl holds list`: list all holds and hold-like bookmarks that zrepl thinks belong to it - client subcommand `zrepl migrate replication-cursor:v1-v2`	2020-02-14 22:00:13 +01:00
Christian Schwarz	9a4763ceee	client/status: notify user if size estimation is imprecise There's plenty of room for improvement here. For example, detect if we're past the last step without size estimation and compute the remaining sum of bytes to be replicated from there on.	2020-02-14 21:42:03 +01:00
Matthias Freund	cca95f613b	client/status: add ETA calculation	2020-01-26 14:45:01 +01:00
Christian Schwarz	f7aa26d418	zrepl status: follow up `c4be60c`: import screen terminfo This is $TERM on FreeBSD and FreeNAS. fixes #204 ref https://github.com/gdamore/tcell/issues/252	2019-09-27 21:31:05 +02:00
Christian Schwarz	b5ff1a9926	snapper + client/status: snapshotting reports	2019-09-27 21:31:00 +02:00
Christian Schwarz	7ba3ae077f	client/status: job filter flag	2019-09-14 13:43:46 +02:00
John	0fc6a63564	Fix typo 'follwing'	2019-06-06 22:04:57 -07:00
Christian Schwarz	5b97953bfb	run golangci-lint and apply suggested fixes	2019-03-27 13:12:26 +01:00
Christian Schwarz	afed762774	format source tree using goimports	2019-03-22 19:41:12 +01:00
Christian Schwarz	2f2e6e6a00	receiving side: placeholder as simple on\|off property	2019-03-20 20:26:30 +01:00
Christian Schwarz	17818439a0	Merge branch 'problame/replication_refactor' into InsanePrawn-master	2019-03-17 17:33:51 +01:00
Christian Schwarz	7584c66bdb	pruner: remove retry handling + fix early give-up Retry handling is broken since the gRPC changes (wrong error classification). Will come back at some point, hopefully by merging the replication driver retry infrastructure. However, the simpler architecture allows an easy fix for the problem that the pruner practically gave up on the first error it encountered. fixes #123	2019-03-13 21:04:39 +01:00
Christian Schwarz	d78d20e2d0	pruner: skip placeholders + FSes without correspondents on source fixes #126	2019-03-13 20:42:37 +01:00
Christian Schwarz	d5250bbf51	client/status: fix wrap for multiline strings with leading space	2019-03-13 18:46:04 +01:00
Christian Schwarz	c87759affe	replication/driver: automatic retries on connectivity-related errors	2019-03-13 15:00:40 +01:00
Christian Schwarz	07b43bffa4	replication: refactor driving logic (no more explicit state machine)	2019-03-13 15:00:40 +01:00
InsanePrawn	c4e23862cd	Added status view for SnapJob.	2018-11-21 04:06:13 +01:00
Christian Schwarz	2db3977408	cli: add 'test placeholder' subcommand for placeholder debugging	2018-11-16 12:21:54 +01:00

1 2

90 Commits