Skip to content

Design proposal: tenant-supplied backup destination and options - #83

Open
Andrey Kolkov (androndo) wants to merge 2 commits into
mainfrom
design-proposal/tenant-backup-destination
Open

Andrey Kolkov (androndo) wants to merge 2 commits into
mainfrom
design-proposal/tenant-backup-destination

Conversation

@androndo

Copy link
Copy Markdown

Overview

Adds a design proposal (design-proposals/tenant-backup-destination/) for two tenant-writable additions to the backup API, symmetric to what the restore side already has:

  • a typed destination on Plan/BackupJob (a target the tenant owns; credentials referenced through a TenantSecret, never inlined);
  • an opaque options blob (*runtime.RawExtension, symmetric to RestoreJobSpec.Options) for driver-specific backup scope (topics/tables/prefixes) and mode (incremental).

Core does not interpret options; it validates destination structurally and passes both through to the strategy driver.

Why

The backup API is admin chooses the configuration, tenant picks a class — a tenant cannot name where a backup goes or what it contains. This is the gap surfaced by the S3 Bucket backup driver review in cozystack/cozystack#4235: a same-store bucket copy provides no durability, and the destination that makes it meaningful is one the tenant owns off-platform, which the API cannot express today.

Status: Draft — feedback on the two open questions (validation mechanism for the TenantSecret rule; shared per-driver options schema) especially welcome.

Adds a design proposal for a typed tenant destination and an opaque
driver-options blob on Plan/BackupJob, symmetric to RestoreJob.Options.
Motivated by the S3 Bucket backup driver (cozystack/cozystack#4235).

Signed-off-by: Andrey Kolkov <androndo@gmail.com>
Assisted-By: LLM
@coderabbitai

coderabbitai Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: ee4cb78c-8fda-4db3-97a9-b08dc7ea29c8

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@lllamnyp Timofei Larkin (lllamnyp) left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for picking this up. The need is real: a bucket backup that lands in the platform's own object store shares the source's failure domain, and today the API has no way to name anything else. I agree with the main calls here: a typed destination rather than one buried in options, credentials by reference only, and failing the run instead of silently falling back to the class default.

I'd like to propose a different shape for the destination half, though: bring back storageRef instead of an inline destination struct. I also have one hard constraint on the credentials half.

Background

The original API in cozystack/cozystack#1640 had a required storageRef (TypedLocalObjectReference) on Plan, BackupJob and Backup. The Plan copied it onto each BackupJob, the driver recorded it on the Backup, and restore read storage from backup.spec.storageRef. The design doc treated Storage as opaque to core ("drivers read Storage to know how/where to store or read artifacts"), with storage drivers pluggable and their kinds left TBD. cozystack/cozystack#1873 removed it in favour of BackupClass. That was the right call at the time: nothing implemented pluggable storage, so the field was cost without benefit. This proposal is exactly the point where it starts paying for itself.

Proposed shape

// PlanSpec / BackupJobSpec
// StorageRef optionally names where backups are written. When omitted, the
// storage configured by the resolved BackupClass strategy applies.
// +optional
StorageRef *corev1.TypedLocalObjectReference `json:"storageRef,omitempty"`

// BackupSpec / status: the storage actually used, recorded by the driver and
// immutable afterwards. Restore and cleanup resolve storage from here.
StorageRef corev1.TypedLocalObjectReference `json:"storageRef"`

The referenced object carries the destination, and its kind decides what that means. Two kinds would cover the motivating cases:

  1. apps.cozystack.io/Bucket, the tenant's own Bucket app. Coordinates and credentials come from the COSI BucketAccess the platform already provisions, so the tenant supplies no secret at all.
  2. An external S3 storage kind (e.g. backups.cozystack.io/S3Storage, namespaced) holding endpoint, bucket, region, prefix, an optional CA, and a reference to the credentials (see below).

options can stay as proposed. It's orthogonal.

Why a reference rather than an inline struct

The inline struct puts the S3 schema into the core backups.cozystack.io types. With it, core validates S3 endpoints, buckets and prefixes, and the proposal lists non-S3 destinations as a non-goal because each new backend would mean a core API change.

A typed reference keeps core out of storage semantics, the same way strategyRef (and now BackupClass) keeps it out of backup mechanics. Core only carries the reference from Plan to BackupJob to Backup. Each storage kind owns its own schema and validation. A new backend, whether a different object store or a third-party storage driver, is a new kind and leaves core untouched. That extensibility was the reason storageRef was a reference in #1640, and this proposal is the first time it has a second kind to separate.

Making the destination its own object also helps in practice:

  • Validate once, reuse many times. Endpoint policy, credential checks and a reachability probe run against one object with a Ready condition, not on every Plan and BackupJob, where failures would only show up at run time.
  • Explicit driver opt-in. Each strategy declares the storage kinds it supports, and admission rejects a Plan whose strategy can't honour the referenced kind. "The driver cannot honor the destination" becomes an admission error, not a failed run.
  • Stable anchor for restore and cleanup. Backup.storageRef points at a live object, not coordinates copied into status. A finalizer can block deleting the storage object while Backups still reference it. That gives the retention proposal (community#77), where drivers delete expired artifacts, a well-defined place to look, and a clear failure mode if the credentials have gone.
  • RBAC on a dedicated resource. "May this tenant send backups off-platform, and to where" becomes RBAC and admission on the storage kind. Admins get that control without editing BackupClasses.
  • The default path stays as simple as #1873 made it. No storageRef means the BackupClass storage, as today.

Credentials

Tenants are not going to get write access to TenantSecret. The write-only TenantSecret from an earlier revision of community#74 has been withdrawn, and community#82 is the path for tenant-supplied secrets. Under #82, a tenant keeps the secret in its own secret store, the spec references an entry by name, and a chart-rendered ExternalSecret materialises it for the consumer without the tenant holding any grant on Secrets.

The external S3 storage kind should take its credentials that way: a reference to an entry in the tenant's store, materialised into a Secret the tenant can't read. The "Tenants gain write access to TenantSecret" item and rollout step 2 should come out, and credentialsSecretRef should become a #82-style reference.

Open questions under either shape

These questions are the same for inline destination and storageRef:

  • Who makes the connection, and from where. CNPG and the Job-based strategies run in the application namespace, so the tenant's own egress position applies. Velero runs in cozy-velero with platform-wide privileges and would need a per-tenant BackupStorageLocation and credential Secret there. An unvetted tenant endpoint reached from that position is an SSRF surface. Error text propagated into status.message would also give a response channel back to the tenant. Endpoints need an admin-owned policy (allowed schemes, no private or link-local ranges), plus egress limits wherever the writer runs outside the tenant namespace.
  • Restore from tenant-writable storage. If the tenant can modify artifacts between backup and restore, the restore driver is applying tenant-controlled input. For Velero that means object manifests (the VM strategies include helmreleases and secrets) applied with Velero's privileges, which is more than tenant RBAC allows. I'd keep Velero-based strategies off tenant-writable storage until restore validates what it applies. The per-strategy opt-in above is the natural place to enforce that.
  • Secrets captured in backups. Velero backups of an app include its labelled Secrets. In tenant-owned storage these become tenant-readable, including any that are deliberately not exposed to the tenant.

Happy to help sketch the storage kinds in more detail if this direction works for you.

@androndo

Copy link
Copy Markdown
Author

Thanks — taking this direction. Agreed on the core moves: a storageRef reference over an inline destination struct (core stays out of storage semantics; a new backend is a new kind; Backup.storageRef as the restore/cleanup anchor, including community#77), options as-is, credentials via community#82 (dropping the TenantSecret-write item and rollout step 2), and a per-strategy opt-in deciding which strategies may target tenant-writable storage at all.

On the tampering open-question specifically, the split looks clean. The core API already has a trust anchor Velero lacks: Backup.status.artifact.checksum, written by the driver at backup time and stored in the Backup CR in-cluster, outside the tenant-writable store (unused today). For single-artifact strategies (Bucket, Job, CNPG, MongoDB) the driver records the snapshot checksum and restore verifies it before applying, so tampering is fail-closed. Velero verifies nothing on restore (velero-io/velero#9187) and applies with its own privileges, so Velero-based strategies stay off tenant-writable storage under the opt-in.

A few things I'd like to settle in this thread before reworking the doc — open under either shape:

  1. Storage-object mutability vs existing backups. Backup.storageRef pointing at a live object is the right anchor, but if that object's spec (endpoint/bucket/prefix/creds) is edited after a backup is taken, existing Backups resolve to different coordinates and restore breaks. A finalizer blocks deleting a referenced storage object but not editing it. We likely need the destination-defining fields immutable while referenced (CEL), or the resolved coordinates snapshotted into Backup.status.

  2. What "durability" the class default actually guarantees. There are two distinct durability stories, and only one needs the external kind: (a) immutability (S3 Object Lock/WORM) plus a separate failure domain on the platform's own backup store, and (b) a tenant-supplied off-platform destination. Worth being explicit about which the default relies on, so we don't imply off-platform durability where only immutability is in play.

  3. Integrity must be hash-on-write. The checksum has to be computed over the bytes the driver streams out, not re-read from the tenant-writable store afterwards — otherwise a parallel overwrite between write and re-read records as valid. Restore stays fail-closed either way, but it's a driver contract worth stating.

  4. Per-strategy support declaration. Where does a strategy declare which storage kinds it supports and that its restore is integrity-checked — a field on the strategy CR, or an admission-side table? That's what the opt-in admission check reads.

  5. Credentials sequencing. The external S3 kind's credentials depend on community#82 landing; the Bucket kind (COSI BucketAccess) does not. Worth sequencing so the Bucket-targeted path isn't gated on design-proposal: tenant-supplied secrets by reference through External Secrets Operator #82.

  6. Tenant-owned storage is not a durability SLA. The tenant owns the bucket and can delete artifacts, disable Object Lock, or rotate credentials; a deleted artifact makes restore fail (availability, not a silent wrong-apply). On storage the platform doesn't control we can promise failure-domain separation, not platform-enforced retention.

  7. options is tenant-supplied, driver-trusted. Core passes it through uninterpreted, so each driver must validate its own options (scope/prefixes) rather than assume well-formed input.

  8. Off-platform durability for privileged-apply strategies, if ever required. Should a later requirement call for Velero/VM backups on tenant-writable off-platform storage, it can't ride this same opt-in: restore there applies tenant-controlled manifests with the restore controller's privileges, so it needs restore-side validation (signed backups / verified-apply) first. Flagging it as a separate track, not a switch to flip on later.

Happy to fold the resolved answers into the doc and sketch the two storage kinds' schemas in the same pass.

…erence

Folds the review consensus into the proposal: the destination becomes a
storageRef to a typed storage object (apps.cozystack.io/Bucket or a new
backups.cozystack.io/S3Storage) rather than an inline S3 struct in the core
types, keeping core out of storage semantics and giving restore and cleanup a
live anchor. Credentials move to the community#82 model instead of a tenant
write grant on TenantSecret. Records the settled calls in Decisions —
reference-over-inline, #82 credentials, per-strategy integrity opt-in with
single-artifact checksum verification and Velero kept off tenant-writable
storage, and failure-domain separation rather than a platform durability SLA —
and narrows Open questions to storage-object mutability, the durability story
the default relies on, the per-strategy declaration mechanism, enforcement, and
off-platform privileged-apply as a separate track.

Assisted-by: LLM
Signed-off-by: Andrey Kolkov <androndo@gmail.com>
@androndo

Copy link
Copy Markdown
Author

Reworked the proposal along these lines and pushed (204c193):

  • The destination is now a storageRef to a typed storage kind — apps.cozystack.io/Bucket (COSI BucketAccess, no secret) or a new namespaced backups.cozystack.io/S3Storage — rather than an inline struct in the core types. The driver records Backup.storageRef (immutable) as the restore/cleanup anchor.
  • Credentials moved to the community#82 model; the TenantSecret write grant and its rollout step are gone.
  • Per-strategy opt-in with integrity: single-artifact strategies verify Backup.status.artifact.checksum on restore (fail-closed), and Velero-style strategies stay off tenant-writable storage.

Those are folded into a new Decisions section. I narrowed Open questions to what's still worth settling here:

  • Storage-object mutability vs existing backups — destination-defining fields immutable while referenced (CEL), or snapshot the coordinates onto Backup.
  • Which durability story the Bucket default relies on (immutability + a separate failure domain vs a tenant off-platform target).
  • How a strategy declares its supported storage kinds and whether its restore is integrity-checked.
  • Enforcement: CEL on the CRDs vs an admission webhook.
  • Off-platform durability for privileged-apply strategies (Velero/VM) — a separate track, since restore there needs validation of what it applies first.

Happy to sketch the two storage kinds' schemas next if the shape looks right.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants