What is this about
I would like to see erasure coding supported in Cozystack's storage stack — for the SeaweedFS (S3) application in particular, and for blockstor as a longer-term goal. Today the stack is replication-only everywhere, and at large scale (5–10 PB clusters) that replication overhead becomes a serious cost driver.
Context
I am planning/operating in the 5–10 PB range, where the current Cozystack redundancy model is expensive:
- SeaweedFS app: the application schema only exposes
replicationFactor (default 2), which is mapped to SeaweedFS replica placement (defaultReplicaPlacement). There is no erasure-coding option anywhere in the app (packages/system/seaweedfs-rd, packages/extra/seaweedfs), even though upstream SeaweedFS supports EC and even has maintenance workers for EC tasks. At replication factor 2, storing 5 PB of usable data requires ~10 PB raw.
- Block storage: LINSTOR/DRBD (and blockstor so far) is synchronous replication only, i.e. 2–3x overhead. LINBIT itself never shipped erasure coding (the 2019 DRBD EC work remained a proof of concept), so this cannot be solved upstream today.
I checked before opening this: a grep -ri erasure over the cozystack repository only matches upstream SeaweedFS chart documentation and a changelog note about an upstream SeaweedFS EC hazard — EC is neither enabled, configurable, nor documented anywhere. The cozystack.io docs likewise mention no erasure-coding capability.
For a 5–10 PB deployment the difference is substantial: an EC scheme with ~1.3–1.5x overhead (e.g. SeaweedFS erasure coding, or an EC-style backend in blockstor later) instead of 2–3x replication saves a large fraction of the raw hardware — potentially millions of euros at that scale — while still tolerating disk/node failures. This matters most for the object-storage tier, where EC is a natural fit (large objects, read-heavy, RGW/S3-style workloads).
What would help
- A statement on whether/when erasure coding is on the roadmap, specifically:
- exposing SeaweedFS EC through the Cozystack SeaweedFS application (EC shards per volume server, EC maintenance workers, placement policy), and
- whether space-efficient redundancy (EC or similar) is considered for blockstor beyond DRBD replication.
- Short term: documentation or guidance on whether enabling upstream SeaweedFS EC on a Cozystack-managed instance is supported/safe (and if not, why — e.g. because of the S3 correctness constraints mentioned in the v1.5.0 SeaweedFS bump).
- If this is worth a design discussion, I'm happy to contribute real-world requirements from a 5–10 PB deployment and help test.
What is this about
I would like to see erasure coding supported in Cozystack's storage stack — for the SeaweedFS (S3) application in particular, and for blockstor as a longer-term goal. Today the stack is replication-only everywhere, and at large scale (5–10 PB clusters) that replication overhead becomes a serious cost driver.
Context
I am planning/operating in the 5–10 PB range, where the current Cozystack redundancy model is expensive:
replicationFactor(default 2), which is mapped to SeaweedFS replica placement (defaultReplicaPlacement). There is no erasure-coding option anywhere in the app (packages/system/seaweedfs-rd,packages/extra/seaweedfs), even though upstream SeaweedFS supports EC and even has maintenance workers for EC tasks. At replication factor 2, storing 5 PB of usable data requires ~10 PB raw.I checked before opening this: a
grep -ri erasureover the cozystack repository only matches upstream SeaweedFS chart documentation and a changelog note about an upstream SeaweedFS EC hazard — EC is neither enabled, configurable, nor documented anywhere. The cozystack.io docs likewise mention no erasure-coding capability.For a 5–10 PB deployment the difference is substantial: an EC scheme with ~1.3–1.5x overhead (e.g. SeaweedFS erasure coding, or an EC-style backend in blockstor later) instead of 2–3x replication saves a large fraction of the raw hardware — potentially millions of euros at that scale — while still tolerating disk/node failures. This matters most for the object-storage tier, where EC is a natural fit (large objects, read-heavy, RGW/S3-style workloads).
What would help