RustFS Unified Storage — Complete Design & Development Plan (EN)
RustFS Unified Storage — Complete Design & Development Plan (EN)
Document: RUSTFS-PLAN-001 (master execution plan)
Status: DRAFT
Version: 1.0.0
Canonical design: RUSTFS-MD-DESIGN-005 (greenfield, no migration)
Supporting: FEAT-001, LLD-001 (I/O stack), LLD-002 (random read),
MD-DESIGN-001/002 (metadata), MD-DESIGN-003 (xl.meta — retired)
Language: English (中文镜像见 rustfs-dev-plan-ZH.md)
1. Project Summary
1.1 Goal
Build a cloud-native unified storage platform on a greenfield codebase that serves the same data through one namespace as:
- File (POSIX) — high-bandwidth, low-latency, optimized for AI "one write, multi fseek/fread" workloads.
- Object (S3) — full S3-compatible API.
- Block — VM / bare-metal disks over iSCSI and NVMe-oF.
1.2 Architecture (final, from DESIGN-005)
Stateless MDS + Transactional KV + Formula-driven layout
Data plane runs client ↔ chunk directly (MDS never on the data path)
Chunk service reuses proven RS / bitrot / disk / heal as libraries
No migration, no xl.meta, no coexistence — the system only sees its own format
1.3 Non-Goals (v1)
✗ Deduplication (deferred — needs a global content index; breaks simplicity)
✗ Backward compatibility with any prior on-disk format (greenfield)
✗ Multi-region active-active (later; v1 is single-region + async replication)
2. Component / Crate Map
Crate / module Responsibility Track
──────────────────── ────────────────────────────────────── ─────
rustfs-kv TransactionalKv trait; embedded A (metadata)
(redb/RocksDB) + TiKV backends
rustfs-inode InodeAttr, Layout, DentValue, Volume, A
ExtentEntry, key codecs
rustfs-placement HRW placement, ClusterMap, offset→chunk A (shared by A+C)
math (shared between MDS and client)
rustfs-mds Stateless metadata service; FS-semantics A
transactions; inode-range allocator; gRPC
chunk-format On-disk chunk format spec + codec B (data)
(self-verifying inline bitrot)
chunk-store Fresh ChunkStore: chunk-key addressing, B
reuses RS/bitrot/disk/heal libs
packing Small-file packing: container mgr B
(append/seal), Slice pointers, pack/
unpack, compaction GC. Containers are
large internal files → reuse chunk-store.
ecstore-core (reuse) Reed-Solomon encode/decode B (existing)
bitrot (reuse) Streaming HighwayHash bitrot B (existing)
disk-io (reuse/new) io_uring disk engine + buffer pool B
heal (reuse) Reconstruction math B (existing)
rustfs-client Client core: placement, cluster-map C (clients)
cache, metadata cache, retry/hedge,
circuit breaker, chunk transport
rustfs-fuse FUSE mount; read-ahead (sequential) + C
random-read path; lease-free caching
s3-frontend S3 API over MDS + ChunkStore C
block-frontend Volume mgr (EXTM, thin/COW), iSCSI, C
NVMe-oF targets
rustfs-rdma RDMA data-plane transport D (platform)
obs / metrics (reuse) OpenTelemetry + Prometheus D
test-harness Conformance + chaos + bench rigs D
3. Development Tracks (parallelizable)
Track A — Metadata rustfs-kv, inode, placement, mds
Track B — Data chunk-format, chunk-store, packing (+ reuse RS/bitrot/disk/heal)
Track C — Clients client, fuse, s3-frontend, block-frontend
Track D — Platform rdma, observability, HA, testing, benchmarks
After Phase 0, A and B run in parallel; C starts when A+B expose stable APIs; D's testing rig runs continuously from Phase 1, with perf/RDMA later.
4. Phased Plan, Milestones, Acceptance
Phase 0 — Design Freeze & Scaffolding (M0)
Resolve open questions (DESIGN-005 §6):
Q1 KV backend: embedded redb/RocksDB first; TiKV for HA.
Q2 inline_threshold (default 16 KiB; per-backend caps).
Q3 On-disk chunk format spec (stripe_unit, inline-bitrot block, chunk-key path).
Q4 Default storage policies (EC k+m for datasets; replication for block/small).
Q5 ChunkStore = fresh thin store (confirmed), reusing RS/bitrot/disk/heal.
Set up:
· Monorepo crate skeletons, CI (build/test/lint/coverage), feature flags
· Protobuf wire contracts (MDS gRPC, ChunkStore gRPC, GetClusterMap)
· chunk-format spec doc + golden test vectors
· Test-harness skeleton (conformance/chaos/bench placeholders)
M0 ACCEPTANCE: all crate skeletons build in CI; protobuf + chunk-format specs
reviewed and frozen; Q1–Q5 decided and recorded.
Effort: ~1 month.
Phase 1 — Metadata Core (M1) [Track A]
Deliverables:
· rustfs-kv: TransactionalKv trait + embedded backend; SSI semantics; CAS
· rustfs-inode: all types + key codecs + (de)serialization
· rustfs-placement: HRW over cluster map; offset→chunk; deterministic dist/path
· rustfs-mds: inode-range allocator; create/lookup/getattr/setattr/mkdir/
rmdir/readdir/rename/unlink/link/symlink/xattr as single KV transactions;
OpenLayout, GetClusterMap, CommitWrite, volume RPCs (stubs ok); gRPC server
· VIDX versioning rows; PEND/DELQ rows defined
M1 ACCEPTANCE:
· All metadata ops pass unit + integration tests against embedded KV
· Cross-directory rename is atomic (transaction test)
· readdir is a single range scan; lookup/getattr single get (perf assertion)
· Crash test: kill MDS mid-op → KV consistent on restart
Effort: ~2.5 months.
Phase 2 — Data Plane / Chunk Service (M2) [Track B]
Deliverables:
· chunk-format: self-verifying on-disk format; codec; golden vectors pass
· disk-io: io_uring engine (per-disk rings, fixed buffers) + tokio::fs fallback
· chunk-store: write_chunk/read_chunk/delete_chunk (chunk-key, byte-range);
RS encode/decode reused; inline bitrot verify on read; heal reconstruction
on corrupt/missing shard
· packing: container manager (open/append/seal), Slice read path, per-container
CMTA (dead_bytes), compaction GC; containers are large internal files reusing
chunk-store + formula layout
· placement integration: chunk_set/path/dist computed identically to MDS
· PEND write-intent on new file_id; GC scrubber for orphans; DELQ chunk GC
M2 ACCEPTANCE:
· Write a chunk, read arbitrary byte ranges back, bit-exact
· Corrupt a shard on disk → read still succeeds via reconstruction; heal repairs
· A client and the MDS compute identical locations for 10^6 random keys
· EC partial read amplification < 1.2x for 4 KiB reads (LLD-002 target)
· Pack 10^6 small files into containers; read each back via Slice; compaction
reclaims space after deletes (no KV bulk data; metadata-only KV verified)
Effort: ~3.5 months (overlaps Phase 1 tail).
Phase 3 — Client + FUSE (POSIX MVP) (M3) ★ key usable milestone [Track C]
Deliverables:
· rustfs-client: placement engine, cluster-map cache, metadata cache
(version-based, close-to-open), retry/backoff, hedged reads, circuit breaker,
chunk transport (gRPC first, io_uring TCP next)
· rustfs-fuse: open/read/write/seek/close/stat/readdir/mkdir/rename/unlink/
truncate/fsync; inline-small-file fast path
· read-ahead engine (sequential + stride detection)
· random-read path: full-chunkmap-on-open, fragment cache (TinyLFU),
high-QD path; posix_fadvise honored
· crash consistency: data-first then CommitWrite; PEND lifecycle
M3 ACCEPTANCE:
· mount; dd 1 GiB write + md5 read-back; ls -la; cp; rename; rm
· "one write, multi fseek/fread": fio randread > target IOPS; p99 within target
· Sequential read approaches aggregate DSS bandwidth (read-ahead working)
· A multi-process reader (simulated DataLoader) keeps reads off the MDS
after open (lease-free version cache holds)
Effort: ~3 months.
Phase 4 — S3 Unified Namespace (M4) [Track C]
Deliverables:
· s3-frontend: PutObject/GetObject/DeleteObject/ListObjects(V2)/HeadObject;
multipart (stage parts → formula chunks on Complete); range GET; conditional
requests; object tagging; bucket ops
· bucket↔directory mapping (/buckets/<bucket>/<key>); ListObjects = readdir
· versioning via VIDX
M4 ACCEPTANCE:
· s3-tests conformance suite passes (core + multipart + versioning subset)
· Data PUT via S3 is readable via the FUSE mount and vice versa (same inode)
· ListObjects on a 10^7-object bucket is O(1)-per-directory (latency proof)
Effort: ~2 months.
Phase 5 — Block Storage (M5) [Track C]
Deliverables:
· block-frontend volume mgr: EXTM extent map; thin provisioning; COW snapshots;
clone; resize
· NBD server (simplest, for testing) → iSCSI target (production) → NVMe-oF
· WAL for write ordering; SYNCHRONIZE_CACHE/flush semantics; UNMAP/TRIM
M5 ACCEPTANCE:
· Create volume; attach via iSCSI; boot a QEMU/KVM VM from it
· Snapshot + clone + resize verified; data consistent after VM fsync
· 4 KiB random write latency within target (LLD-001)
Effort: ~3 months.
Phase 6 — HA & Scale-Out (M6) [Track A + D]
Deliverables:
· TiKV backend for rustfs-kv (clustered, distributed transactions)
· Multiple stateless MDS behind LB / client round-robin; session failover
· Cluster-map epoch changes → formula relocation → online rebalancing
(copy-then-switch, rate-limited); node add/decommission; heal on node loss
M6 ACCEPTANCE:
· Kill an MDS instance under load → clients fail over, no errors surfaced
· Add a node → data rebalances online with < 10% client I/O impact
· KV-backed metadata ops scale ~linearly with MDS instance count (mdtest)
Effort: ~2.5 months.
Phase 7 — Performance & Hardening (M7) [Track D]
Deliverables:
· RDMA data-plane transport (READ pull for reads, WRITE push for writes);
transport negotiation (RDMA → io_uring TCP → gRPC)
· io_uring tuning: SEND_ZC, IOPOLL random-read rings, NUMA pinning
· random-read: access-plan hinting API + framework shim; DSS hot-fragment tier
· FUSE-over-io_uring (kernel ≥ 6.14) with fallback
· (optional) GPUDirect Storage path
M7 ACCEPTANCE (targets from LLD-001/002):
· 4 KiB random read p50 < 15 µs (RDMA), p99 < 50 µs
· Random read > 500 K IOPS/client (RDMA, EC partial)
· Shuffled-epoch training with access plan: > 5x vs no plan; GPU util > 90%
· Sequential aggregate bandwidth scales linearly with DSS nodes
Effort: ~3.5 months.
Phase 8 — Production Readiness (M8 / GA) [Track D]
Deliverables:
· Observability: metrics, tracing, dashboards for MDS/ChunkStore/client
· IAM/policy integration; encryption at rest (reuse); audit
· Ops tooling: cluster bootstrap, node lifecycle, backup/restore of KV
· Kubernetes CSI driver (RWX FUSE, RWO block)
· Soak (30-day) + chaos + Jepsen-style consistency tests green
· Documentation (admin, API, tuning) — bilingual
M8 ACCEPTANCE: production readiness review passed; GA criteria met.
Effort: ~3 months (overlaps Phase 7).
5. Milestone Timeline (rough, team of ~8–10)
Quarter Milestones in flight Usable outcome
──────── ───────────────────────────────────────── ────────────────────────
Q1 M0 freeze; M1 metadata core; M2 start —
Q2 M2 chunk done; M3 FUSE MVP ★ POSIX usable (fseek/fread)
Q3 M4 S3; M5 block start unified file+object
Q4 M5 block done; M6 HA block + clustered HA
Q5 M7 performance hits AI perf targets
Q6 M8 production readiness GA
Critical path to POSIX MVP: M0 → M1 → M2 → M3 ≈ 7–9 months
Critical path to GA: ≈ 18–24 months
Parallelism: Track A and B run concurrently after M0; Track C (FUSE) starts mid Phase 2 against mocked then real APIs; Track D testing runs from M1 onward.
6. Dependency Graph (build order)
┌── rustfs-kv ──┐
M0 ──────┤ ├── rustfs-mds ──┐
└── rustfs-inode┘ │
│ ├── rustfs-client ── rustfs-fuse (M3)
rustfs-placement (shared) ───────┤ ├── s3-frontend (M4)
│ └── block-frontend (M5)
┌── chunk-format ──┐ │
M0 ──────┤ ├── chunk-store ──┘
└── disk-io ───────┘ (reuses RS/bitrot/heal)
│
rustfs-rdma ─────┴── perf (M7)
TiKV backend ──────────────────────── HA (M6)
7. Testing & Validation Strategy
Level What When
───────────── ───────────────────────────────────────── ──────────────
Unit per-crate logic, codecs, transactions every phase, CI gate
Integration MDS + ChunkStore + client end-to-end from M2
Conformance POSIX (pjdfstest), S3 (s3-tests), M3 / M4 / M5
block (fio, libiscsi/SCSI compliance)
Chaos kill MDS, kill chunk node, partition, from M2; gate for M6
crash mid-write (PEND/GC correctness)
Consistency Jepsen-style on KV txns + rename/create M1 + M6
Performance fio (IOPS/latency), IOR/mdtest (parallel), M3 baseline; M7 targets
real GPU training run (utilization)
Soak 30-day continuous load + fault injection M8
Regression CI perf + correctness gates per merge continuous
8. Risk Register
Risk Mitigation
──────────────────────────────────── ──────────────────────────────────────────
FUSE overhead caps POSIX throughput FUSE-over-io_uring (M7); native client
later; measure early in M3
io_uring kernel-version fragmentation trait-abstracted disk-io with tokio::fs
fallback; feature-detect
RDMA absent on commodity clusters transport negotiation → io_uring TCP
KV (metadata) becomes the bottleneck version-based client cache; bulk file
data NEVER in KV (packing, §M2); TiKV
sharding; benchmark in M1/M6
Billions of small files (AI datasets, packing into large containers (M2); KV
node_modules, pip/conda, build tmp) holds metadata + Slice pointers only;
client-side write batching; raise inline
cap only for ≤~1 KB; compaction GC sized
for high churn (temp/dependency trees)
Formula placement vs rebalancing cluster-map epochs + HRW (minimal movement);
copy-then-switch online rebalance (M6)
Crash window between data & KV commit PEND write-intent + GC scrubber; verified
in M2/M3 chaos tests
Chunk on-disk format wrong → costly freeze format in M0 with golden vectors;
version the format header for future
ChunkStore fresh-build scope creep reuse RS/bitrot/disk/heal as libs strictly;
keep the wrapper thin (M2 scope guard)
EC partial-write amplification (block) policy routing: replication for block/
random-write; EC for write-once datasets
Packing compaction can lag under heavy size compactor throughput to churn; per-
delete/temp-file churn container dead_bytes threshold triggers;
rate-limit vs client I/O (like rebalance)
9. Definition of Done (per milestone)
M0 specs frozen, CI green on skeletons, Q1–Q6 recorded
M1 metadata ops correct + atomic + crash-safe on embedded KV
M2 chunk read/write bit-exact, self-verifying, heal-on-corrupt, formula parity;
small-file packing (pack/read/compact) works; KV proven metadata-only
M3 POSIX mount serves the fseek/fread workload at target IOPS/latency
M4 S3 conformance + bidirectional file/object visibility + O(1) listing
M5 VM boots from block volume; snapshot/clone/resize correct
M6 MDS failover + online rebalance + linear metadata scaling
M7 AI perf targets met (RDMA latency, IOPS, planned-shuffle throughput)
M8 observability + CSI + soak/chaos/consistency green; PRR passed → GA
10. Team Shape (suggested, ~8–10 engineers)
2 Metadata (Track A): KV abstraction, MDS, placement
2 Data (Track B): chunk-format, chunk-store, packing, disk-io/io_uring
2–3 Clients (Track C): FUSE, S3, block frontends
1–2 Platform (Track D): RDMA, observability, HA, CSI
1 Test/Perf engineering (harness, conformance, chaos, benchmarks) — from day 1
1 Tech lead / architect (owns specs, cross-track APIs, design freeze)
11. First Actions (next 2 weeks)
1. Freeze the on-disk chunk format (Q3): write chunk-format spec + golden vectors.
2. Lock the protobuf wire contracts (MDS + ChunkStore + GetClusterMap).
3. Decide KV: stand up embedded redb/RocksDB behind TransactionalKv; spike TiKV.
4. Stand up the monorepo crate skeletons + CI + the test-harness shell.
5. Spike rustfs-placement (HRW + cluster map) and prove client/MDS parity on a
unit test — it underpins the whole "formula" thesis and must be nailed early.
12. Summary
A greenfield, single-model unified storage: stateless MDS over a transactional
KV (metadata only — never bulk data), pure formula layout, client-direct data
plane, a thin chunk store reusing the proven RS/bitrot/disk/heal code, and
first-class small-file packing for the billions of small files in AI datasets and
code/dependency trees. Built in 8 milestones across 4 parallel tracks; POSIX MVP
in ~7–9 months, GA in ~18–24 months. No migration, no coexistence, no legacy
format — ever.
Chinese mirror: rustfs-dev-plan-ZH.md. Executes the canonical design RUSTFS-MD-DESIGN-005 with support from FEAT-001, LLD-001, LLD-002.