RustFS 文档 文档

RustFS Unified Storage — Complete Design & Development Plan (EN)

RustFS Unified Storage — Complete Design & Development Plan (EN)

Document:   RUSTFS-PLAN-001  (master execution plan)
Status:     DRAFT
Version:    1.0.0
Canonical design:  RUSTFS-MD-DESIGN-005 (greenfield, no migration)
Supporting:        FEAT-001, LLD-001 (I/O stack), LLD-002 (random read),
                   MD-DESIGN-001/002 (metadata), MD-DESIGN-003 (xl.meta — retired)
Language:   English  (中文镜像见 rustfs-dev-plan-ZH.md)

1. Project Summary

1.1 Goal

Build a cloud-native unified storage platform on a greenfield codebase that serves the same data through one namespace as:

  • File (POSIX) — high-bandwidth, low-latency, optimized for AI "one write, multi fseek/fread" workloads.
  • Object (S3) — full S3-compatible API.
  • Block — VM / bare-metal disks over iSCSI and NVMe-oF.

1.2 Architecture (final, from DESIGN-005)

Stateless MDS  +  Transactional KV  +  Formula-driven layout
Data plane runs client ↔ chunk directly (MDS never on the data path)
Chunk service reuses proven RS / bitrot / disk / heal as libraries
No migration, no xl.meta, no coexistence — the system only sees its own format

1.3 Non-Goals (v1)

✗ Deduplication (deferred — needs a global content index; breaks simplicity)
✗ Backward compatibility with any prior on-disk format (greenfield)
✗ Multi-region active-active (later; v1 is single-region + async replication)

2. Component / Crate Map

Crate / module        Responsibility                          Track
────────────────────  ──────────────────────────────────────  ─────
rustfs-kv             TransactionalKv trait; embedded          A (metadata)
                      (redb/RocksDB) + TiKV backends
rustfs-inode          InodeAttr, Layout, DentValue, Volume,    A
                      ExtentEntry, key codecs
rustfs-placement      HRW placement, ClusterMap, offset→chunk  A (shared by A+C)
                      math (shared between MDS and client)
rustfs-mds            Stateless metadata service; FS-semantics A
                      transactions; inode-range allocator; gRPC
chunk-format          On-disk chunk format spec + codec        B (data)
                      (self-verifying inline bitrot)
chunk-store           Fresh ChunkStore: chunk-key addressing,  B
                      reuses RS/bitrot/disk/heal libs
packing               Small-file packing: container mgr        B
                      (append/seal), Slice pointers, pack/
                      unpack, compaction GC. Containers are
                      large internal files → reuse chunk-store.
ecstore-core (reuse)  Reed-Solomon encode/decode               B (existing)
bitrot (reuse)        Streaming HighwayHash bitrot              B (existing)
disk-io (reuse/new)   io_uring disk engine + buffer pool       B
heal (reuse)          Reconstruction math                      B (existing)
rustfs-client         Client core: placement, cluster-map      C (clients)
                      cache, metadata cache, retry/hedge,
                      circuit breaker, chunk transport
rustfs-fuse           FUSE mount; read-ahead (sequential) +    C
                      random-read path; lease-free caching
s3-frontend           S3 API over MDS + ChunkStore             C
block-frontend        Volume mgr (EXTM, thin/COW), iSCSI,      C
                      NVMe-oF targets
rustfs-rdma           RDMA data-plane transport                D (platform)
obs / metrics (reuse) OpenTelemetry + Prometheus               D
test-harness          Conformance + chaos + bench rigs         D

3. Development Tracks (parallelizable)

Track A — Metadata    rustfs-kv, inode, placement, mds
Track B — Data        chunk-format, chunk-store, packing (+ reuse RS/bitrot/disk/heal)
Track C — Clients     client, fuse, s3-frontend, block-frontend
Track D — Platform    rdma, observability, HA, testing, benchmarks

After Phase 0, A and B run in parallel; C starts when A+B expose stable APIs; D's testing rig runs continuously from Phase 1, with perf/RDMA later.


4. Phased Plan, Milestones, Acceptance

Phase 0 — Design Freeze & Scaffolding (M0)

Resolve open questions (DESIGN-005 §6):
  Q1 KV backend: embedded redb/RocksDB first; TiKV for HA.
  Q2 inline_threshold (default 16 KiB; per-backend caps).
  Q3 On-disk chunk format spec (stripe_unit, inline-bitrot block, chunk-key path).
  Q4 Default storage policies (EC k+m for datasets; replication for block/small).
  Q5 ChunkStore = fresh thin store (confirmed), reusing RS/bitrot/disk/heal.

Set up:
  · Monorepo crate skeletons, CI (build/test/lint/coverage), feature flags
  · Protobuf wire contracts (MDS gRPC, ChunkStore gRPC, GetClusterMap)
  · chunk-format spec doc + golden test vectors
  · Test-harness skeleton (conformance/chaos/bench placeholders)

M0 ACCEPTANCE: all crate skeletons build in CI; protobuf + chunk-format specs
               reviewed and frozen; Q1–Q5 decided and recorded.
Effort: ~1 month.

Phase 1 — Metadata Core (M1) [Track A]

Deliverables:
  · rustfs-kv: TransactionalKv trait + embedded backend; SSI semantics; CAS
  · rustfs-inode: all types + key codecs + (de)serialization
  · rustfs-placement: HRW over cluster map; offset→chunk; deterministic dist/path
  · rustfs-mds: inode-range allocator; create/lookup/getattr/setattr/mkdir/
    rmdir/readdir/rename/unlink/link/symlink/xattr as single KV transactions;
    OpenLayout, GetClusterMap, CommitWrite, volume RPCs (stubs ok); gRPC server
  · VIDX versioning rows; PEND/DELQ rows defined

M1 ACCEPTANCE:
  · All metadata ops pass unit + integration tests against embedded KV
  · Cross-directory rename is atomic (transaction test)
  · readdir is a single range scan; lookup/getattr single get (perf assertion)
  · Crash test: kill MDS mid-op → KV consistent on restart
Effort: ~2.5 months.

Phase 2 — Data Plane / Chunk Service (M2) [Track B]

Deliverables:
  · chunk-format: self-verifying on-disk format; codec; golden vectors pass
  · disk-io: io_uring engine (per-disk rings, fixed buffers) + tokio::fs fallback
  · chunk-store: write_chunk/read_chunk/delete_chunk (chunk-key, byte-range);
    RS encode/decode reused; inline bitrot verify on read; heal reconstruction
    on corrupt/missing shard
  · packing: container manager (open/append/seal), Slice read path, per-container
    CMTA (dead_bytes), compaction GC; containers are large internal files reusing
    chunk-store + formula layout
  · placement integration: chunk_set/path/dist computed identically to MDS
  · PEND write-intent on new file_id; GC scrubber for orphans; DELQ chunk GC

M2 ACCEPTANCE:
  · Write a chunk, read arbitrary byte ranges back, bit-exact
  · Corrupt a shard on disk → read still succeeds via reconstruction; heal repairs
  · A client and the MDS compute identical locations for 10^6 random keys
  · EC partial read amplification < 1.2x for 4 KiB reads (LLD-002 target)
  · Pack 10^6 small files into containers; read each back via Slice; compaction
    reclaims space after deletes (no KV bulk data; metadata-only KV verified)
Effort: ~3.5 months (overlaps Phase 1 tail).

Phase 3 — Client + FUSE (POSIX MVP) (M3) ★ key usable milestone [Track C]

Deliverables:
  · rustfs-client: placement engine, cluster-map cache, metadata cache
    (version-based, close-to-open), retry/backoff, hedged reads, circuit breaker,
    chunk transport (gRPC first, io_uring TCP next)
  · rustfs-fuse: open/read/write/seek/close/stat/readdir/mkdir/rename/unlink/
    truncate/fsync; inline-small-file fast path
  · read-ahead engine (sequential + stride detection)
  · random-read path: full-chunkmap-on-open, fragment cache (TinyLFU),
    high-QD path; posix_fadvise honored
  · crash consistency: data-first then CommitWrite; PEND lifecycle

M3 ACCEPTANCE:
  · mount; dd 1 GiB write + md5 read-back; ls -la; cp; rename; rm
  · "one write, multi fseek/fread": fio randread > target IOPS; p99 within target
  · Sequential read approaches aggregate DSS bandwidth (read-ahead working)
  · A multi-process reader (simulated DataLoader) keeps reads off the MDS
    after open (lease-free version cache holds)
Effort: ~3 months.

Phase 4 — S3 Unified Namespace (M4) [Track C]

Deliverables:
  · s3-frontend: PutObject/GetObject/DeleteObject/ListObjects(V2)/HeadObject;
    multipart (stage parts → formula chunks on Complete); range GET; conditional
    requests; object tagging; bucket ops
  · bucket↔directory mapping (/buckets/<bucket>/<key>); ListObjects = readdir
  · versioning via VIDX

M4 ACCEPTANCE:
  · s3-tests conformance suite passes (core + multipart + versioning subset)
  · Data PUT via S3 is readable via the FUSE mount and vice versa (same inode)
  · ListObjects on a 10^7-object bucket is O(1)-per-directory (latency proof)
Effort: ~2 months.

Phase 5 — Block Storage (M5) [Track C]

Deliverables:
  · block-frontend volume mgr: EXTM extent map; thin provisioning; COW snapshots;
    clone; resize
  · NBD server (simplest, for testing) → iSCSI target (production) → NVMe-oF
  · WAL for write ordering; SYNCHRONIZE_CACHE/flush semantics; UNMAP/TRIM

M5 ACCEPTANCE:
  · Create volume; attach via iSCSI; boot a QEMU/KVM VM from it
  · Snapshot + clone + resize verified; data consistent after VM fsync
  · 4 KiB random write latency within target (LLD-001)
Effort: ~3 months.

Phase 6 — HA & Scale-Out (M6) [Track A + D]

Deliverables:
  · TiKV backend for rustfs-kv (clustered, distributed transactions)
  · Multiple stateless MDS behind LB / client round-robin; session failover
  · Cluster-map epoch changes → formula relocation → online rebalancing
    (copy-then-switch, rate-limited); node add/decommission; heal on node loss

M6 ACCEPTANCE:
  · Kill an MDS instance under load → clients fail over, no errors surfaced
  · Add a node → data rebalances online with < 10% client I/O impact
  · KV-backed metadata ops scale ~linearly with MDS instance count (mdtest)
Effort: ~2.5 months.

Phase 7 — Performance & Hardening (M7) [Track D]

Deliverables:
  · RDMA data-plane transport (READ pull for reads, WRITE push for writes);
    transport negotiation (RDMA → io_uring TCP → gRPC)
  · io_uring tuning: SEND_ZC, IOPOLL random-read rings, NUMA pinning
  · random-read: access-plan hinting API + framework shim; DSS hot-fragment tier
  · FUSE-over-io_uring (kernel ≥ 6.14) with fallback
  · (optional) GPUDirect Storage path

M7 ACCEPTANCE (targets from LLD-001/002):
  · 4 KiB random read p50 < 15 µs (RDMA), p99 < 50 µs
  · Random read > 500 K IOPS/client (RDMA, EC partial)
  · Shuffled-epoch training with access plan: > 5x vs no plan; GPU util > 90%
  · Sequential aggregate bandwidth scales linearly with DSS nodes
Effort: ~3.5 months.

Phase 8 — Production Readiness (M8 / GA) [Track D]

Deliverables:
  · Observability: metrics, tracing, dashboards for MDS/ChunkStore/client
  · IAM/policy integration; encryption at rest (reuse); audit
  · Ops tooling: cluster bootstrap, node lifecycle, backup/restore of KV
  · Kubernetes CSI driver (RWX FUSE, RWO block)
  · Soak (30-day) + chaos + Jepsen-style consistency tests green
  · Documentation (admin, API, tuning) — bilingual

M8 ACCEPTANCE: production readiness review passed; GA criteria met.
Effort: ~3 months (overlaps Phase 7).

5. Milestone Timeline (rough, team of ~8–10)

Quarter   Milestones in flight                        Usable outcome
────────  ─────────────────────────────────────────  ────────────────────────
Q1        M0 freeze; M1 metadata core; M2 start        —
Q2        M2 chunk done; M3 FUSE MVP                    ★ POSIX usable (fseek/fread)
Q3        M4 S3; M5 block start                         unified file+object
Q4        M5 block done; M6 HA                          block + clustered HA
Q5        M7 performance                                hits AI perf targets
Q6        M8 production readiness                       GA

Critical path to POSIX MVP:  M0 → M1 → M2 → M3   ≈ 7–9 months
Critical path to GA:         ≈ 18–24 months

Parallelism: Track A and B run concurrently after M0; Track C (FUSE) starts mid Phase 2 against mocked then real APIs; Track D testing runs from M1 onward.


6. Dependency Graph (build order)

            ┌── rustfs-kv ──┐
   M0 ──────┤               ├── rustfs-mds ──┐
            └── rustfs-inode┘                │
                    │                        ├── rustfs-client ── rustfs-fuse (M3)
            rustfs-placement (shared) ───────┤                 ├── s3-frontend (M4)
                                             │                 └── block-frontend (M5)
            ┌── chunk-format ──┐             │
   M0 ──────┤                  ├── chunk-store ──┘
            └── disk-io ───────┘   (reuses RS/bitrot/heal)
                                             │
                            rustfs-rdma ─────┴── perf (M7)
            TiKV backend ──────────────────────── HA (M6)

7. Testing & Validation Strategy

Level          What                                       When
─────────────  ─────────────────────────────────────────  ──────────────
Unit           per-crate logic, codecs, transactions       every phase, CI gate
Integration    MDS + ChunkStore + client end-to-end         from M2
Conformance    POSIX (pjdfstest), S3 (s3-tests),            M3 / M4 / M5
               block (fio, libiscsi/SCSI compliance)
Chaos          kill MDS, kill chunk node, partition,        from M2; gate for M6
               crash mid-write (PEND/GC correctness)
Consistency    Jepsen-style on KV txns + rename/create      M1 + M6
Performance    fio (IOPS/latency), IOR/mdtest (parallel),   M3 baseline; M7 targets
               real GPU training run (utilization)
Soak           30-day continuous load + fault injection     M8
Regression     CI perf + correctness gates per merge        continuous

8. Risk Register

Risk                                  Mitigation
────────────────────────────────────  ──────────────────────────────────────────
FUSE overhead caps POSIX throughput    FUSE-over-io_uring (M7); native client
                                       later; measure early in M3
io_uring kernel-version fragmentation  trait-abstracted disk-io with tokio::fs
                                       fallback; feature-detect
RDMA absent on commodity clusters      transport negotiation → io_uring TCP
KV (metadata) becomes the bottleneck   version-based client cache; bulk file
                                       data NEVER in KV (packing, §M2); TiKV
                                       sharding; benchmark in M1/M6
Billions of small files (AI datasets,  packing into large containers (M2); KV
node_modules, pip/conda, build tmp)    holds metadata + Slice pointers only;
                                       client-side write batching; raise inline
                                       cap only for ≤~1 KB; compaction GC sized
                                       for high churn (temp/dependency trees)
Formula placement vs rebalancing       cluster-map epochs + HRW (minimal movement);
                                       copy-then-switch online rebalance (M6)
Crash window between data & KV commit   PEND write-intent + GC scrubber; verified
                                       in M2/M3 chaos tests
Chunk on-disk format wrong → costly     freeze format in M0 with golden vectors;
                                       version the format header for future
ChunkStore fresh-build scope creep      reuse RS/bitrot/disk/heal as libs strictly;
                                       keep the wrapper thin (M2 scope guard)
EC partial-write amplification (block)  policy routing: replication for block/
                                       random-write; EC for write-once datasets
Packing compaction can lag under heavy  size compactor throughput to churn; per-
delete/temp-file churn                  container dead_bytes threshold triggers;
                                       rate-limit vs client I/O (like rebalance)

9. Definition of Done (per milestone)

M0  specs frozen, CI green on skeletons, Q1–Q6 recorded
M1  metadata ops correct + atomic + crash-safe on embedded KV
M2  chunk read/write bit-exact, self-verifying, heal-on-corrupt, formula parity;
    small-file packing (pack/read/compact) works; KV proven metadata-only
M3  POSIX mount serves the fseek/fread workload at target IOPS/latency
M4  S3 conformance + bidirectional file/object visibility + O(1) listing
M5  VM boots from block volume; snapshot/clone/resize correct
M6  MDS failover + online rebalance + linear metadata scaling
M7  AI perf targets met (RDMA latency, IOPS, planned-shuffle throughput)
M8  observability + CSI + soak/chaos/consistency green; PRR passed → GA

10. Team Shape (suggested, ~8–10 engineers)

2  Metadata (Track A): KV abstraction, MDS, placement
2  Data (Track B): chunk-format, chunk-store, packing, disk-io/io_uring
2–3 Clients (Track C): FUSE, S3, block frontends
1–2 Platform (Track D): RDMA, observability, HA, CSI
1  Test/Perf engineering (harness, conformance, chaos, benchmarks) — from day 1
1  Tech lead / architect (owns specs, cross-track APIs, design freeze)

11. First Actions (next 2 weeks)

1. Freeze the on-disk chunk format (Q3): write chunk-format spec + golden vectors.
2. Lock the protobuf wire contracts (MDS + ChunkStore + GetClusterMap).
3. Decide KV: stand up embedded redb/RocksDB behind TransactionalKv; spike TiKV.
4. Stand up the monorepo crate skeletons + CI + the test-harness shell.
5. Spike rustfs-placement (HRW + cluster map) and prove client/MDS parity on a
   unit test — it underpins the whole "formula" thesis and must be nailed early.

12. Summary

A greenfield, single-model unified storage: stateless MDS over a transactional
KV (metadata only — never bulk data), pure formula layout, client-direct data
plane, a thin chunk store reusing the proven RS/bitrot/disk/heal code, and
first-class small-file packing for the billions of small files in AI datasets and
code/dependency trees. Built in 8 milestones across 4 parallel tracks; POSIX MVP
in ~7–9 months, GA in ~18–24 months. No migration, no coexistence, no legacy
format — ever.

Chinese mirror: rustfs-dev-plan-ZH.md. Executes the canonical design RUSTFS-MD-DESIGN-005 with support from FEAT-001, LLD-001, LLD-002.