RustFS 文档 文档

RustFS Metadata — Final Greenfield Design (EN)

RustFS Metadata — Final Greenfield Design (EN)

Document:   RUSTFS-MD-DESIGN-005  (canonical / single source of truth)
Status:     DRAFT
Version:    1.0.0
Premise:    No historical data. Greenfield build. No migration, ever.
Folds in:   RUSTFS-MD-DESIGN-002 (core) + 004 (no-coexistence), and RETIRES the
            migration/coexistence content of CHG-001 and DESIGN-003 entirely.
Language:   English  (中文镜像见 rustfs-final-greenfield-ZH.md)

1. New Premise and What It Removes

There is no historical data. The system is built fresh. Therefore:

REMOVED ENTIRELY (no longer part of any plan):
  ✗ All migration (no rustfs-migrate tool, no re-stripe, no Option A/B choice)
  ✗ xl.meta in every form — not at runtime, not as a migration input, not at all
  ✗ Every legacy-decode open item (O1 old-bitrot, O5 reproduce-old-EcDist, ...)
  ✗ Any backward-compatibility scaffolding

RESULT: the design is purely forward-looking. The only data the system ever sees
        is data it wrote itself, in its own format.

2. Reframed Principle: "Minimal Change" → "Maximal Reuse"

The earlier goal "minimal change to the RustFS chunk" existed to avoid breaking running data. With no historical data, it is restated:

Reuse the proven, hard-to-get-right algorithms (Reed-Solomon, bitrot, disk I/O, heal reconstruction). Build a clean, formula-native chunk service on top of them. Carry over none of the object-store / xl.meta semantics.

REUSE as libraries (do not rewrite):  ecstore-core (RS encode/decode),
                                       bitrot reader/writer, disk I/O layer,
                                       heal reconstruction math, quorum logic.

BUILD fresh (clean, no legacy):        the chunk service wrapper, on-disk chunk
                                       format, formula placement, KV metadata.

DROP forever:                          SetDisks object semantics, xl.meta,
                                       object→set name hashing, DDir, EcDist
                                       storage, versioned-object-dir layout.

Because we define the on-disk chunk format from scratch, we make it self-verifying (inline streaming bitrot) by design — so metadata never needs to carry per-block hashes, and the decode path needs only the layout params.


3. Canonical Architecture (the whole thing, one page)

        CLIENT                         STATELESS MDS            TRANSACTIONAL KV
  ┌──────────────────┐          ┌────────────────────┐      ┌──────────────────┐
  │ metadata RPC stub │──control─▶ FS semantics        │──txn─▶ INOD / DENT /    │
  │ placement engine  │          │ inode-range alloc   │      │ VIDX / VOLU /    │
  │ cluster-map cache │◀─layout──│ placement engine    │      │ EXTM / PEND/DELQ │
  │ chunk I/O ────────┼─data─┐   │ (NO data path)      │      │ (TiKV / embedded)│
  └──────────────────┘      │   └────────────────────┘      └──────────────────┘
                            ▼
                   ┌──────────────────────────────────────────┐
                   │ CHUNK SERVICE (fresh wrapper)             │
                   │  · formula-native, chunk-key addressed    │
                   │  · self-verifying on-disk chunk format    │
                   │  reuses: RS · bitrot · disk I/O · heal     │
                   └──────────────────────────────────────────┘

3.1 Metadata = KV rows

INOD | inode_id          → InodeAttr { size, mode, uid, gid, times, kind, nlink,
                                       body, sys_meta }
                           body = Inline(bytes)               # ≤ ~1 KB only (rare)
                                | Slice{container_id,off,len} # small file → packed
                                | OwnLayout(Layout)           # large file → own chunks
DENT | parent_ino | name → { child_ino, kind }
VIDX | bucket | key      → [ { version_id, inode_id, mtime, delete_marker } ]
EXTM | volume_id | extent → ExtentEntry        (block volumes — the only index)
VOLU | volume_id         → Volume
CMTA | container_id      → ContainerMeta { layout, capacity, used, dead_bytes,
                                           sealed }            (small-file packing)
IALOC                     → batched inode counter
PEND | file_id           → write intent (crash GC)
DELQ | file_id           → delete tombstone (chunk GC)
MPUP/MPRT | upload_id...  → multipart staging

KV holds metadata only — never bulk file data. Inline is reserved for degenerate ≤ ~1 KB cases; it is NOT the small-file strategy (see §3.5). Billions of small files are normal for AI datasets and code/dependency trees (node_modules, pip/conda) — their metadata (inode + dentry rows) scales in the KV exactly as 3FS demonstrates, but their data must never go inline.

3.2 Layout = pure formula (one descriptor per file, inline in inode)

struct Layout {
    file_id: u64, ec: EcParams, stripe_unit: u32, chunk_size: u64,
    placement_seed: u64, cluster_map_epoch: u64,
}
// All locations computed, nothing stored per chunk:
//   set    = HRW(file_id, chunk_index, seed, cluster_map[epoch])
//   path   = deterministic(file_id, version, chunk_index)
//   dist   = deterministic(file_id, chunk_index)
//   hashes = inline in the shard (self-verifying)

3.3 Data plane = client ↔ chunk directly (MDS never in the path)

open():  client gets Layout + cluster-map epoch once → caches (layout is
         immutable for write-once files → cache for the file's lifetime)
read():  client computes location locally → ReadChunk(chunk_key, off, len)
         direct to the chunk server. Zero MDS traffic.
write(): client → WriteChunk direct to chunk server (RS encode inside the
         service); then MDS CommitWrite(size, mtime) — metadata only.

3.4 The clean primitives kept simple

Tiny files (≤~1KB) → inode.body = Inline (rare; degenerate optimization only).
Small files        → PACKED into containers (§3.5) — the answer for billions of
                     small files. inode.body = Slice{container_id, off, len}.
Large files        → inode.body = OwnLayout (the §3.2 formula path).
Versioning         → VIDX; each version its own inode + deterministic chunk paths.
Multipart          → stage parts; on Complete write formula chunks (re-stripe only
                     if part boundaries misalign — uncommon).
Block volumes      → EXTM extent map (thin/COW), the single enumerated index.
Crash consistency  → data-first, then KV commit; PEND intent + GC scrubber.

3.5 Small-file packing (first-class — billions of small files)

AI datasets (millions of tiny samples/tokens) and code/dependency trees (node_modules, pip/conda, build artifacts) routinely have hundreds of millions to billions of small files. Putting their data in the KV is fatal; sharding each tiny file with EC/replica wastes the per-shard fixed floor. The answer is packing, which composes cleanly with the formula layout:

Container  = a large internal object (e.g. 256 MiB) living in a reserved internal
             namespace. It IS a normal large file internally → its own OwnLayout,
             formula-placed, EC'd. Reuses the large-file chunk path entirely.

Write      = append the small file's bytes into the current open container for the
             writer; reserve (container_id, offset, length); commit the inode with
             body = Slice{container_id, offset, length}. Seal the container at its
             size limit → sealed containers are immutable → ideal for EC.

Read       = inode → Slice{container_id, offset, length} → container's Layout
             (cached; containers are few and hot) → compute chunk location →
             direct byte-range read from the chunk server. MDS off the data path.
             A small slice lands within one EC data-shard fragment → EC partial
             read at ~1x amplification (reuses LLD-002 random-read path).

Delete/GC  = tombstone the slice; track per-container dead_bytes (CMTA). A
             background compactor repacks live slices from sparse containers into
             new ones and frees the old container (log-structured GC).

Scale check: 1e9 × 32 KiB files → data packs into ~131 K containers (256 MiB each) that EC efficiently (~1.5x, not the small-file shard floor); metadata is ~1e9 inode + dentry rows in the KV (~hundreds of GB, distributed — the KV's job). This is the Haystack / SeaweedFS model, adapted to our formula layout.

Costs (accepted, see PLAN-001 Track B): compaction GC; client-side batching for fast bulk small-file creation; per-container append-offset reservation (bounded by using one open container per writer/session).


4. The Chunk Service: Reuse Algorithms, Build a Clean Wrapper

One design decision to record:

DECISION: build a FRESH thin ChunkStore that calls the reused libraries directly,
          rather than retrofitting the existing object-oriented SetDisks.

Why: SetDisks couples EC + distribution + xl.meta + object-level quorum/locking.
     With no legacy data, retrofitting it would carry object-store assumptions
     into the chunk layer. A fresh ChunkStore that composes ecstore-core (RS),
     the disk layer, and the bitrot/heal modules is cleaner and formula-native.

Cost: slightly more new code than a retrofit — but it is the thorough-iteration
      choice and avoids inheriting object semantics we will never use.

ChunkStore surface (chunk-key addressed, byte-range, self-verifying):
  write_chunk(chunk_key, offset, data, sync) -> ()
  read_chunk(chunk_key, offset, length)      -> bytes   (verifies inline bitrot)
  delete_chunk(chunk_key)                    -> ()
  // RS encode/decode, shard placement (by formula), bitrot, heal: reused libs

5. Build Plan (no migration — just build)

B1  rustfs-kv: TransactionalKv trait + embedded backend (redb/RocksDB) + TiKV.
B2  rustfs-inode: InodeAttr, Layout, DentValue, Volume, ExtentEntry types.
B3  rustfs-placement: shared HRW placement + cluster-map (used by MDS & client).
B4  rustfs-mds: stateless service, inode-range allocator, FS-semantics txns,
                gRPC (namespace + OpenLayout + GetClusterMap + CommitWrite + vol).
B5  chunk-service: fresh ChunkStore reusing RS/bitrot/disk/heal; self-verifying
                   on-disk chunk format; chunk-key addressing.
B5b packing: container manager (append/seal/Slice), small-file pack/unpack,
                   compaction GC. Reuses B5 (containers are large internal files).
B6  rustfs-client / FUSE: placement engine, cluster-map cache, direct chunk I/O,
                   read-ahead (sequential) + random-read path (LLD-002);
                   client-side small-file write batching into container appends.
B7  S3 + block (iSCSI/NVMe-oF) front ends over the same MDS + ChunkStore.
B8  HA: point MDS at TiKV → distribution + HA for free. No custom Raft.

No phases for seam/dual-write/cutover/migrate — those concepts are gone.


6. The Few Real Open Questions Left

Q1  KV backend: TiKV (Rust-native, clustered, recommended) vs embedded
    redb/RocksDB (single-node). Ship embedded first, add TiKV for HA.
Q2  inline cap (degenerate ≤~1 KB only — bulk small files go to packing, §3.5,
    NOT inline). Keep inline tiny so the KV never holds bulk data.
Q3  Define the on-disk chunk format: stripe_unit, inline-bitrot block size,
    chunk-key path scheme. (Free choice — no legacy constraint.)
Q4  Default storage policy by SIZE × ACCESS-MODE:
      tiny ≤~1KB → Inline; small → packed containers (§3.5); large ≥~1MiB → EC
      own-layout; random/partial-write (block volumes, mutable files) → replica
      (factor = m+1 so all tiers tolerate the same failures).
Q5  ChunkStore: confirm the fresh-wrapper decision vs SetDisks retrofit (§4).
Q6  Container sizing + packing locality (pack by writer/temporal, optionally by
    directory subtree) + compaction trigger thresholds (dead_bytes ratio).

These are forward design choices, not compatibility constraints.


7. Summary

Premise   : no historical data → greenfield → no migration, no xl.meta, anywhere.
Principle : reuse proven RS/bitrot/disk/heal; build a clean formula-native shell.
Metadata  : stateless MDS over a transactional KV; INOD/DENT/VIDX rows. KV holds
            metadata ONLY — never bulk file data.
Layout    : pure formula; zero per-chunk metadata; self-verifying shards.
Small files: PACKED into large containers (Slice pointer in the inode); billions
            of small files scale; never inline bulk data in the KV.
Data plane: client ↔ chunk directly; MDS never on the data path.
Chunk      : fresh thin ChunkStore reusing the hard algorithms; no object/xl.meta.
HA        : TiKV provides distribution + consistency; no custom consensus.
Coexistence/migration : none. The system only ever sees its own format.

This is the end state of the simplification: a single native model, the proven math reused as libraries, and nothing carried forward.


Chinese mirror: rustfs-final-greenfield-ZH.md. Canonical design; retires the migration/coexistence content of RUSTFS-CHG-001 and RUSTFS-MD-DESIGN-003.